Conformal Prediction: Calibrated Forecast Intervals

A point forecast tells you what the model expects. It does not tell you how much risk you are taking. Conformal prediction adds a calibrated range around that number, built from the errors the model actually made on data it never saw — so coverage can be measured instead of assumed.

Calibrated Intervals Coverage Held-Out Errors Adaptive Widths

TL;DR: A point forecast is not a plan. Split conformal prediction wraps any forecasting model in a prediction interval built from the errors it made on a held-out calibration set, so a 90% interval targets 90% long-run coverage — measured, not assumed. Basic split conformal gives every prediction the same width, so peak versus off-peak demand needs an adaptive method such as conformalized quantile regression. And no interval replaces evaluation: check coverage by horizon, segment, and future period before anyone plans capacity around it.

Visual Summary

The Problem

A point forecast is often reported as if it were a fact. The model says “Expected demand: 1,200 units”, and that single number moves into a plan, a purchase order, and a staffing roster. The operational question underneath it is different: what range should we plan for?

The usual fix is to assume an error distribution — normally distributed, a standard deviation, a backtested average error — and derive a range from it. That assumption is the weak point. It is rarely checked against the errors the model actually makes, and a single average error hides the fact that peak periods and quiet periods do not behave the same way.

The result is a number that looks precise and a plan that inherits every unexamined assumption behind it. When the miss lands, nobody can say whether the model was wrong or the range was never calibrated.

The Approach

Conformal prediction does not require a special model. It takes a model you already trust for its point estimate and wraps it in a range built from measured errors. The basic split conformal workflow has four steps.

Step 1 — Train the model

Fit the forecasting model on historical training data. Conformal prediction does not change this model or require it to satisfy any distribution assumption.

Step 2 — Hold out calibration data

Reserve a separate set that was not used to fit the model. For forecasting, keep it chronologically later than the training data so the interval is built from errors the model genuinely could not have seen.

Step 3 — Measure calibration errors

Score the held-out set as actual minus predicted and take the error quantile you need, typically the upper residual quantile at the coverage level you care about.

Step 4 — Create the future interval

Apply that calibrated error quantile to each future prediction to produce an interval at the target coverage, such as 980 to 1,420 units around a 1,200 unit forecast.

The interval itself is then the point prediction plus or minus the calibrated error quantile:

ŷ ± q(calibration errors)

For a forecast of 1,200 units with a calibrated 90% interval, that becomes 980 to 1,420 units. The important detail for time series is step two: calibration data must be held out chronologically, not split at random, because a random split leaks future information into the interval.

What 90% Coverage Actually Means

A 90% interval is a coverage target, not a promise about one forecast. Across many comparable future predictions, approximately nine out of ten actual outcomes may fall inside their interval. Any given forecast can still land outside it, and that is expected behaviour rather than a defect.

Coverage is a long-run marginal property. It says nothing about a single prediction, and it does not hold automatically inside every small subgroup. That is why the interval is only useful if you actually test it:

Test by horizon

Check coverage separately for each forecast horizon you actually use. An interval that averages out to target while failing at the horizon that drives planning is not calibrated for that decision.

Test by segment

Verify coverage inside the business segments that matter. A method can hold on aggregate and miss badly on a small but critical subgroup, which is where the cost concentrates.

Test by time period

Re-test on realistic future periods rather than a single historical slice, then re-check after deployment. Calibration is a property that can decay as the data shifts.

If coverage holds on average but breaks down at long horizons or in a key segment, the interval is not usable for the decision you built it for.

Where Basic Conformal Breaks

Split conformal prediction has one limitation that matters operationally: it applies a single calibrated error quantile to every prediction, so every interval comes out the same width. A calm Tuesday and a peak-demand Friday receive identical ranges even though their uncertainty is not identical.

Basic split conformal

One calibrated error quantile creates one common interval width. Simple and cheap, and the right default when uncertainty is genuinely stable across the range you forecast.

Adaptive conformal

Conformalized quantile regression or locally scaled conformal scores use a conditional uncertainty model, so intervals stay narrow in stable periods and widen when conditions are volatile. More work to build and to validate.

This is the distinction worth being precise about. Standard split conformal gives one common width; an adaptive method gives intervals that respond to local conditions. Neither is universally better — the adaptive approach costs more to build and to validate. The point is only that you should not assume a conformal wrapper widens intervals during peak periods when the base method does not.

Outcome

Once a calibrated interval is attached to the forecast, the conversation changes from whether a number is right to how much risk a decision can absorb. Capacity gets a staffing range instead of a single target, inventory gets a buffer sized to the interval, and the downside case becomes a scenario you planned for rather than a surprise.

It is equally important what a calibrated interval does not do. It does not make a weak model strong, it does not repair a flawed evaluation, and it does not remove the need to monitor whether coverage still holds once the data shifts. Backtest coverage on realistic future periods, by horizon and by segment, and keep checking it after deployment.

Key Takeaway

Design insight: Conformal prediction turns forecast uncertainty from an assumption into a measurement. It wraps an existing model in an interval built from errors on held-out calibration data, so a 90% interval targets 90% long-run coverage instead of promising certainty. Coverage is marginal rather than per-forecast, and basic split conformal applies one width everywhere. When uncertainty varies by context, adaptive methods such as conformalized quantile regression widen intervals where they should. Always backtest coverage by horizon and segment before decisions depend on it.

FAQ

What does a 90% conformal prediction interval actually mean?

It is a coverage target, not a guarantee for any single forecast. Across many comparable future predictions, roughly nine out of ten actual outcomes may fall inside their interval. Coverage is a long-run marginal property, so an individual forecast can still land outside its interval, and a method can hit 90% on average while missing badly inside a particular horizon or business segment.

How do you build a prediction interval with split conformal prediction?

Fit the forecasting model on training data, then hold out a separate calibration set that the model never saw. Score that set as actual minus predicted and take the error quantile matching your coverage target. Apply it to each future prediction as prediction plus or minus that quantile. For forecasting, the calibration data must be chronologically later than the training data, because a random split leaks future information into the interval.

Why do basic conformal intervals have the same width everywhere?

Split conformal uses a single calibrated error quantile for every prediction, so one common width is applied across the whole forecast range. That is a reasonable default when uncertainty is genuinely stable, but it understates risk in volatile periods and overstates it in calm ones. When uncertainty varies by context, you need an adaptive method such as conformalized quantile regression or locally scaled conformal scores.

How do you check that conformal prediction intervals are well calibrated?

Backtest coverage on realistic future periods rather than one historical slice, and check it separately by forecast horizon and by business segment. Then re-check after deployment, because calibration decays as data shifts. If coverage holds on average but breaks down at the horizon that drives planning, the interval is not usable for that decision.

Comments

Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding Coding Agents Open Source AI Aider Cline Continue OpenHands Goose Try & Buy Frameworks Machine Learning Feature Engineering MLForecast Time Series Decomposition Forecasting LightGBM XGBoost Catboost Clustering Segmentation NLP LLMs Web App R Markdown SQL Oracle DB SAS-Guide SAS E-Miner Dataiku BigQuery GCP Python R CRISP-DM Hypothesis Testing ANOVA Data Analytics Dimensionality Reduction Recommendation System Network Analysis Geospace Analysis Spatial Data Embedding Sampling Techniques Decision Rules Data Storytelling CVM Churn Fraud Detection Sentiment Analysis Topic Modeling IBM Watson PowerBI Looker Studio VBA Statistical Learning Ensemble Modeling Stacking Cross-Validation Profiling ABT Construction Plumber Tidyverse Shiny Prophet Deep Learning Scikit-Learn JSON SAS Programming Git VS Code CSS Styling Automated Reporting Outlier Detection Temporal Clustering Startup Survival Pre-Valuation Modeling K-Means Decision Trees Data Science Predictive Modeling SVM LDA Text Classification Weight Prediction Pattern Recognition Real-Time Detection Community Detection Pipeline Automation Data Quality Checks Data Reliability Specification Mapping Business Strategy Marketing Campaigns Try & Buy Frameworks KPI Dashboards Network Quality Sales Analytics Mentoring Statistics Lecturer Remote Work Hybrid Work Consulting Contract Full-Time Freelance Sofrecom Orange Group Tunisia Telecom Kiota Intelligence VC Analytics Series A Prediction Production ML Applied AI Prompt Engineering Business Forecasting Decision Systems Graph Analytics Household Detection Multi-SIM Detection FTTH Forecasting Audit Extraction Infrastructure Classification Pydantic GPT-4 OpenAI API Base64 Classification Zindi Codementor LAAS-CNRS ESSAI MIT xPRO Tunisia ML Competition Cell Tower Analysis Uber Logistics Uber Cape Town Necessary Condition Analysis Behavioral Signals Spike Smoothing Observation Unit Design Dendrogram Ward Clustering VIF Target Encoding Coding Agents Open Source AI Aider Cline Continue OpenHands Goose Try & Buy Frameworks