DLinear: Make a Simple Forecasting Baseline Fail First
One of the healthiest lessons in forecasting: a simple model can be a very strong baseline. DLinear made that lesson hard to ignore, because it reaches competitive accuracy with decomposition and linear projection — and no attention blocks at all.
TL;DR: More architecture does not automatically create more forecast value. In time series forecasting, complexity is a cost you pay in latency, tuning effort, and operational burden, so it should solve a proven limitation rather than a fashionable one. Start with a statistical benchmark, then a decomposed linear model such as DLinear, and only keep the deeper architecture when a fair point-in-time backtest shows a clear, stable advantage by horizon. If the simple baseline has not failed yet, you do not yet have a forecasting problem that requires a transformer.
Visual Summary
The Problem
One of the healthiest lessons in forecasting is this: a simple model can be a very strong baseline. DLinear made that lesson hard to ignore, because its design is almost embarrassingly straightforward.
No attention blocks. No deep transformer stack. No architecture theatre. That is not an argument that transformers are useless for time series — it is an argument that complexity is not automatically a source of forecast accuracy.
The Approach
At a high level, DLinear does three things: it decomposes the historical series into a trend component and a residual component, projects each component from the input window into the forecast horizon, and adds both forecasts back together.
Lower capacity
Fewer parameters may reduce overfitting when the available data is limited.
Direct projection
The model maps historical windows straight to forecast horizons, without an attention mechanism in between.
Clear baseline
A simple structure is easier to inspect, debug, and explain when a forecast looks wrong.
Efficient operation
Lower tuning effort and lower inference cost, which matters when forecasts run on a schedule for many series.
That is also why DLinear is useful as a baseline rather than as a universal winner. My practical sequence is: start with a simple statistical benchmark, add a linear or decomposed-linear baseline, compare stronger machine-learning or deep-learning models fairly, and keep the complexity only when it delivers stable value across backtests.
When More Capacity Earns Its Place
Transformers can still be the right call. They become valuable when the problem contains enough signal to justify higher model capacity — so it is worth being precise about when that is.
Consider more complex models when
Many related series share signal, long-range dependencies matter, multivariate inputs add real context, interactions are complex, and there is enough data to support the extra capacity.
Use caution when
Data is limited, gains are unstable across folds, latency or tuning budgets are constrained, the simpler baselines are already strong, or operational complexity is hard to justify.
For many forecasting problems the real gains come from somewhere else entirely: valid point-in-time features, external drivers, better data quality, sensible forecast horizons, robust backtesting, and clear business constraints — not from adding attention by default.
Outcome
A model comparison is only useful when the forecast boundary is fair. So every candidate runs on the same rolling-origin evaluation: the same past history, the same forecast origins, the same holdout, and no future information crossing the origin. I compare error by horizon, stability across folds, latency, and cost — not one headline metric.
After that, the decision is rarely "which model is more advanced?" It is "does the more complex model improve the forecast enough to justify the added cost, latency, tuning effort, and operational complexity?" Simple models do not always win. But they deserve the opportunity to.
Key Takeaway
Design insight: More architecture does not automatically create more forecast value. In time series forecasting, complexity is a cost you pay in latency, tuning effort, and operational burden, so it should solve a proven limitation rather than a fashionable one. Start with a statistical benchmark, then a decomposed linear model such as DLinear, and only keep the deeper architecture when a fair point-in-time backtest shows a clear, stable advantage by horizon. If the simple baseline has not failed yet, you do not yet have a forecasting problem that requires a transformer.
Related
RevIN: When a Neural Forecast Learns Only One Scale →Zero-Shot Forecasting Changes the Baseline →Foundation Models Raise the Baseline →Turning Decomposition Into a Forecasting Strategy →The Hardest Forecasting Decision Is Which Model to Use →Seed Sensitivity Is Part of Model Selection →Route Structurally Inactive Series Before Forecasting →MLforecast Forecasting Pipeline →Conformal Prediction: When the Model Is Uncertain →StatsForecast, MLForecast, NeuralForecast: One Ecosystem →A Strong AutoML Baseline Can Beat Hand-Tuned Models →An LLM-Derived Feature Is Still a Feature →Semi-Automatic Forecasting Pipeline →
FAQ
What is the key takeaway from "DLinear: Make a Simple Forecasting Baseline Fail First"?
More architecture does not automatically create more forecast value. In time series forecasting, complexity is a cost you pay in latency, tuning effort, and operational burden, so it should solve a proven limitation rather than a fashionable one. Start with a statistical benchmark, then a decomposed linear model such as DLinear, and only keep the deeper architecture when a fair point-in-time backtest shows a clear, stable advantage by horizon. If the simple baseline has not failed yet, you do not yet have a forecasting problem that requires a transformer.
What is DLinear and how does it work?
DLinear is a linear neural forecasting architecture for time series. It decomposes the historical series into a trend component and a residual component, projects each component from the input window into the forecast horizon with a linear layer, then adds both forecasts together. It uses no attention mechanism and no recurrent state, which makes it a very cheap and very interpretable baseline.
Is a linear model like DLinear really competitive with transformers for forecasting?
Often, yes, on a per-parameter or per-unit-cost basis. A lower-capacity model may reduce overfitting when data is limited, the direct projection is easy to inspect, and tuning and inference are cheap. But a simple model is not universally better: transformers can earn their place with long-range dependencies, many related series sharing signal, rich multivariate inputs, complex cross-variable interactions, and enough data to support the extra capacity. The only fair answer comes from a rolling-origin backtest on the same origins and the same holdout, compared on error by horizon, stability across folds, latency, and cost.
Who wrote this and what is it about?
This was written by Mahmoud Trigui, Senior Data Scientist. DLinear shows why a decomposed linear model is a strong time series baseline, and when extra model capacity is actually worth its cost.