Forecast Encompassing Test: Does One Prediction Contain the Other’s Information?
Learn what a forecast-encompassing test asks, how forecast-error regressions reveal incremental information, and why the result differs from an RMSE ranking or an equal-accuracy test.
In this guideA forecast with a higher RMSE can still add information
Short summary
A forecast-encompassing test asks whether one forecast contains the useful predictive information in another. If forecast A encompasses forecast B under a stated test, adding B should not improve the relevant forecast combination. That is a different question from asking which forecast had the smaller average error. The regression and calculations below are teaching devices for point forecasts under squared-error loss. Formal encompassing procedures differ in their null restrictions, forecast design, and inference. A non-rejection is not proof that two models are equivalent, and a statistically useful forecast is not automatically a profitable trading rule.
A forecast with a higher RMSE can still add information
Suppose a benchmark predicts next month’s return and a second model predicts the same return. The benchmark may have the lower root mean squared forecast error (RMSE), yet its errors might be systematically related to the gap between the two predictions. If so, the second model could help correct the benchmark on some dates even though it performed worse on its own.
Forecast encompassing describes this information relationship. If A encompasses B, B does not contribute additional information that improves the chosen forecast combination, given the evaluation setup. If neither encompasses the other, both may contain distinct information. One forecast can also encompass the other in one direction without the reverse being true.
The idea grew from econometric model evaluation: a model should be judged by whether it accounts for relevant information available in a rival, not only by fitting its own historical sample. Chong and Hendry proposed a limited-information forecast-encompassing test that can use the competing forecasts without needing all of the other model’s internal data. Their paper also discusses the test’s derivation and drawbacks. {source:chongHendry1986ForecastEncompassing}
Match the forecasts before comparing their information
The test only has a clear interpretation if each pair refers to the same target, forecast horizon, origin, units, and information cutoff. For an evaluation origin t, let y be the later realized target, f_A the forecast from model A, and f_B the forecast from model B. Define the errors as e_A = y − f_A and e_B = y − f_B. A five-day return forecast cannot be paired with a one-day return, and a close-to-close forecast should not silently include a different overnight interval from its rival.
Use forecasts that could actually have been produced at the stated time. A chronological out-of-sample design estimates each model from information available at an origin, records its prediction, and later matches that prediction to the realized target. Keep the training window, re-estimation schedule, data vintage, missing-value rules, and forecast horizon. Replacing old data with revised values or selecting a model after looking repeatedly at the evaluation period changes what “out of sample” means.
Both series should use the same evaluation origins. If one model has missing predictions during volatile periods, calculate comparisons on the shared dates or explain a different design. When targets overlap, adjacent forecast errors can share realized observations; that affects uncertainty estimates even though the forecasts are correctly aligned. The Diebold–Mariano guide compares paired forecast losses, while the forecast-combination guide explains how weights are chosen. Neither makes a mismatched sample informative.
Define what “A encompasses B” means for the question
For point forecasts evaluated with squared error, consider a linear combination that starts from forecast A and can add the difference between B and A:
d_t = f_B,t − f_A,t f_C,t = f_A,t + w d_t
When w is zero, the combination uses A alone. A positive or negative nonzero w means the second forecast changes the combined prediction in the direction or opposite direction of the forecast gap. If the weight is restricted to the interval from zero to one, the result is a convex average; an unrestricted regression can estimate weights outside that range.
Under the corresponding population assumptions, A encompasses B when B contributes no incremental information to the forecast combination being studied. In the population regression with an intercept, the slope is β = Cov(e_A,t, d_t) ÷ Var(d_t), when Var(d_t) is positive; β = 0 therefore means no linear association after allowing for a constant error. Without an intercept, the moment E[e_A,t d_t] = 0 is equivalent to zero covariance when either E[e_A,t] = 0 or E[d_t] = 0. Some procedures also impose a separate zero-mean-error restriction. Other procedures test directional alternatives or make different assumptions about bias, estimated model parameters, and the information set.
This is a definition tied to a loss function and a class of combinations. It does not say that B is useless for every horizon, target, state of the market, nonlinear rule, or other loss. A test of linear encompassing can miss nonlinear information. State the forecast pair, combination class, loss, and null before interpreting a result.
Use a forecast-error regression as an intuition
One teaching regression is
e_A,t = α + β d_t + u_t
Here α captures a constant average error after the difference is considered, and β measures a linear relation between A’s error and the forecast gap. If the forecasts are unbiased under the chosen setup, the no-incremental-information restriction often includes β = 0. When a bias correction is allowed, β = 0 asks whether B adds linear information conditional on a constant offset. A joint test of α = 0 and β = 0 asks more: it also imposes the mean-error restriction.
The regression’s slope is not a universal test statistic. Original and later procedures differ in whether the regression is written with forecast errors, forecast levels, or a transformed difference; they may use one-sided or two-sided alternatives and distinct reference distributions. Harvey, Leybourne, and Newbold show that a natural least-squares encompassing test can have a null distribution that is not robust to nonnormal forecast errors, and they discuss alternative tests. Use a procedure whose assumptions match the forecast construction rather than applying an ordinary regression t-test automatically. {source:harveyLeybourneNewbold1998ForecastEncompassing}
<!-- learn:illustration -->

Work through a small hypothetical calculation
Assume six forecast origins, all measured in basis points. Let A’s forecast be zero at each origin, let B’s difference from A be the following values, and let the realized outcomes equal the listed A errors:
| Origin | A forecast | B forecast | Realized value | A error | Difference B − A |
|---|---|---|---|---|---|
| 1 | 0.0 | −4.0 | −1.6 | −1.6 | −4.0 |
| 2 | 0.0 | −2.0 | −1.2 | −1.2 | −2.0 |
| 3 | 0.0 | −1.0 | 1.5 | 1.5 | −1.0 |
| 4 | 0.0 | 1.0 | 0.5 | 0.5 | 1.0 |
| 5 | 0.0 | 2.0 | 1.3 | 1.3 | 2.0 |
| 6 | 0.0 | 4.0 | 1.5 | 1.5 | 4.0 |
The difference series has mean zero and sum of squares 42 bp². The mean A error is 2 ÷ 6 = 0.333… bp. The cross-product of the difference and demeaned A error is 16.4 bp². The least-squares slope with an intercept is therefore 16.4 ÷ 42 ≈ 0.3905. In this constructed sample, the fitted intercept is about 0.333 bp and the second forecast receives an additional linear weight of about 0.39 in the error regression.
This arithmetic does not establish that B truly adds information. With only six deliberately chosen observations, it supplies no credible p-value, standard error, or out-of-sample conclusion. Estimating and evaluating a combination on the same six observations also creates selection optimism. In practice, estimate any combination weight using a training segment, then assess the frozen rule on later observations or use a formal procedure designed for the forecast-generation method.
Separate encompassing from accuracy rankings and other forecast tests
An RMSE comparison summarizes the square root of each forecast’s mean squared error. A paired equal-accuracy test asks whether a selected loss difference has a nonzero mean. An encompassing test asks whether one forecast’s error still contains information related to the other forecast after the chosen combination or transformation. These questions can produce different answers. A forecast can rank second by RMSE and still improve a combination; a forecast with lower RMSE does not necessarily encompass its rival.
The Diebold–Mariano test focuses on equal expected predictive loss under its assumptions. Forecast encompassing is related to forecast combination, but it is not simply “pick the forecast with a smaller RMSE.” The Mincer–Zarnowitz regression instead tests a calibration relationship between realized outcomes and a forecast’s level and scale. These tests have different null hypotheses; one cannot substitute for another just because all use forecast errors or regressions.
If the question is whether an extra predictor improves a nested forecast model, use a test designed for that nested-model null. For the one-step-ahead nested-linear-model design they study, Clark and McCracken show that equal-accuracy and encompassing procedures have asymptotic and finite-sample behavior that differs from the familiar non-nested setting. {source:clarkMcCrackenNestedForecasts2001} The Clark–West guide explains its adjusted squared-error comparison. It is not a generic replacement for every encompassing test.
Choose inference for the forecast-generation process
The forecasts themselves may depend on parameters estimated from data. Their estimation uncertainty can affect the sampling behavior of an encompassing statistic. West studies forecast-encompassing tests when forecasts depend on estimated regression parameters; the appropriate procedure depends on the model and the way the forecasts were generated. {source:west2001EstimatedForecastsEncompassing}
Overlapping multi-step forecasts, persistent forecast differences, changing volatility, and repeated parameter estimation can make the test moments serially dependent or heteroskedastic. In those cases, an independent-observation standard error may be misleading. A heteroskedasticity-and-autocorrelation-consistent covariance estimate can be one option when its conditions fit, while some forecast designs require test-specific critical values; neither choice is automatic and neither alone corrects every source of dependence. Newey and West describe a positive semi-definite HAC covariance estimator. {source:neweyWest1987} State the horizon, kernel, bandwidth, estimation window, and any finite-sample correction.
Small samples, nonnormal forecast errors, outliers, breaks, and forecast instability can also change the size and power of a test. A failure to reject may mean the sample is too noisy to distinguish the forecasts; it does not confirm equivalence. If neither forecast encompasses the other, both may add information, but the evidence can also reflect low power or a misspecified linear combination.
Translate a statistical result into a trading question carefully
For a technical trading signal, the predicted target might be next-session return, realized volatility, direction, or a tail quantile. The encompassing setup must match the target. A squared-error combination for expected returns does not automatically apply to a probability forecast, volatility forecast, or value-at-risk estimate; each has a different loss and suitable comparison methods.
Even a reliable improvement in a forecast statistic does not demonstrate a tradable edge. A position rule adds choices about thresholds, sizing, leverage, execution timing, spread, fees, market impact, funding, borrow, and risk limits. Those choices can alter realized profit and loss, and selecting them after testing many forecast pairs creates another multiple-testing problem. Keep model evaluation separate from the trading strategy’s untouched final evaluation.
Report the target and horizon, forecast origin and common dates, estimation window, data vintage, loss and combination class, null and alternative, test implementation, standard-error method, and any model-selection steps. Show the individual loss metrics as context, then explain the narrower incremental-information result. That record lets another researcher reproduce what “encompasses” meant in the particular study.
A practical sequence for an encompassing analysis
- State whether forecasts are nested or non-nested, and choose the target, horizon, loss, and candidate combination before inspecting the evaluation results.
- Generate both forecasts using only information available at each origin. Preserve the matching sample, data vintage, model estimates, and any missing-date rule.
- Define the null precisely: for example, whether it concerns zero incremental linear information, a bias restriction, or a one-sided weight restriction.
- Select the published statistic and reference distribution that fit the forecast-generation process, including estimated parameters, overlapping horizons, and error dependence.
- Report the regression or moment estimate, uncertainty, sample size, forecast accuracy context, and sensitivity to reasonable inference choices.
- Treat any discovered combination weight or model selection as a new hypothesis. Freeze it and evaluate it on data that did not choose it before making a trading claim.
Common questions
Q1Does the forecast with lower RMSE always encompass the other?
No. RMSE ranks each forecast’s average squared error. Encompassing asks whether a rival contributes information that improves a specified combination. The two results can differ.
Q2Does failing to reject prove that two forecasts are equivalent?
No. A non-rejection can reflect limited power, noisy errors, an unsuitable test, or a sample too small to distinguish the forecasts. Equivalence needs a separately defined margin and test.
Q3Can a significant encompassing result prove a trading strategy is profitable?
No. It is evidence about a forecast relationship under its target, sample, loss, and inference method. Trading profitability also depends on a pre-specified decision rule and net execution results on genuinely held-out data. Primary research - Chong and Hendry, “Econometric Evaluation of Linear Macro-Economic Models” (1986) - Harvey, Leybourne, and Newbold, “Tests for Forecast Encompassing” (1998) - Clark and McCracken, “Tests of Equal Forecast Accuracy and Encompassing for Nested Models” (2001)00071-9) - West, “Tests for Forecast Encompassing When Forecasts Depend on Estimated Regression Parameters” (2001) - Newey and West, “A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix” (1987)
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
What does the null that forecast A encompasses forecast B generally represent?
Choose an answer to see the explanation