Clark–West Test for Nested Forecasts: Formula, Example, and Limits
Learn why nested forecast models need an adjusted MSPE comparison, how the Clark–West statistic works, and what a one-sided result can and cannot establish.
In this guideWhy nested forecasts need a separate comparison
Short summary
The Clark–West test adjusts a squared-error comparison when one forecast model nests a smaller benchmark. Extra estimated parameters can add noise to the larger model’s forecasts even when they add no true predictive information. The adjustment changes the loss difference used for inference; it does not change either forecast or guarantee that the larger model is useful.
Why nested forecasts need a separate comparison
Suppose a restricted model forecasts next month’s return using a constant, while a larger model adds a valuation ratio. The restricted model is nested inside the larger one: setting the added coefficient to zero returns the benchmark. If that coefficient is truly zero, estimating it can still introduce sampling noise into the larger model’s predictions. Its observed mean squared prediction error (MSPE) may then be higher even when the benchmark and larger model have equal population predictive content.
The Clark–West adjustment targets this estimation-noise issue in an out-of-sample squared-error comparison of nested forecasts. It asks whether the larger model has predictive value beyond the restricted benchmark under that setup. This is a narrower question than the general comparison in the Diebold–Mariano guide, which compares paired forecast losses and is not automatically valid with the standard reference distribution for nested models. Clark and McCracken study the different asymptotic and finite-sample behavior of nested forecast accuracy and encompassing tests; Clark and West develop the adjusted MSPE procedure and its approximate inference. Clark and McCracken (2001)00071-9); Clark and West (2007).
Establish the nesting and preserve an out-of-sample design
Let the restricted benchmark forecast the same target (y_{t+h}) as the larger alternative. The larger model must contain the benchmark specification as a special case, typically by setting one or more added coefficients to zero. Merely being more complicated, using more variables, or having a better in-sample fit does not establish this nesting relationship.
For every forecast origin, generate both predictions using only information available then. Estimate the models on a chronological training window, issue forecasts over a later evaluation window, and use the same target, horizon, units, origins, and realized observations for both. A recursive or rolling re-estimation scheme can be appropriate, but state which one you use and retain the forecasts actually made at each origin. Repeatedly checking the evaluation window to choose predictors, transformations, or horizons turns that window into part of model selection rather than untouched evidence. The CAViaR guide illustrates why a tail-quantile forecast also needs to be evaluated against a loss suited to its target, rather than being folded into a mean-forecast comparison.
The Clark–West adjustment concerns parameter-estimation noise under a nested-model null; it does not repair look-ahead, mismatched samples, revised data used as if known earlier, or undisclosed model search. Report the initial estimation window, number of forecast origins, horizon, data vintage, and any rolling or recursive updates so a reader can tell what was genuinely out of sample.
Compute the adjusted loss difference
Write f₀,t+h for the restricted forecast, f₁,t+h for the larger forecast, and eⱼ,t+h = yₜ₊ₕ − fⱼ,t+h for forecast error j. For squared-error loss, the period-by-period Clark–West adjusted difference is
dCW,t+h = (e₀,t+h)² − [(e₁,t+h)² − (f₀,t+h − f₁,t+h)²]
Equivalently, it is the restricted model’s squared error minus the larger model’s squared error, plus the squared distance between their forecasts. The last term estimates the extra forecast noise associated with the additional fitted parameters under the null. Clark and West describe the adjustment as lowering the larger model’s sample MSPE by this forecast-gap term before comparing it with the restricted model. Clark and West (2007).
Average the adjusted differences over the common evaluation origins: mean(dCW) = (1/P) Σ dCW,t+h. If outcomes and forecasts are measured in basis points, each squared-error and adjusted-difference value is in basis-points squared. A positive adjusted mean points toward the larger model under the convention above. It is an adjusted comparison score, not the larger model’s realized MSPE and not a forecast to use for trading.
State the one-sided hypothesis before reading the result
The usual Clark–West question is whether the larger, nested model has a lower expected MSPE than the restricted benchmark. The null represents no incremental predictive improvement from the extra model component; the one-sided alternative favors the larger model. With the formula above, evidence for that alternative appears as a sufficiently positive mean adjusted difference relative to its standard error.
In the adjusted-loss notation, the working hypotheses are H₀: E[dCW] = 0 against H₁: E[dCW] > 0. That positive direction follows from defining restricted-model loss first. Reverse the subtraction and the sign reverses, so state the convention alongside the result.
The test statistic is often written as the sample mean of dCW divided by its estimated standard error. Clark and West recommend an approximately normal one-sided test for the settings they study, using a standard or autocorrelation-consistent standard error as appropriate. West (1996) develops inference for moments of smooth functions of out-of-sample forecasts and errors, including settings where forecasts depend on estimated parameters. The direction must be chosen before looking at results: a positive Clark–West statistic is not a license to switch to a two-sided or differently signed hypothesis after seeing the sample.
The reported p-value is conditional on the nested relationship, forecast construction, target, loss, inference method, and sampling assumptions. It does not give the probability that the larger model is true, nor does it establish that an added predictor causes the target to move.
Work through a four-period example
Take four hypothetical observations in basis points. The actuals are [0, 1, −1, 0], the restricted forecasts are [−1, 0, 0, 0], and the nested forecasts are [−1.5, 0.5, −0.5, 0.5]. The restricted model’s squared errors are [1, 1, 1, 0], with mean MSPE 0.75 bp². The larger model’s squared errors are [2.25, 0.25, 0.25, 0.25], also with mean MSPE 0.75 bp².
The squared forecast gaps (f₀ − f₁)² are [0.25, 0.25, 0.25, 0.25]. Applying the adjusted difference e₀² − e₁² + (f₀ − f₁)² gives [−1, 1, 1, 0] bp². Its sample mean is 0.25 bp². The raw MSPEs tie, while the adjusted mean is positive because the correction accounts for the larger model’s forecast-estimation noise.
These four periods illustrate the arithmetic only. They are far too few to estimate uncertainty credibly or establish statistical significance. The positive adjusted mean is not evidence of a real forecasting edge on its own, and the values are hypothetical rather than market observations. <!-- learn:illustration -->

Account for dependent forecast errors and horizons
For one-step forecasts, adjusted loss differences may still be serially dependent because the target, predictors, or estimation window evolve over time. For overlapping multi-step forecasts, adjacent origins share realized target periods, so their adjusted differences can be correlated by construction. An ordinary standard error that treats every origin as independent can understate uncertainty.
Estimate the standard error from the adjusted-difference series with an inference method that reflects the sampling design, such as a heteroskedasticity-and-autocorrelation-consistent long-run variance estimate when its conditions fit. Clark and West explicitly note autocorrelation-consistent errors for autocorrelated forecast errors. State the forecast horizon, covariance estimator, kernel and bandwidth if used, and any resampling or critical-value method. The block-bootstrap guide explains why resampling dependent time-series observations requires preserving their ordering structure; a generic bootstrap does not automatically implement the Clark–West null.
The required dependence treatment is design-specific. Longer horizons, persistent targets, repeated parameter estimation, and volatility changes can create dependence beyond the simple overlap pattern. Compare reasonable inference settings and explain them rather than selecting the one that gives the smallest p-value.
Treat normal critical values as an approximation, not a guarantee
Nested-model forecast tests can have nonstandard null distributions, especially when both the estimation and evaluation samples grow together. Clark and McCracken’s analysis shows that equal-accuracy and encompassing procedures for nested forecasts have different asymptotic and finite-sample behavior from the familiar non-nested comparison. Clark and West motivate an adjusted statistic with useful approximately normal inference in studied designs, but that label is not a universal finite-sample guarantee. Clark and McCracken (2001)00071-9); Clark and West (2007).
Small evaluation samples, many added parameters, misspecified benchmarks, data-driven predictor selection, and the choice between rolling and recursive estimation can all affect size and power. Where the sample or design differs materially from the paper’s settings, consider design-appropriate critical values or a carefully specified simulation/bootstrap procedure. Report the estimation and evaluation sample sizes and make clear which reference distribution produced the p-value.
Distinguish nearby tests and limit forecast claims
The Diebold–Mariano test asks whether two forecast loss series have equal expected loss under a general paired-loss setup. The Clark–West adjustment is tailored to squared-error comparisons where one model nests another and estimated extra parameters distort the raw MSPE difference under the null. Applying the adjustment to unrelated, non-nested models changes the question without justification. The original DM paper allows a broad range of loss measures and discusses forecast errors that need not be Gaussian or independent; it is a useful neighbor in the literature, not a synonym for the nested-model procedure. Diebold and Mariano (1995).
Neither method resolves researcher degrees of freedom from trying many targets, horizons, benchmarks, and predictor sets. If the final comparison emerged from a broad search, keep a record of the candidate family and use a method designed for the search-level question. White’s Reality Check and Hansen’s SPA guide covers family-wide inference for searched trading rules; the Clark–West test by itself does not make a selected forecast comparison confirmatory.
A significant one-sided result is evidence, under its stated assumptions, that the larger model improves expected squared forecast accuracy for the tested target and evaluation design. It does not show that the extra predictor is causal, that a signal will remain stable, or that a strategy using the forecast will earn positive returns. Forecast improvements may be too small to matter for decisions, and a model’s strongest gains may occur in periods that are costly to trade.
Trading profitability needs a separate rule, untouched evaluation data, and realistic spreads, fees, slippage, market impact, financing, borrow, and risk limits. Causal claims require an identification strategy suited to the question; a post-sample forecast comparison is not a causal design. State what the Clark–West result does establish and leave those other claims to evidence that actually tests them.
Report the design so the comparison can be reproduced
Name the restricted and larger models, identify the exact restrictions that make them nested, and define the target, horizon, squared-error units, estimation window, forecast-origin schedule, evaluation dates, and common-sample rule. Explain whether coefficients are re-estimated recursively or with a rolling window and how the information set and data vintages are preserved.
Report both raw MSPEs, the mean Clark–West adjusted difference, its sign convention, standard error, one-sided alternative, reference distribution or critical values, p-value, and dependence estimator. Also disclose added parameter count, sample sizes, predictor and horizon searches, benchmark selection, and whether the evaluation window influenced model design. That record lets readers distinguish an adjusted nested-model test from a general forecast comparison and judge its remaining uncertainty.
Common questions
Q1Does the Clark–West test replace the Diebold–Mariano test?
No. Clark–West addresses a nested-model squared-MSPE comparison. Diebold–Mariano is a general paired-loss framework, and its usual reference distribution should not be assumed to handle nested forecasts automatically.
Q2Does a positive adjusted mean prove the larger model is better?
No. It points toward the larger model under the stated sign convention, but inference also needs a suitable standard error, a credible out-of-sample design, and enough observations.
Q3Can I use Clark–West for models that are not nested?
Not by default. The correction is motivated by estimation noise under a nested-model null. For a non-nested comparison, choose a test matched to that forecast design and loss.
Q4Does a significant Clark–West result imply a profitable trading strategy?
No. It tests forecast accuracy for one target and evaluation design. Trading results also depend on the decision rule, execution, fees, financing, risk, and further out-of-sample evidence. Primary research - Clark and West, “Approximately Normal Tests for Equal Predictive Accuracy in Nested Models” (2007) - Clark and McCracken, “Tests of Equal Forecast Accuracy and Encompassing for Nested Models” (2001)00071-9) - West, “Asymptotic Inference about Predictive Ability” (1996) - Diebold and Mariano, “Comparing Predictive Accuracy” (1995)
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
Why does the Clark–West adjustment add the squared gap between the two forecasts?
Choose an answer to see the explanation