Giacomini–White Test: Conditional Forecast Accuracy
Learn how the Giacomini–White test checks whether forecast accuracy changes with known conditions, how to choose instruments, and what its result can show
In this guideAverage accuracy and conditional accuracy answer different questions
Short summary
The Giacomini–White (GW) framework asks whether the relative loss of two forecasts can be related to information available when the forecast is made. It can reveal differences hidden by an average score, but its answer depends on the forecasts, evaluation window, loss function, and conditioning variables chosen in advance.
Average accuracy and conditional accuracy answer different questions
Suppose forecast A has lower average loss than forecast B over a test sample. That average may hide a useful pattern: A may do better in quiet markets while B does better when volatility is high. If the question is which forecast performed better on average, an unconditional comparison is appropriate. If the question is whether information known at the forecast date helps predict which forecast will perform better, the target is conditional predictive ability.
Giacomini and White formulate this as a comparison of forecasting methods under a specified loss function and information set. Their framework can be used with point, interval, probability, or density forecasts when the loss is defined for the forecast task. The method includes the estimation procedure used to generate each forecast, so a rolling window, fixed training sample, or weighting rule is part of the object being evaluated—not a harmless detail to change after seeing results (Giacomini and White, 2006).
The distinction is related to, but not identical to, a test for a structural break. A conditional test asks whether relative forecast loss is systematically associated with chosen information. A break test asks whether parameters or a relationship changed at an unknown or specified time. The Diebold–Mariano guide covers an unconditional average-loss comparison; a conditional question needs explicit conditioning information.
Align forecasts, outcomes, and information first
Let (Y_{t+1}) be the target observed after forecast origin (t). Forecasts ŷA,t+1 and ŷB,t must refer to the same target, horizon, units, and information cutoff. For a lower-is-better loss (L), define the loss difference as
dₜ₊₁ = L(Yₜ₊₁, ŷA,ₜ₊₁) − L(Yₜ₊₁, ŷB,ₜ₊₁)
With this sign convention, a positive value means A incurred more loss and favors B on that origin; a negative value favors A. Keep one row per forecast origin and score both forecasts against the same realization. Preserve the forecasts as they were available at the time, along with the data vintage and model version. Do not rebuild historical forecasts using revised inputs or information that arrived later.
The conditioning variables must also be known at the forecast origin. Examples include a lagged volatility estimate, a preannounced event indicator, a lagged loss difference, or a market-state variable computed only from past observations. A variable measured after (t) cannot explain which forecast would have been better at (t+1) without look-ahead bias.
Match the score to the target. Squared error is one choice for a conditional-mean forecast, but it does not automatically suit a volatility, quantile, or density forecast. If one model predicts variance and another predicts standard deviation, first express them as forecasts of the same quantity. A test cannot repair a mismatch in the question being scored.
For volatility forecasts, score both against the same realized-variance definition and sampling interval. If one forecast covers intraday prices while another also targets overnight returns, a loss gap may reflect different targets rather than forecasting skill. Decide how closed-market moves, sampling frequency, missing observations, and the realized measure will be handled, then apply those choices consistently. If the realized proxy is noisy, that measurement error affects both loss scores and belongs in the interpretation.
State the conditional null and choose test functions
Let (mathcal{G}_t) represent the information used to condition the comparison. The idealized null is
H₀: E[dₜ₊₁ | 𝒢ₜ] = 0
It says that the expected loss difference is zero given the specified information. The one-step GW test turns this conditional statement into a set of moments using a preselected vector of test functions (h_t), each measurable using information available at time (t):
H₀,ₕ: E[hₜ dₜ₊₁] = 0
A constant in (h_t) includes the average loss difference. Other elements can represent a market state, a lagged relative loss, or a nonlinear transform chosen before examining the test results. For example, (h_t=(1,S_t)'), where (S_t) is a binary high-volatility indicator known at (t), tests both a constant moment and whether the loss difference is associated with that state.
A finite set of test functions tests only those moments. Rejecting says at least one selected moment is inconsistent with zero under the test assumptions; it does not identify a universally best model or prove that every possible state variable matters. Failing to reject does not prove equal conditional performance. A test using a constant alone is an unconditional moment comparison, not a broad search over every possible condition.
Choose the loss, instruments, transformations, sample, and forecast horizon before inspecting which model wins. Adding many state indicators, thresholds, and lags after seeing the results creates another search problem. The conditional test does not correct for that selection by itself.
Fix the timing used to construct each state variable. If “high volatility” means the top 20 percent of volatility in the full sample, future observations help set the threshold for past origins. Estimate a threshold only from information available at each origin or use a rule announced before that date. For a continuous state variable, define its units and any centering or scaling in advance. Several nearly identical indicators can make moment conditions redundant and the covariance matrix hard to invert, so use the smallest set that has a clear interpretation.
Build the Wald statistic and interpret its reference distribution
For (q) test functions, form the vector (g_{t+1}=h_t d_{t+1}). Its sample mean is
ḡ = (1/n) Σₜ gₜ₊₁
The one-step GW statistic has the Wald form
GW = n ḡ′ Ω̂⁻¹ ḡ
where (hatOmega) estimates the covariance matrix of the moment vector. Under the null and the paper's regularity conditions, the statistic has an asymptotic chi-square reference distribution with (q) degrees of freedom. The number of test functions therefore affects the reference distribution and can affect power. Adding many weak or redundant instruments may make the covariance estimate unstable without adding useful information.
In the one-step setup, the null implies a martingale-difference structure for the moment vector, which allows the original test to use a sample second-moment estimate. For other horizons, overlapping outcomes, or departures from that setup, the variance estimator must match the test's assumptions and dependence structure. Follow the formulation for the chosen implementation rather than automatically copying a HAC bandwidth from another test. General HAC covariance methods, including the estimator of Newey and West (1987), are tools for settings that require them, not a universal bandwidth prescription for GW. A result is only as credible as the covariance estimate and forecast design behind it.
Also check whether Ω̂ is estimable from the selected moments. Nearly duplicate test functions, an indicator with very few observations in one state, or too many interactions can make the matrix singular or unstable. Its inverse can then change sharply after a small data revision and distort the test statistic. Each function should have a reason tied to a condition or forecast-performance difference; report their count and moment covariance. Dropping an unhelpful function after looking at the result is another selection step and should be disclosed.
The p-value is evidence against the stated moment restrictions under the reference distribution. It is not the probability that A or B is the true model, and it is not the chance that a trading rule will be profitable. Report the statistic, degrees of freedom, p-value, and the exact moments tested so readers can tell what a rejection means.
<!-- learn:illustration -->

A two-state example shows what the average can hide
Consider eight hypothetical one-step origins. Let (S_t=0) mark a calm state and (S_t=1) a high-volatility state. Suppose the loss differences (d_{t+1}=L_A-L_B) are −2, −2, −2, and −2 during the four calm origins, and +2, +2, +2, and +2 during the four high-volatility origins. Negative differences favor A; positive differences favor B.
The average across all eight origins is zero, so this small example shows no unconditional loss difference. The calm-state average is −2 and the high-volatility average is +2. The relative ranking reverses with the state. A test with (h_t=(1,S_t)') includes a constant moment and a high-volatility moment that can reflect this pattern.
These eight values are a teaching illustration, not enough observations for a reliable inferential result. In real data, the test must account for how the forecasts were estimated, whether the state variable was fixed in advance, how many origins are available, and how the loss-difference moments vary over time. A visual state split alone is not evidence that the pattern will persist.
Treat the estimation window as part of the method
The original GW asymptotics preserve forecast-estimation uncertainty by keeping the maximum estimation-window size finite while the number of out-of-sample forecasts grows. Rolling windows and a fixed estimation sample fit that core setup. The window size and any data-dependent window rule are part of each forecasting method and should be set and reported as such.
This matters when the forecast models are re-estimated. A comparison that ignores parameter estimation may miss uncertainty created by the way each method learns from past data. GW can accommodate forecasts from nested or nonnested models under its stated conditions, but this is not a guarantee for every expanding-window implementation or every estimator. In the original one-step framework, a conventional expanding window is outside the finite-window assumption. Use a result derived for the actual forecast construction if the setup differs.
The test also does not choose a suitable window for you. A short window may adapt to recent conditions but use fewer observations; a long or expanding window may stabilize estimates while retaining stale relationships. Compare these choices through a preplanned evaluation, and include window selection itself in the recorded procedure. Giacomini and Rossi (2010) develop separate tests for how relative forecast performance evolves in unstable environments; that time-path question is related but not the same statistic as the basic GW conditional test.
Nor does a conditional rejection by itself give a trading rule that selects the right forecast in real time. The instrument may be associated with the loss difference in the evaluation data without yielding a stable or actionable switching rule. To use current information for forecast selection, define that rule with past observations and evaluate it on a later, untouched sequence of origins, including its own estimation and decision uncertainty.
Separate GW from DM, nested-model tests, and instability tests
The Diebold–Mariano test asks whether two forecasts have the same average loss over the evaluation sample. GW asks whether selected information is associated with their relative loss, while allowing an unconditional comparison as a special moment restriction. A model can have the same average score yet have condition-dependent relative performance, as the two-state example shows.
If the forecasts come from nested models and the task is specifically an equal-MSPE comparison, a nested-model method such as the Clark–West adjustment addresses a different null and estimation-noise issue. If many trading rules or forecasts were searched and the question is whether any candidate has superior predictive ability, use a family-level procedure such as White's Reality Check or Hansen's SPA test. A conditional test on one selected pair does not undo a broader search.
Finally, evidence that relative accuracy changes over time does not establish a particular break date. A test for local instability or reversals may be better when the research question concerns the path of relative performance rather than its association with chosen instruments. Giacomini and Rossi (2010) study such forecast comparisons. Choose the method to match the question instead of treating every rejection as evidence for a regime-switching trading strategy.
Account for low power and repeated testing
Conditional tests can need substantial data. Splitting a sample into calm and volatile states reduces the observations supporting each comparison. Adding more functions increases the number of moments and the degrees of freedom. If several conditions are only weakly related to relative loss, the test may have little power even when conditional differences exist.
McCracken's Federal Reserve working paper studies the existence, size, and power of the GW test in a simple quadratic-loss example. Its simulations find that the test can be accurately sized while power is typically low in the settings examined (McCracken, 2020). That result is a warning about interpretation, not a universal numerical forecast for every application. A non-rejection can reflect limited information or low power; it should not be described as proof that the forecasts are interchangeable.
Repeatedly trying alternative state definitions, thresholds, horizons, loss functions, and evaluation dates inflates the chance of a chance finding. Record all candidate tests and separate exploratory analysis from a pre-specified confirmation sample. If the goal is to claim that one among a family of strategies is predictive, apply an appropriate multiple-testing or data-snooping method to that family. More conditional instruments do not automatically create more convincing evidence.
When several forecast models remain under comparison, a family-level procedure can answer a different question from the pairwise conditional test. The model confidence set guide explains how an iterative comparison can retain a set of models that are not distinguishable at the chosen level; it does not replace pre-specifying the conditions in a GW analysis.
Report the design so another reader can reproduce it
Name both forecast methods, how each is estimated, whether the models are nested, the target and horizon, and the evaluation dates. Give the exact loss function and sign convention for (d_{t+1}). Describe each element of (h_t), how it is constructed, when it becomes observable, and whether it was chosen before looking at the results.
Report the forecast-origin schedule, estimation-window rule, data vintage, common-sample rule, number of out-of-sample origins, number of test functions, covariance estimator, reference distribution, test statistic, degrees of freedom, and p-value. State whether horizons overlap and how dependence is handled. Include average losses and the corresponding loss-difference summaries, not only the test result.
List the other instruments, thresholds, windows, losses, and forecast pairs that were tried. If the same observations selected the model or conditioning variables and tested them, say so. A result can be informative without being confirmatory; transparent reporting lets readers judge its scope and design a separate validation.
Common questions
Q1Is the Giacomini–White test the same as a Diebold–Mariano test?
No. Diebold–Mariano focuses on equal average loss. The one-step GW framework tests selected conditional moments and includes an unconditional moment as one possible restriction.
Q2Does rejection identify which forecast to use in every market regime?
No. It shows that at least one chosen moment is inconsistent with zero under the test assumptions. The result depends on the selected instruments, sample, and loss, and does not guarantee performance in every state.
Q3Can I keep adding volatility thresholds until the test rejects?
That creates a search over conditioning variables. Pre-specify the variables where possible, report all tried specifications, and use a separate confirmation sample or an appropriate family-level correction when making a broad predictive claim.
Q4Does conditional predictive ability mean a trading strategy is profitable?
No. The test compares forecast loss. A trading claim requires a defined decision rule and a separate evaluation of execution, fees, financing, market impact, and risk. Primary research - Giacomini and White, “Tests of Conditional Predictive Ability” (2006) - Giacomini and Rossi, “Forecast Comparisons in Unstable Environments” (2010) - McCracken, “Tests of Conditional Predictive Ability: Existence, Size, and Power” (Federal Reserve Bank of St. Louis Working Paper, 2020) - Diebold and Mariano, “Comparing Predictive Accuracy” (1995) - Newey and West, “A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix” (1987)
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
With dₜ₊₁ = L(A) − L(B), what does a positive loss difference mean?
Choose an answer to see the explanation