Skip to content
All option guides
Forecast evaluation14 min read

Mincer–Zarnowitz Regression: Test Forecast Bias and Calibration

Learn how the Mincer–Zarnowitz regression tests a forecast’s intercept and scale, why zero average error is not the same as calibration, and what the test cannot prove

In this guideZero average error can still hide a calibration problem

Short summary

A Mincer–Zarnowitz (MZ) regression asks whether realized values line up with the level and scale of their forecasts in a linear calibration check. Its familiar null sets the intercept to zero and the forecast slope to one. That joint restriction is more informative than checking average error alone, but it is not a complete test of whether a forecaster used every available piece of information well.

Zero average error can still hide a calibration problem

Suppose a strategy model predicts a one-period return, or an economist predicts next quarter’s growth. Over the evaluation sample, the forecast errors may average to zero. That is useful evidence against a persistent average overprediction or underprediction, but it does not show that forecast values move one-for-one with later outcomes. Forecasts could be too extreme, too compressed, or linearly shifted while positive and negative errors cancel.

The MZ regression makes that second question visible by regressing the realized outcome on the forecast. Mincer and Zarnowitz introduced this framework in their evaluation of economic forecasts (1969 chapter). It is a diagnostic for point forecasts of a numeric target, especially forecasts intended to estimate a conditional mean under squared-error loss. It is not a comparison of which of two models has lower average loss; the [Diebold–Mariano guide](/en/learn/diebold-mariano-forecast-accuracy-test-explained) covers that separate question.

The distinction matters in trading research. A return forecast can have little average bias and still be poorly scaled; a volatility forecast may be systematically too large in one range and too small in another. A regression check can flag a linear mismatch, but its result depends on the evaluation sample, target definition, information timing, and the inference method. It does not establish that a forecast can be turned into a profitable trading rule.

Align forecast origins, targets, and data vintages

For each forecast origin \(t\), pair the forecast \(f_{t+h\mid t}\) with the later realized target \(y_{t+h}\) it was issued to predict at a fixed horizon \(h\). To keep the equations compact, the matched observations are written below as \((f_t,y_t)\); that subscript labels each pair, while the forecast was made at an earlier information cutoff. Both values must refer to the same variable, horizon, units, and timestamp convention. If the forecast is for a five-day return but the outcome is a one-day return, the regression no longer asks whether the forecast was calibrated for its intended target.

Use forecasts that were genuinely available at each origin. Keep the forecast value, model version, data vintage, and information cutoff. For macroeconomic targets, early releases can later be revised; decide whether calibration is judged against the first release or a later finalized value, and preserve that choice. Replacing historical inputs with revised data can give an evaluation access to information unavailable to the original forecaster.

For a trading series, align close-to-close, intraday, and overnight components consistently. Overlapping multi-period targets also create dependence between adjacent forecast errors. A common evaluation sample is important: if one model is missing precisely during volatile dates, comparing different date sets can look like calibration when the difference is really sample composition. The same target and origin schedule should be used across the forecast series being assessed.

Separate generated forecasts from fitted values. If the model was estimated on the same observations used in the MZ regression, apparent calibration can partly reflect in-sample fitting. For a claim about future forecasting, construct a chronological out-of-sample sequence, re-estimate only with information available at each origin, and keep any final confirmation period untouched by model or threshold selection.

Fit the Mincer–Zarnowitz regression and state the null

The standard levels regression is

yₜ = α + β fₜ + uₜ

where (yₜ) is the realized target, (fₜ) is the forecast, and (uₜ) is the regression residual. The familiar joint calibration restriction is

H₀: α = 0 and β = 1

If both restrictions hold, the fitted linear relationship between outcome and forecast is the identity line. A constant offset is not needed, and a one-unit change in the forecast corresponds to a one-unit change in the fitted outcome. Estimate the regression, then test the two restrictions jointly with a Wald or equivalent F test using a covariance estimate appropriate for the sampling design. Testing the intercept and slope separately is not the same as testing their joint restriction.

The regression residual (uₜ) is not automatically the raw forecast error (yₜ − fₜ). They are equal under the joint null, but away from it the estimated intercept and slope absorb a linear shift and rescaling. Report the actual forecast error as well as the regression coefficients, so readers can see both the average miss and the calibration relation.

The MZ restriction is often called a forecast unbiasedness or rationality test. Be precise about the definition. Zero unconditional mean error is the condition (E[yₜ − fₜ] = 0); the restrictions (α = 0, β = 1) impose a particular linear relationship and are generally stronger. Holden and Peel show that the conventional joint restriction is sufficient but not necessary for forecast unbiasedness under their formulation (1990). Do not treat the labels “unbiased” and “MZ calibrated” as interchangeable without stating the restriction being tested.

<!-- learn:illustration -->

Slate-blue and amber glass spheres align along a clear diagonal plane, with small offsets suggesting forecast calibration
Conceptual view of forecasts and realized values compared against a calibration line; not market data, a test result, or a trading signal.

A small example separates mean bias from the joint restriction

Consider three hypothetical forecasts and later outcomes:

Forecast (fₜ)Realized value (yₜ)Error (yₜ − fₜ)
11.50.5
22.00.0
32.5−0.5

The average forecast is 2, the average outcome is also 2, and the average forecast error is zero. Yet the values follow (yₜ = 1 + 0.5 fₜ), so the MZ regression has intercept (α = 1) and slope (β = 0.5), not (0, 1). The errors cancel on average even though the linear calibration restriction fails: the realized values move only half as much as the forecasts in this constructed example.

This is a hand-built arithmetic illustration, not market data or an inferential sample. With three noiseless points, it makes no sense to report a meaningful p-value or claim that a real forecast system has failed. In an actual sample, the joint test uses estimation uncertainty; a large coefficient deviation may be imprecise, while a small one may be precisely estimated with enough observations. The example only shows why the average error and the MZ restrictions answer different questions.

A recalibration fit such as (1 + 0.5 fₜ) would reproduce these three outcomes exactly because it was constructed from them. That does not show it will work later. If recalibration is part of the forecast procedure, estimate it on an earlier training segment and evaluate the adjusted forecast on later data that did not determine those coefficients.

Read the intercept, slope, and R-squared together

The intercept (α) is the fitted outcome at a zero forecast, so its practical meaning depends on whether zero is in or near the forecast range. The slope (β) describes how strongly fitted outcomes vary with forecast values. With the intercept accounted for, a slope below one means the forecast changes more per unit than the fitted outcome; a slope above one means the fitted outcome changes more. Neither coefficient alone is a full description when the other is also unrestricted.

A high (R²) does not establish calibration. For example, a noiseless relationship (yₜ = 2 + 1.1 fₜ) has (R² = 1), but its intercept and slope do not meet the MZ null. (R²) summarizes how much variation is explained by the linear fit; the joint test asks whether that fit is the identity line. Conversely, a noisy forecast may have a low (R²) even when the joint restriction is not rejected. Calibration and precision are related but distinct properties.

Do not infer the calibration result from the two individual t-statistics alone. The intercept and slope estimates can be correlated, so the joint Wald/F statistic uses their covariance. Report (α), (β), their uncertainty, the joint statistic and p-value, the sample size, and the average raw error. Also show a scatterplot or binned comparison when helpful, with the identity line and an honest note about any bins chosen after inspecting the data. A linear test can miss curved miscalibration, changing regimes, or tail-specific errors.

Calibration is weaker than full forecast efficiency

Under squared-error loss, an optimal point forecast is the conditional expectation of the target given the forecaster’s information set. That requirement implies that forecast errors should not be predictable using information that was already available when the forecast was made. An MZ regression only checks a linear relationship involving the forecast itself and a constant. It does not test every other variable in that information set.

A stronger diagnostic adds pre-specified information variables (zₜ) to a forecast-error regression, for example

yₜ − fₜ = δ₀ + δ₁ zₜ + vₜ

If a variable known at the forecast origin predicts later errors, the forecast may have left usable information on the table. The variables must be timed correctly and chosen with care; testing many lags, market states, and transformations creates another multiple-search problem. Failure to find predictability for a short list does not prove efficiency relative to the full information set.

Zarnowitz’s evaluation of survey forecasts discusses bias, serial correlation, sample size, and the power of forecast tests (NBER Working Paper 1070, 1983). For this reason, describe MZ as a linear calibration diagnostic or a weak forecast-efficiency check, not as proof that the forecaster used all relevant information optimally. The [Giacomini–White guide](/en/learn/giacomini-white-conditional-predictive-ability-test-explained) covers a different conditional comparison based on selected loss moments.

Respect estimation, overlap, and finite-sample uncertainty

Forecast errors may be heteroskedastic, serially correlated, or overlapping, especially for rolling multi-step forecasts. The usual OLS coefficient formulas do not by themselves guarantee valid standard errors for the joint restriction. Use an inference procedure justified for the forecast horizon, dependence pattern, and estimation scheme; report it rather than assuming errors are independent. There is no single automatic bandwidth or correction that fits every design.

Estimated forecasts add another complication. If the same short sample estimates model parameters, constructs a calibration adjustment, and tests that adjustment, ordinary regression uncertainty may not reflect the full procedure. Nested forecasts can also make the null distribution nonstandard for comparisons of predictive accuracy. Use a result derived for the actual forecast construction, or evaluate a fully specified calibration step on a later holdout. Do not select the regression form, target, and sample after trying many variants and then treat the last p-value as confirmatory.

A non-rejection is not proof that the forecast is calibrated; it may reflect low power or a wide confidence region. A rejection is also conditional on the target, sample, linear form, and covariance estimate. Check sensitivity to reasonable evaluation choices, but disclose those checks and distinguish them from a pre-specified test. This matters when forecasts are highly persistent or the evaluation sample contains only a few independent episodes.

Choose a test that matches the forecast question

MZ tests a linear calibration restriction for a forecast and realized target. The [Diebold–Mariano test](/en/learn/diebold-mariano-forecast-accuracy-test-explained) instead asks whether two forecasting procedures have equal average loss under a chosen score. One forecast can be well calibrated yet less accurate by that score, or more accurate on average while failing an MZ calibration check.

If the question is whether accuracy changes with information known at the forecast origin, the [Giacomini–White test](/en/learn/giacomini-white-conditional-predictive-ability-test-explained) evaluates selected conditional loss moments. If many forecasts or trading rules were searched and the question concerns a family of candidates, a family-level procedure such as the [Model Confidence Set](/en/learn/model-confidence-set-forecast-comparison-explained) or [White Reality Check/Hansen SPA](/en/learn/white-reality-check-vs-hansen-spa-test-trading-rules-explained) addresses a different selection problem. None of these methods substitutes for a clearly defined target and honest forecast origins.

The standard MZ levels regression is for point forecasts of a numeric target. Quantile, probability, interval, and density forecasts need calibration tests and scoring rules matched to those forecast objects. A good result on a mean forecast should not be carried over to a tail-risk forecast or a volatility distribution without a separate argument.

Report the calibration design so it can be reproduced

State the forecast target, horizon, units, information cutoff, evaluation dates, data vintage, and whether the forecasts were genuinely out of sample. Give the exact regression equation, the definition of forecast error, and the restrictions tested. Clarify whether “unbiasedness” means zero average error or the joint MZ restriction (α = 0, β = 1).

Report coefficient estimates, standard errors or confidence intervals, the joint Wald/F statistic, degrees of freedom, p-value, number of forecast origins, mean forecast error, and (R²) with its limited interpretation. Explain the covariance or resampling method, especially if horizons overlap or errors are serially dependent. Disclose model estimation, recalibration, selection, and any alternative targets, samples, or transformations that were tried.

A transparent report lets another reader distinguish average bias, linear calibration, information efficiency, and comparative accuracy. Those are separate empirical claims. Passing one regression restriction does not validate every use of the forecast, and a forecast evaluation by itself is not evidence of trading profitability after costs.

Common questions

Q1Is the Mincer–Zarnowitz test just a test of average forecast error?

No. Average error checks whether forecast and outcome means differ. The conventional MZ test jointly restricts the intercept to zero and the slope to one. Holden and Peel show that this joint restriction is sufficient but not necessary for unbiasedness under their formulation.

Q2Does passing the MZ test prove a forecast is efficient?

No. It tests a linear relationship involving the forecast, not whether all information available at the forecast origin is unable to predict the errors. Efficiency requires a stronger information-set condition.

Q3Can I use the same regression for a VaR or quantile forecast?

Not without changing the calibration definition and inference. The standard levels regression is for numeric point forecasts; quantile and tail forecasts need tests and scores designed for those objects.

Q4Does calibration prove that a trading strategy is profitable?

No. Calibration concerns the relation between forecasts and outcomes. A trading claim needs a pre-defined decision rule and a separate out-of-sample evaluation that includes execution, fees, financing, market impact, and risk. Primary research - Mincer and Zarnowitz, “The Evaluation of Economic Forecasts” (1969) - Holden and Peel, “On Testing for Unbiasedness and Efficiency of Forecasts” (1990) - Zarnowitz, “Rational Expectations and Macroeconomic Forecasts” (NBER Working Paper 1070, 1983) - Diebold and Mariano, “Comparing Predictive Accuracy” (1995) - Giacomini and White, “Tests of Conditional Predictive Ability” (2006)

Sources and further reading

Report an issue

We’ll prepare an email with this article link. Mark receives the report only after you send it

Quick check

Read the guide? Check yourself with 3 questions

Question 1 / 3

Question 01

What is the conventional joint null in the Mincer–Zarnowitz levels regression?

Choose an answer to see the explanation

Options glossary

Clear definitions of essential option terms, from calls, puts, and option chains to IV, Greeks, open interest, and max pain

Browse the options glossary