Engle–Ng Sign and Size Bias Tests for Volatility Models
Learn how the Engle–Ng sign, negative-size, positive-size, and joint tests check whether a fitted volatility model misses patterns in past shocks
In this guideWhy check shock signs and sizes after fitting a volatility model?
Short summary
Engle–Ng sign and size bias tests check whether a fitted volatility model leaves predictable patterns in squared standardized residuals. The tests isolate shock sign and size; they diagnose selected forms of misspecification, not a trading edge or a universal leverage effect
Why check shock signs and sizes after fitting a volatility model?
A volatility model can capture clustering in the sizes of returns and still miss how different shocks change future variance. A symmetric GARCH model, for example, gives equal-sized positive and negative innovations the same variance impact. If negative and positive shocks leave different patterns in the model's standardized residuals, that restriction may be too strong for the sample.
Engle and Ng's sign and size bias tests turn this into a post-fit diagnostic. They ask whether the sign of the previous innovation, or its size within the negative or positive branch, helps explain today's squared normalized residual after the volatility model has been estimated. The original paper defines the news impact curve and applies the diagnostics in a comparison of volatility models using daily Japanese stock returns. Its empirical ranking is specific to that sample, not a claim that one model wins for every market (Engle and Ng, 1993).
The word “bias” here means a remaining sign- or size-related pattern that the fitted variance model does not explain. It does not mean that a parameter estimate is statistically biased. Start with a clear GARCH model, then use these tests to ask whether its response to past shocks is adequate for the stated sample and purpose.
Define the innovation and standardized residual
Write a return as
\[ r_t = \mu_t + \epsilon_t, \]
where \(\mu_t\) is the conditional mean and \(\epsilon_t\) is the innovation after the mean model is fitted. Let \(h_t = \operatorname{Var}(\epsilon_t\mid\mathcal F_{t-1})\) be the fitted conditional variance. The standardized residual is
\[ \hat z_t = \frac{\hat\epsilon_t}{\sqrt{\hat h_t}}. \]
Under a well-specified variance model, \(\hat z_t^2\) should not be predictably larger or smaller after particular past shocks, beyond what the maintained model already accounts for. The tests use \(\hat z_t^2\) as the diagnostic outcome because standardization removes the variance scale predicted by the fitted model. Squared raw residuals can retain ordinary volatility clustering even when the model has accounted for it.
Fit a defensible conditional-mean equation first. An omitted mean pattern can contaminate the innovations and appear as a variance pattern. The residual, conditional variance, forecast origin, return units, and estimation sample should all be defined before interpreting any p-value.
Build the sign and size regressors
Define an indicator for a negative prior innovation and its complement:
\[ S^-_{t-1}=\mathbf 1(\hat\epsilon_{t-1}<0), \qquad S^+_{t-1}=1-S^-_{t-1}. \]
The sign-bias regressor is \(S^-_{t-1}\). The negative-size regressor is \(S^-_{t-1}\hat\epsilon_{t-1}\), and the positive-size regressor is \(S^+_{t-1}\hat\epsilon_{t-1}\). They ask three different questions: do negative shocks behave differently by sign; does the magnitude of negative shocks matter; and does the magnitude of positive shocks matter? State how zero returns are classified and which innovation scale the implementation uses. Size coefficients depend on that scale.
A news impact curve summarizes how a volatility specification maps a new innovation into the next conditional variance while holding the other state inputs fixed. Symmetric GARCH implies a curve centered symmetrically around zero; asymmetric specifications relax that shape. The curve describes the model's restriction, not a causal response measured directly from the market. Engle and Ng propose the diagnostics to look for selected patterns left outside that restriction.
Sign bias: do negative shocks leave a different pattern?
The sign-bias test asks whether the sign of the previous innovation helps predict the current squared standardized residual while retaining the other terms in the maintained null variance model. Its auxiliary regression is
\[ \hat z_t^2 = a + b_s S^-_{t-1} + \boldsymbol{\beta}^{\top}\mathbf z_{ot}^* + v_t. \]
The individual null is \(b_s=0\). Each auxiliary regression retains \(\mathbf z_{ot}^*\), the vector of normalized derivatives of the maintained conditional variance with respect to nuisance parameters, evaluated under the null. These are score/derivative terms, not necessarily the raw variance regressors; the sign indicator is the added term being tested. A significant coefficient is evidence that negative and nonnegative shocks are followed by different residual variance patterns not captured by the fitted model.
The test does not say that negative returns always raise volatility. It tests a difference in the residual pattern conditional on the maintained model, and its coefficient reflects the indicator coding and sample. It also does not distinguish every possible nonlinear response. A rejection is a reason to inspect the model and data, not to label all negative news as a cause.
Negative size bias: does the magnitude of negative shocks matter?
The negative-size test focuses on the negative branch and asks whether the magnitude of a negative innovation adds explanatory power after retaining the other terms in the maintained null variance model:
\[ \hat z_t^2 = a + b_- S^-_{t-1}\hat\epsilon_{t-1} + \boldsymbol{\beta}^{\top}\mathbf z_{ot}^* + v_t. \]
The individual null is \(b_-=0\). The regression retains \(\mathbf z_{ot}^*\), the other terms from the maintained null variance specification. Because the innovation in this regressor is negative when \(S^-_{t-1}=1\), interpret the coefficient with the exact coding and units in view. The test's main diagnostic question is whether negative-shock size adds explanatory information the variance model did not include.
For example, a symmetric model may respond to all shocks through their square, yet the sample may show a different residual pattern after large negative shocks. The test can flag that gap; it does not estimate a causal effect of bad news or establish that a negative-shock model will forecast better. Engle, Ng, Glosten, Jagannathan, and Runkle compare asymmetric variance specifications in the original study and related work, but those comparisons remain tied to their model assumptions and data (Glosten, Jagannathan, and Runkle, 1993).
<!-- learn:illustration -->

Positive size bias: does the magnitude of positive shocks matter?
The positive-size test checks the other branch while retaining the other terms in the maintained null variance model:
\[ \hat z_t^2 = a + b_+ S^+_{t-1}\hat\epsilon_{t-1} + \boldsymbol{\beta}^{\top}\mathbf z_{ot}^* + v_t. \]
Its individual null is \(b_+=0\). With \(\mathbf z_{ot}^*\) retained from the maintained null specification, this term tests whether positive-innovation size adds explanatory power. A significant result can occur even if the sign-bias test is insignificant: the response may depend on shock size within one branch without a simple average difference between negative and nonnegative shocks.
Do not read “positive size bias” as a claim that positive returns make volatility rise or fall. The coefficient is attached to a particular regressor and its sign convention; significance identifies a remaining association under the auxiliary regression. Inspect fitted values, residual construction, and the volatility model before translating it into prose about good or bad news.
Nelson's EGARCH specification and threshold models provide examples of variance dynamics that can accommodate asymmetric responses without using the same parameterization (Nelson, 1991; Zakoian, 199490039-6)). They are possible model alternatives to evaluate, not automatic remedies or winners.
Joint test: are any of the three terms missing?
The three terms can be added jointly to an auxiliary regression that retains the other terms in the maintained null variance model:
\[ \hat z_t^2 = a + b_1 S^-_{t-1} + b_2 S^-_{t-1}\hat\epsilon_{t-1} + b_3 S^+_{t-1}\hat\epsilon_{t-1} + \boldsymbol{\beta}^{\top}\mathbf z_{ot}^* + v_t. \]
The joint null is \(b_1=b_2=b_3=0\), conditional on the maintained null-model terms \(\mathbf z_{ot}^*\). Engle and Ng note that omitting \(\mathbf z_{ot}^*\) from the joint regression gives a conservative test: its size is no greater than the nominal size, and its power may be reduced. Under the maintained volatility specification and standard large-sample conditions, the Engle–Ng joint LM test is asymptotically compared with a chi-square distribution with three degrees of freedom. A common form uses \(LM=M R^2\), where \(M\) is the number of usable auxiliary-regression rows and \(R^2\) is calculated for the stated test regression. Follow the precise regression and finite-sample convention used by the software or paper you report.
For a wholly hypothetical illustration, suppose the specified joint regression has \(M=400\) usable rows and \(R^2=0.012\). Then \(LM=400\times0.012=4.8\). The 5% reference cutoff for chi-square(3) is about 7.81, so this calculation would not reject at that level. These inputs are invented; the result is not an empirical finding and does not prove that the variance model is symmetric.
Interpret the results with power and collinearity in mind
A small p-value says the tested residual pattern would be unusual under the stated null and its reference distribution. It supports checking for a missing sign or size response. It does not identify the correct alternative, prove a structural leverage mechanism, or show that a volatility forecast has economic value. A large p-value means the test did not detect the selected pattern; it is not proof of symmetry.
The auxiliary regressors can be correlated, particularly the sign indicator and the negative-size term. Engle and Ng note that this collinearity can weaken the power of the individual tests. The joint test may better reflect whether the three terms matter together, but it too depends on sample size, model estimation, and the maintained specification. Multiple tests across assets, model variants, lag choices, or sample windows increase the chance of finding an isolated small p-value.
Outliers, jumps, breaks, an omitted conditional mean, or an inconsistent innovation scale can all affect the result. Prespecify the return definition, estimation window, model, sign convention, and inference method. If the sample is short or the fitted model is complex, report the limits of asymptotic inference and consider an appropriate resampling procedure rather than treating textbook critical values as exact.
Use the diagnostic alongside other model checks
The ARCH LM test asks whether selected lags of squared residuals explain current squared residuals; the ARCH-LM guide covers that broader conditional-magnitude diagnostic. Engle–Ng adds specific sign and branch-size terms. A Ljung–Box test asks about residual autocorrelation, a different question. None of these tests alone certifies a complete volatility model.
After a rejection, an asymmetric model such as GJR-GARCH, EGARCH, or a threshold specification may be worth estimating. Compare it with the maintained model using a common chronological sample, recheck standardized residuals, and assess its forecasts separately. QLIKE and squared-error losses score forecasts after they are made; they do not replace a residual specification test. If a forecast informs a trading rule, evaluate the rule out of sample after fees, spread, market impact, financing, and risk limits. A sign or size bias is not a buy or sell signal.
Common questions
Q1Do sign and size bias tests measure bias in the estimated GARCH coefficients?
No. “Bias” refers to a remaining sign- or size-related pattern in squared standardized residuals. It is not a direct test that the fitted coefficient estimates are biased.
Q2Does a nonsignificant joint test prove the volatility model is symmetric?
No. It means the test did not detect the selected terms in that sample under its assumptions. Limited power, collinearity, a short sample, or a different omitted pattern can still matter.
Q3Is the Engle–Ng test the same as the ARCH LM test?
No. ARCH LM tests whether lagged squared residuals jointly explain current squared residuals. Engle–Ng tests add specific sign and positive- or negative-branch size terms.
Q4Does a rejection mean negative returns are a profitable trading signal?
No. It is a post-fit specification diagnostic. It does not predict return direction or establish a strategy's profitability after execution costs.
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
What does the sign-bias test add to a fitted volatility model check?
Choose an answer to see the explanation
Options glossary
A time-series model that updates conditional variance from past squared shocks and prior variance; it models volatility persistence, not return direction.
Read the deeper guideIV rankThe current IV's position between a lookback-period low and high; one outlier can distort it, so it should not be read as a standalone signal.
Read the deeper guide