Skip to content
All option guides
First-stage evidence and inference when identification is weak14 min read

Weak Instrument Diagnostics: First-Stage F and Robust Inference

Learn how first-stage F statistics, partial R-squared, Stock–Yogo criteria, robust diagnostics, and Anderson–Rubin inference answer different weak-instrument questions.

In this guideWhy instrument strength matters

Short summary

A first-stage statistic describes how much excluded instruments explain an endogenous regressor after included controls. It does not establish instrument validity, and “F above 10” is not a universal guarantee. The relevant diagnostic depends on the number of endogenous regressors and the error assumptions; weak-IV-robust inference can remain wide or unbounded.

Why instrument strength matters

Instrumental variables use variation in an endogenous regressor \(X\) that is predicted by excluded instruments \(Z\), conditional on included controls \(W\).

If \(Z\) predicts little of the remaining variation in \(X\), the data contain little information about the structural effect from this source.

With weak relevance, finite-sample 2SLS can be biased toward OLS and its sampling distribution can depart substantially from a normal approximation. A conventional Wald interval may then have poor coverage even when its standard error looks precise.

Staiger and Stock develop weak-instrument asymptotics that explain these failures and motivate nonstandard confidence regions (1997 paper).

Strength is only one part of identification. A strong first stage cannot show that \(Z\) is independent of the structural error or that it affects the outcome only through \(X\).

The instrumental-variables guide discusses those separate assumptions and the effect being identified.

The first stage isolates conditional instrument variation

For one endogenous regressor and \(q\) excluded instruments, write the first stage as

\[ X_t = W_t'\delta + Z_t'\pi + v_t. \]

Here \(W_t\) contains included exogenous controls, often including an intercept, while \(Z_t\) contains excluded instruments. The standard first-stage F test evaluates \(H_0:\pi=0\), a joint restriction that the excluded instruments add no explanatory power after \(W_t\).

This is a conditional relevance question, not a test of exclusion or independence. Report which controls are included, how many excluded instruments are tested, the sample and covariance assumptions, and the statistic used.

A coefficient’s individual t statistic is not a substitute for the joint test when there are several excluded instruments.

Partial R-squared and the conventional F statistic

The partial \(R^2\) measures the share of residual variation in \(X\), after removing \(W\), explained by the residualized excluded instruments.

Let \(R^2_p\) denote that partial \(R^2\), \(n\) the sample size, \(k\) the number of included regressors, including an intercept if used, and \(q\) the number of excluded instruments.

Under the conventional homoskedastic linear regression calculation, the incremental F statistic is

\[ F = \frac{R^2_p/q}{(1-R^2_p)/(n-k-q)}. \]

The numerator measures the fit improvement per excluded-instrument restriction; the denominator estimates residual variance using the full first-stage residual degrees of freedom. This formula is not a universal robust-statistic identity.

A robust covariance estimator changes the test calculation and reference distribution.

Partial \(R^2\) and F answer related but different descriptive questions: partial \(R^2\) is a scale-free fit measure, while F also reflects sample size and the restriction count.

A small partial \(R^2\) may produce a large F in a very large sample without making every finite-sample IV approximation reliable.

Why F above 10 is not a universal threshold

The familiar value 10 is a rule of thumb, not a pass/fail theorem. Staiger and Stock discuss it as a practical guide in a particular weak-instrument framework; it does not guarantee valid inference for every sample, estimator, number of instruments, or error process.

Stock and Yogo define instrument weakness in terms of bounds on relative 2SLS bias or Wald-test size distortion and tabulate critical values for those chosen criteria. The critical value therefore depends on the criterion and model dimensions.

Their conventional first-stage and Cragg–Donald results rely on homoskedastic, independent-error assumptions.

A number from those tables cannot be carried unmodified to every robust setting (author-hosted chapter page).

A value above 10 does not prove exclusion, independence, or a causal interpretation. A value below 10 does not prove that an instrument is useless or that a particular estimate is invalid.

Treat the statistic as a diagnostic under stated assumptions, then choose inference that matches the design.

The Hansen J guide explains why an overidentification test does not replace a strength assessment.

<!-- learn:illustration -->

A fine gold filament links a small amber crystal to a large blue glass sphere within a broad haze
Conceptual illustration of a weak first-stage link and surrounding uncertainty; not data or a diagnostic result.

Multiple endogenous regressors need a joint strength assessment

If there is more than one endogenous regressor, separate first-stage F statistics can miss a weakly identified linear combination.

Each regressor may look individually predicted even when the instrument-induced variation does not span the combinations needed to estimate all structural coefficients precisely.

For the homoskedastic linear IV setup, the Cragg–Donald minimum-eigenvalue statistic summarizes this joint identification information. Stock–Yogo critical values for it target stated bias or size-distortion criteria.

The assumptions matter: the familiar critical values are not automatically valid after arbitrary heteroskedasticity, serial dependence, or clustering.

The number of excluded instruments, the number of endogenous regressors, and the rank of the first-stage coefficient matrix all matter. State the full setup and use a test designed for it; do not infer joint strength by averaging the individual first-stage F values.

Robust errors call for design-matched diagnostics

Financial observations may be heteroskedastic, serially dependent, or clustered by firm, venue, or policy unit. The classical first-stage F and Cragg–Donald calibrations do not become robust merely because the second-stage standard errors are clustered.

For one endogenous regressor, Montiel Olea and Pflueger develop an effective-F test for weakness that allows heteroskedasticity, autocorrelation, and clustering under their stated conditions.

It scales a conventional first-stage F using covariance information, then compares the result with criterion-specific critical values (2013 paper).

The effective F is not the same quantity as every software-reported robust Wald F.

The Kleibergen–Paap rk Wald F is widely reported with robust covariance estimators, but Stock–Yogo critical values were derived for different conventional statistics and assumptions. Do not compare these values as if they were on a common calibrated scale.

Report the estimator, statistic, covariance estimator, clustering level or dependence treatment, and the critical-value framework.

Newey–West or cluster-robust covariance estimation addresses sampling uncertainty under its assumptions; it does not by itself cure weak identification.

Anderson–Rubin inference can remain valid with weak relevance

For a candidate coefficient value \(\beta_0\) in a single-endogenous-regressor model, form the outcome adjusted by that candidate, \(Y-X\beta_0\), and test whether the excluded instruments jointly explain it after included controls.

Under the null and instrument exogeneity, this is an Anderson–Rubin test of \(\beta=\beta_0\).

Because the test does not divide by an estimated first-stage coefficient, its validity need not rely on that coefficient being large.

The original Anderson–Rubin construction uses the model’s distributional assumptions; in applied work, the covariance estimator and reference distribution must also match the sampling design. A robust label on a test is not a license to ignore few clusters or serial dependence.

Inverting the test over candidate values gives a confidence set. With weak information, that set may be wide, disjoint, or unbounded; this is a finding about limited identifying information, not a software failure. Anderson–Rubin tests can also have low power.

The GMM overidentification guide covers a different question and should not be treated as weak-IV-robust coefficient inference.

The original paper is Anderson and Rubin (1949).

A worked partial-F calculation

Consider a hypothetical first stage with \(n=200\) observations, \(k=3\) included regressors including the intercept, \(q=2\) excluded instruments, and partial \(R^2_p=0.04\).

Under the conventional homoskedastic formula, the residual degrees of freedom are \(200-3-2=195\).

\[ F = \frac{0.04/2}{(1-0.04)/195} = \frac{0.02}{0.96/195} = 4.0625. \]

The arithmetic says that the conventional joint statistic is 4.0625 under these inputs and assumptions. It does not, by itself, prove invalidity, determine the Stock–Yogo criterion-specific conclusion, or establish the result under heteroskedasticity or clustering.

A robust analysis needs its own design-matched statistic and calibration.

If the same two instruments are used in a fuzzy regression-discontinuity design, the first-stage discontinuity is part of the design-specific relevance analysis.

The F calculation above is not a substitute for the RDD assumptions, bandwidth choices, or inference appropriate to that design.

Report the diagnostic and the identifying argument separately

State the endogenous regressors, excluded instruments, included controls, sample, first-stage coefficients, partial \(R^2\), and the exact strength statistic.

For several endogenous regressors, identify the joint statistic and its assumptions; for robust settings, name the covariance and critical-value method instead of writing only “F = …”.

If the statistic suggests weak identification, report a weak-IV-robust test or confidence set for the target coefficient. Explain whether the set is bounded and how its shape changes the substantive conclusion.

Do not replace an uninformative interval with a conventional Wald interval simply because it is easier to summarize.

Finally, defend the instrument’s independence and exclusion paths using institutional and economic evidence. Relevance diagnostics, balance checks, placebo outcomes, and an overidentification test can reveal problems but cannot prove every assumption.

A difference-in-differences design uses different identifying assumptions; a stronger first stage does not make it interchangeable with that design.

Common questions

Q1Does a first-stage F above 10 prove that an instrument is valid?

No. It is at most a relevance diagnostic under a particular calibration. It cannot establish independence or exclusion.

Q2Is partial R-squared the same as the first-stage F statistic?

No. Partial R-squared describes incremental fit. The conventional F also depends on sample size, restriction count, and residual degrees of freedom.

Q3Can I compare a robust F directly with Stock–Yogo critical values?

Only when the statistic and assumptions match the calibration. Robust statistics often require different critical values and interpretation.

Q4What should I report when the instrument appears weak?

Report the design-matched diagnostic and a weak-IV-robust test or confidence set, including whether that set is wide, disconnected, or unbounded.

Sources and further reading

Report an issue

We’ll prepare an email with this article link. Mark receives the report only after you send it

Quick check

Read the guide? Check yourself with 3 questions

Question 1 / 3

Question 01

What does the conventional first-stage F test assess?

Choose an answer to see the explanation

Options glossary