ADF vs. Phillips–Perron: How Unit-Root Tests Differ
Compare the ADF and Phillips–Perron unit-root tests: their shared null, deterministic terms, lag and bandwidth corrections, finite-sample limits, and careful interpretation
In this guideBoth tests place a unit root in the null
Short summary
The augmented Dickey–Fuller (ADF) and Phillips–Perron (PP) tests usually start from the same question: does a series have a unit root? Their main difference is how they handle serial dependence in the test regression. ADF adds lagged differences; PP adjusts a simpler regression statistic using an estimate of long-run covariance. Both need a deliberate deterministic specification and nonstandard critical values, and neither test automatically tells you to difference a series.
Both tests place a unit root in the null
For an autoregressive level equation, a unit root means that a shock need not decay from the level. Writing the first difference as a regression on the lagged level makes the hypothesis visible: if the coefficient on the lagged level is zero, the level process has a root at one under the specified model. A negative coefficient is evidence in the direction of a stationary alternative, subject to the maintained assumptions and deterministic terms.
The null is therefore not “the series is stationary.” In common ADF and PP implementations, the null is a unit root; the alternative is stationarity around the chosen deterministic component. A test statistic that is not sufficiently negative to reject the null does not prove that the series has a unit root. It may simply contain too little information to distinguish a unit root from a persistent process whose autoregressive root is close to one.
This is different from a test whose null is stationarity, such as KPSS, and from a direct test of cointegration. The stationarity and unit-root guide introduces those broader questions; the Engle–Granger guide treats the special case of a stationary combination of nonstationary series.
Choose deterministic terms before reading a critical value
A unit-root regression can include no deterministic term, a constant, or a constant and linear time trend. These are not cosmetic switches. A constant permits a drift in the level under the unit-root null and a nonzero mean under a stationary alternative. A time trend allows the stationary alternative to fluctuate around a deterministic trend. The model and its null distribution change with the selected terms.
Including an unnecessary trend can reduce power because the test estimates extra parameters. Omitting a real drift or trend can distort the interpretation and size of the test. Choose the deterministic specification from the data-generating question and the plotted series, state it explicitly, and use the matching critical values. Do not choose whichever option produces the preferred rejection.
The same discipline applies to a log price, a spread, an interest rate, or a return. A positive trend that is plausible for a price level may make no sense for a return. The test cannot select the economic meaning of the target for you. Deterministic components are part of the hypothesis, not a preprocessing detail to hide after estimation.
ADF absorbs serial dependence with lagged differences
The augmented Dickey–Fuller regression can be written as
Δyₜ = dₜ′δ + γyₜ₋₁ + Σᵢ₌₁ᵖ φᵢΔyₜ₋ᵢ + εₜ
Here, dₜ contains the chosen deterministic terms, and the null is H₀: γ = 0. The extra lagged differences are intended to make the remaining disturbance adequately well behaved for the test approximation. They model short-run dynamics parametrically inside the regression; they are not additional unit-root restrictions.
The lag order p is consequential. Too few lags can leave serial correlation in εₜ, undermining the intended approximation. Too many consume degrees of freedom and can weaken the test, especially in a short sample. A Monte Carlo study by Harris (1992)90022-Q) examines operational choices involving ADF lag structure, size, and power. A practical analysis can start with a stated maximum lag, compare a defensible information-criterion rule with residual diagnostics, and disclose whether the conclusion changes under nearby lag choices. A small selected lag is not proof that residual dependence is harmless.
Lag selection itself can be difficult in some processes. Ng and Perron show that when errors have a moving-average root near −1, conventional AIC or BIC choices can select too few augmentation lags in the setting they study; their modified information criteria are designed for unit-root procedures with GLS detrending and related tests. This is a warning about selection under particular dynamics, not a theorem that one criterion or a longer lag always wins.
PP keeps a shorter regression and adjusts its statistic
The Phillips–Perron test begins with a Dickey–Fuller-style regression without the ADF sum of lagged differences:
Δyₜ = dₜ′δ + γyₜ₋₁ + uₜ
It then corrects the test statistic using estimates of short-run variance and long-run covariance. A common long-run variance estimate takes the form
Ω̂ₗ = γ̂₀ + 2Σⱼ₌₁ˡ K(j/(ℓ+1))γ̂ⱼ
where γ̂ⱼ are residual autocovariances, K is a kernel weight, and ℓ is a bandwidth or truncation choice. Exact definitions differ across PP statistics and software conventions. The conceptual contrast is stable: ADF expands the regression with lagged changes; PP applies a nonparametric dependence correction to the statistic from a shorter regression.
PP is not simply “ADF with robust standard errors.” Its correction changes the unit-root statistic and its asymptotic treatment; ordinary heteroskedasticity-robust regression output does not make an ordinary t critical value valid. The bandwidth and kernel govern how much dependence the correction captures. A bandwidth that is too small can miss relevant correlation; an overly broad choice can make a noisy long-run variance estimate less reliable.
<!-- learn:illustration -->

Their reference distributions are not ordinary t distributions
Under the unit-root null, the lagged level is generated by a stochastic trend rather than a stationary regressor. The resulting DF-family statistics have nonstandard limiting distributions. The conventional t table is therefore not the reference distribution, even for a large sample. Critical values depend on whether the regression has no constant, a constant, or a constant plus trend, as well as on the statistic and implementation.
PP statistics are transformed to account for dependence, but their unit-root null is still evaluated against the appropriate nonstandard reference law. Use critical values or p-values from a recognized implementation that matches the deterministic terms and statistic. Do not compare an ADF statistic from one specification with PP critical values from another, or infer rejection by treating a displayed coefficient t-ratio as an ordinary Student t statistic.
Asymptotic calibration does not guarantee good behavior in every finite sample. Schwert’s Monte Carlo investigation reports sensitivity of unit-root procedures to model misspecification and different finite-sample behavior across methods. A reported p-value should be read with the sample length, data-generating features, lag or bandwidth setting, and regression specification in view.
Finite-sample choices can change size and power
“Size” is the probability of rejecting a true unit-root null; “power” is the probability of rejecting under a specified stationary alternative. A test can look decisive in one sample while being mis-sized under its actual dependence structure. The ADF lag order and the PP long-run variance correction target related dependence problems through different approximations, so a result from one method is not a guaranteed robustness check for the other.
For ADF, inspect whether residual autocorrelation remains after adding lags and whether a longer lag materially changes inference. For PP, report the kernel and bandwidth rule and examine reasonable alternatives. These are sensitivity checks, not a license to search many specifications and keep the most favorable p-value. The Ng–Perron results also caution that lag choice interacts with the disturbance process and deterministic detrending; a mechanical information criterion may not settle every design.
Both tests can have low power against highly persistent but stationary alternatives. A failure to reject is not affirmative evidence that differencing is required. When the sample is short, the series is near-integrated, or the conclusion is central to a model, consider complementary evidence and alternative procedures with clearly stated assumptions rather than treating ADF and PP as independent votes.
Breaks and transformations complicate the unit-root story
A level shift, trend change, policy regime, exchange-rate peg change, or market-structure transition can make a process look highly persistent even when its behavior is better described by stationary segments around a changing deterministic path. Perron’s analysis of major historical breaks demonstrates why a unit-root conclusion can depend on whether a break is allowed. Standard ADF and PP specifications do not automatically discover every break.
Plot the series, document known institutional changes, and ask whether a break-aware test or segmented model addresses the actual question. Searching for a break after seeing the same sample can itself affect inference, so use a method with critical values appropriate to the search or treat the exercise as exploratory. The Bai–Perron guide covers regression-break estimation; it is a separate procedure rather than a drop-in ADF or PP correction.
The transformation matters as well. A price level, log price, price spread, and return are different variables with different economic interpretations. Differencing a unit-root-like level may yield a more stable series, but differencing a stationary level can discard information and create an unnecessarily noisy specification. For pairs research, cointegration analysis can help distinguish separate stochastic trends from a potentially stationary spread. None of these diagnostics alone establishes a profitable strategy.
A shared equation shows what changes and what does not
Suppose a hypothetical level follows
yₜ = yₜ₋₁ + μ + εₜ
with a constant drift μ and serially dependent innovations εₜ. Subtracting yₜ₋₁ gives Δyₜ = μ + εₜ, so the coefficient γ on yₜ₋₁ is zero in the corresponding differenced regression. That is the unit-root null representation. This algebra does not provide a p-value: the inferential reference law remains nonstandard and depends on how the deterministic terms and dependence are handled.
If a practitioner fits ADF, the lagged differences attempt to represent the short-run dependence in the regression. If the practitioner fits PP, the base equation is shorter and the statistic uses an estimated long-run covariance correction. If the data instead have a broken deterministic trend, both standard implementations can be misspecified even when their software reports a precise-looking statistic.
The comparison therefore has three separate layers: the substantive null, the regression/deterministic specification, and the correction for residual dependence. Holding the first two fixed makes ADF and PP useful alternative diagnostics; changing all three until one rejects does not. A difference between their outcomes is a prompt to inspect assumptions and tuning, not a contest whose winner automatically determines the model.
Report the setup and use the result as evidence, not an order
Report the variable and transformation, sample dates, observation frequency, deterministic terms, ADF lag rule and selected order, or PP statistic, kernel and bandwidth. State the null and alternative, statistic, matching critical value or p-value, and any residual-autocorrelation or break checks. If results are sensitive to defensible settings, say so; do not present only the specification that supports a preferred narrative.
A useful workflow is to define the target and economic model first, choose deterministic terms, run a well-specified ADF or PP test, diagnose serial dependence and possible breaks, and compare the conclusion with complementary evidence. The variance-ratio guide examines a different implication of random-walk behavior. It does not share the same null or replace a unit-root test.
Neither ADF nor PP is a command to difference once a p-value crosses a threshold. Differencing is a modeling decision about the series and the question: forecasting changes, estimating a long-run relation, testing a trading spread, or preserving levels can lead to different choices. A test can inform that decision, but the final transformation should also respect the data-generating story, stability, forecast objective, and any cointegrating structure.
Common questions
Q1Do ADF and PP test the same null?
Usually, yes: both test a unit-root null against a stationary alternative defined around selected deterministic terms. Their statistics handle serial dependence differently, but the deterministic specification and critical values still need to match.
Q2Is PP just ADF with robust standard errors?
No. PP modifies the unit-root statistic using a nonparametric estimate of long-run covariance. A generic robust standard error does not reproduce the PP correction or make an ordinary t distribution appropriate.
Q3Which test should I trust if their results disagree?
Neither result automatically wins. Check deterministic terms, the ADF lag order, the PP kernel and bandwidth, residual dependence, sample size, and possible breaks. Report sensitivity and use complementary evidence suited to the question.
Q4Does failure to reject mean I should difference the series?
No. It is one piece of evidence about persistence under a particular model. Choose transformations based on the variable, objective, stability, and any long-run relation; over-differencing can remove useful information. Primary research - Dickey and Fuller, “Distribution of the Estimators for Autoregressive Time Series with a Unit Root” (1979) - Harris, “Testing for Unit Roots Using the Augmented Dickey–Fuller Test: Some Issues Relating to the Size, Power and the Lag Structure of the Test” (1992)90022-Q) - Phillips and Perron, “Testing for a Unit Root in Time Series Regression” (1988) - Schwert, “Tests for Unit Roots: A Monte Carlo Investigation” (1989) - Ng and Perron, “Lag Length Selection and the Construction of Unit Root Tests with Good Size and Power” (2001) - Perron, “The Great Crash, the Oil Price Shock, and the Unit Root Hypothesis” (1989)
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
What is the usual null hypothesis shared by ADF and PP unit-root tests?
Choose an answer to see the explanation
Options glossary
Clear definitions of essential option terms, from calls, puts, and option chains to IV, Greeks, open interest, and max pain
Browse the options glossary