Pesaran–Timmermann Test: Does a Directional Forecast Add Information?
Learn how the Pesaran–Timmermann test compares directional hits with a baseline based on actual and forecast up-rates, and why serial dependence changes inference.
In this guideA 64% hit rate can be unimpressive or useful depending on the baseline
Short summary
The Pesaran–Timmermann (PT) test asks whether forecasts contain information about which way an outcome moves. For a binary up/down forecast, it compares the observed hit rate with the agreement rate implied by the separate actual-up and forecast-up shares. That baseline is not necessarily 50%. The usual test also relies on assumptions about dependence across forecast origins, and statistical significance does not establish that a trading rule is profitable.
A 64% hit rate can be unimpressive or useful depending on the baseline
Suppose a market rises on 60% of the dates in an evaluation sample. A forecaster calls “up” on 70% of those dates. Even if the calls and outcomes were unrelated, the two series would agree fairly often: both would say up on some dates, and both would say down on others. Comparing the hit rate mechanically with 50% would ignore that imbalance.
The PT framework makes that comparison explicit. It asks whether the forecast direction and realized direction line up more often than their separate frequencies would imply under independence. This is a test of directional predictive information, not a score for the size of returns, the timing of a trade, or the value of an entire strategy. Pesaran and Timmermann introduced the predictive-performance test using contingency tables; for the binary 2×2 case, their test is asymptotically equivalent to testing independence {source:pesaranTimmermannDirectionalTests1992}.
Put each matched forecast into a two-by-two table
Let \(Y_t=1\) when the realized outcome is classified as up and \(Y_t=0\) when it is classified as down. Let \(Z_t=1\) when the forecast calls up and \(Z_t=0\) when it calls down. Each row in the evaluation should pair a forecast made at origin \(t\) with the outcome for the horizon that forecast was intended to predict.
The four cells count up/up, up/down, down/up, and down/down pairs. If \(n_{11}\) and \(n_{00}\) are the two agreement cells and \(n\) is the number of matched, usable pairs, the observed directional hit rate is
\[ \widehat{P}=\frac{n_{11}+n_{00}}{n}. \]
This setup is binary. Decide in advance how to handle a zero return, a zero forecast, or a neutral band. For example, a study might classify a return at or below zero as “not up” and a forecast at or below its threshold as “not up.” Excluding ties is another possible rule, but it changes the sample and must be applied consistently. Do not count zero-zero as a third kind of match while using a two-category baseline formula.
The forecast origin and target horizon must match as well. If \(Z_t\) predicts the sign of the next five-session return, \(Y_t\) must label that same five-session return from the same origin, not tomorrow’s one-session return. A different target would put unlike events in the table. Daily forecasts of overlapping five-session outcomes are valid pairs for a descriptive table, but they create the dependence issue discussed below.
Derive the independence baseline from both up-rates
Let \(\hat p_y\) be the sample share of realized outcomes classified as up, and let \(\hat p_z\) be the sample share of forecasts classified as up. If actual and forecast labels were independent, the expected share of agreements would be
\[ \widehat{P}^{*}=\hat p_y\hat p_z+(1-\hat p_y)(1-\hat p_z). \]
The first product is the probability of an up/up match under independence; the second is the down/down match. The baseline therefore changes with both marginal shares. It can be above or below 50%, so a forecast that always predicts the more common direction may have an apparently respectable hit rate without adding information.
An extreme case makes this visible. If a classifier calls “up” on every date, then \(\hat p_z=1\), the observed hit rate is just the actual-up share \(\hat p_y\), and the independence baseline is also \(\hat p_y\). The excess agreement is zero by construction. Such a constant forecast can look accurate when up outcomes are common, but it does not distinguish which dates will rise; its estimated variance may also be degenerate, so a conventional PT p-value may not exist.
Consider 100 hypothetical matched dates. Actual outcomes are up on 60 dates and forecasts call up on 70. Suppose the table is:
| Realized direction | Forecast up | Forecast down | Total |
|---|---|---|---|
| Up | 47 | 13 | 60 |
| Down | 23 | 17 | 40 |
| Total | 70 | 30 | 100 |
There are 64 agreements, so \(\widehat P=0.64\). The independence baseline is \(0.60\times0.70+0.40\times0.30=0.54\). The observed hit rate is 10 percentage points above that baseline. These invented counts illustrate the calculation; the gap alone is not a valid uncertainty estimate or proof of useful trading performance.
Scale the hit-rate gap by its estimated uncertainty
For the standard binary PT statistic, define
\[ \widehat v=\frac{\widehat P^{*}(1-\widehat P^{*})}{n} \]
and
\[ \widehat w=\frac{(2\hat p_y-1)^2\hat p_z(1-\hat p_z)+(2\hat p_z-1)^2\hat p_y(1-\hat p_y)}{n}. \]
Then
\[ PT=\frac{\widehat P-\widehat P^{*}}{\sqrt{\widehat v-\widehat w}}. \]
Under the null and the standard large-sample assumptions, this statistic has an asymptotic standard-normal reference distribution. The numerator is the excess agreement over the independence baseline. The denominator accounts for estimation of that baseline from the same sample rather than treating it as a fixed 50% target. The formula and its notation are documented by the statsmodels implementation {source:statsmodelsPesaranTimmermannAPI}.
The expression above is the leading asymptotic variance form documented by statsmodels. The fuller finite-sample delta variance in the original treatment adds \(4\hat p_y\hat p_z(1-\hat p_y)(1-\hat p_z)/n^2\) to \(\widehat w\). That smaller-order term does not change the first-order asymptotics, but it can shift a finite-sample statistic. When results depend on the implementation, state which formula and software version you used rather than comparing p-values as if every package used identical finite-sample arithmetic.
The formula is not a cure for a weak design. A tiny sample, nearly constant forecasts, empty cells, or a zero or invalid estimated variance can make the usual approximation unreliable or undefined. In such cases, do not force a p-value out of the expression; report the marginal shares and table, and choose an inference procedure suited to the data and sample size.
For the hypothetical table above, the leading-formula components are \(\widehat v=0.54(0.46)/100=0.002484\) and \(\widehat w=[0.2^2(0.7)(0.3)+0.4^2(0.6)(0.4)]/100=0.000468\). The standard error is \(\sqrt{0.002484-0.000468}\approx0.0449\), giving \(PT\approx0.10/0.0449=2.23\). A two-sided standard-normal reference would give a p-value near 0.026. This is a calculation from invented counts under the leading asymptotic formula; the fuller finite-sample term changes the number slightly, and serial dependence can invalidate the normal reference. Do not present it as an empirical result.

Read a rejection as evidence about direction, not as a trading verdict
A positive numerator means observed directions agree more often than the independence baseline; a negative one means less often. A pre-specified one-sided test can ask whether agreement is greater than the baseline. A two-sided test asks whether directional association differs in either direction. The alternative and significance level should be chosen before inspecting results.
A small p-value is evidence against the stated null under the test assumptions. It is not the probability that the null is true, nor the probability that the forecast will work next period. It does not say that returns were large enough to cover spread, fees, market impact, funding, or borrow costs. Nor does it correct for trying many assets, horizons, thresholds, model variants, or forecast signs and then reporting only the winner. If candidate strategies were searched, the White Reality Check guide explains why selection must be reflected in inference.
Conversely, failing to reject is not proof that the directions are independent. A short or imbalanced sample may have little power to detect an association. A negative PT numerator also does not prove that a forecast contains no structure at all: it can indicate systematic disagreement. Reversing the forecast sign after seeing that result is a new model choice and must be evaluated on fresh data or included in the original search procedure.
Report the actual-up share, forecast-up share, observed hit rate, independence baseline, percentage-point lift, alternative hypothesis, and uncertainty method together. A p-value without the table margins and effect size makes it hard to distinguish a small but precisely estimated lift from a larger, noisier one.
Serial dependence can make the textbook reference distribution misleading
The ordinary PT statistic is commonly presented for serially independent directional observations. Financial labels may instead arrive in runs. Daily observations can share overlapping multi-day targets; volatility regimes can persist; and rolling forecasts can reuse most of the same estimation data. Then the \(n\) pairs are not \(n\) independent pieces of information. The usual standard error can be too small, making rejection too frequent.
Pesaran and Timmermann later developed tests for dependence among serially correlated multi-category variables using dynamically augmented regressions and a finite-order Markov framework {source:pesaranTimmermannSerialCategories2009}. Work focused on directional forecasts also finds that conventional independence-based PT procedures can be oversized when serial correlation is present, and studies dependence-robust alternatives including bootstrap methods {source:blaskowitzHerwartzDirectionalSerialCorrelation2014}. The appropriate method depends on how forecasts are generated, the forecast horizon, and the dependence structure; a generic resampling recipe is not automatically valid.
State whether forecast origins overlap, how far apart they are, and what dependence method supplies the uncertainty estimate. If you report the textbook statistic as a descriptive benchmark, label its independence assumption. Do not interpret a nominal 5% result as having a 5% false-rejection rate when the design violates that assumption.
Build the evaluation around forecasts that could have been made
For each origin, preserve the forecast exactly as it was available at that time, the forecast horizon, the information cutoff, the realized outcome, and the rule that converts each into an up/down label. Align timestamps, instruments, price or return definition, and missing-value treatment before forming the table. A forecast based on revised macroeconomic data or a signal tuned on the evaluation period is not an untouched out-of-sample forecast.
Choose the classification threshold before seeing the holdout results. If a model emits a probability, the threshold converts it into a binary direction; searching thresholds on the same sample changes the question and creates selection bias. Keep training, validation, and final evaluation roles separate, and rerun the full selection process inside any resampling procedure that is meant to account for model search.
The PT question differs from comparing magnitudes of forecast errors. The Diebold–Mariano guide compares paired loss differences, while the Giacomini–White guide asks whether relative predictive ability varies with pre-specified information. For nested forecast models, see the Clark–West adjustment guide. These tests answer different questions and should not be swapped just because they return familiar test statistics.
Directional significance does not establish a profitable strategy
The PT test reduces each outcome to a direction label. A one-tick rise and a large rally count as the same up result; a small miss and a severe loss count equally as direction errors. A forecast can therefore improve its hit rate while losing money if its misses are larger than its wins, its signal arrives after a price move, or trading costs consume the edge.
For example, imagine 100 equal-size hypothetical trades with 64 correct directions earning an average of 0.10% each and 36 incorrect directions losing an average of 0.25% each. Their simple gross sum is \(64(0.10\%)-36(0.25\%)=-2.6\) percentage points before costs, despite a 64% hit rate. This deliberately simplified arithmetic ignores compounding and is not market evidence; it shows why hit rate alone does not determine expectancy.
To evaluate a trade, specify entry and exit timing, position size, risk limits, and treatment of unfilled orders before calculating returns. Include bid–ask spread, commissions, slippage, market impact, financing, and venue-specific costs where relevant. Compare net returns with an appropriate benchmark out of sample and account for the same multiple-testing search that generated the rule. The PT result can be one piece of evidence about directional association; it is not a substitute for that evaluation.
Common questions
Q1Does the Pesaran–Timmermann test compare accuracy with 50%?
No. It compares observed agreement with the independence baseline calculated from the separate actual-up and forecast-up shares. That baseline can be above or below 50%.
Q2Can I use the usual PT p-value when forecast horizons overlap?
Not without addressing dependence. Overlapping targets and persistent labels can make the standard independence-based uncertainty estimate too small. Use a method whose assumptions fit the forecast horizon and dependence structure.
Q3Does a significant PT result mean the trading strategy is profitable?
No. The test uses direction labels and does not measure move size, entry and exit prices, position sizing, fees, slippage, or funding. Evaluate net out-of-sample returns separately.
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
If actual outcomes are up 60% of the time and forecasts call up 70% of the time, what is the independence agreement baseline?
Choose an answer to see the explanation
Options glossary
Clear definitions of essential option terms, from calls, puts, and option chains to IV, Greeks, open interest, and max pain
Browse the options glossary