Skip to content
All option guides
Estimate integrated variance when high-frequency prices contain noise14 min read

Two-Scale Realized Volatility: Estimating Variance with Microstructure Noise

See how two-scale realized volatility uses fine and sparse price grids to reduce independent microstructure-noise bias, with a worked example, assumptions, and limits

In this guideWhy more high-frequency observations can make a variance estimate worse

Short summary

Two-scale realized volatility (TSRV) combines a noisy estimate from every observation with an average over many sparser sampling grids. Under a stated additive, independent-noise model, the fine-grid estimate helps estimate and remove the sparse-grid noise bias. The method estimates integrated variance; it does not turn noisy observations into exact prices or predict the next return.

Why more high-frequency observations can make a variance estimate worse

Realized variance is often calculated by adding squared returns. If observed log prices followed the efficient price perfectly, using more observations could capture more of its path. Actual transaction prices and quotes also reflect bid–ask bounce, discreteness, latency, and other market microstructure effects. At very short intervals, these effects can be large compared with the efficient price move between two observations.

The naive realized variance can therefore rise as observations become more frequent, even when the underlying integrated variance over the session has not changed. This is not evidence that the asset’s economic volatility necessarily rose. It can be a consequence of squaring noisy price differences and summing them many times. The general [realized-volatility calculation guide](/en/learn/realized-volatility-calculation) explains the ordinary sum-of-squared-returns measure; this article focuses on correcting one particular source of high-frequency bias.

Two-scale realized volatility, introduced by Zhang, Mykland, and Aït-Sahalia in their original two-time-scale study, uses two resolutions for different purposes. A sparse grid reduces the number of noise-contaminated increments. An average across shifted sparse grids avoids relying on one arbitrary starting point. The full fine grid supplies a scaled noise estimate that can be subtracted. The logic depends on assumptions and finite-sample choices, so the result is an estimator with uncertainty, not a cleaned observation of the true path.

Separate the efficient log price from observation noise

Let \(X_t\) denote the latent efficient log price, modeled over the window as a continuous process with local variance rate \(\sigma_t^2\). The target is its integrated variance,

\[ IV_T=\int_0^T\sigma_t^2\,dt. \]

At observation time \(t_i\), suppose the recorded log price is

\[ Y_{t_i}=X_{t_i}+\epsilon_{t_i}, \]

where \(\epsilon_{t_i}\) is observation noise. The classical derivation uses a baseline in which noise is mean zero, independent across observation times, independent of the efficient-price process, and has constant variance \(\omega^2\). It also assumes a regular sampling grid and a suitable continuous efficient-price model over the window. These are modeling conditions, not universal descriptions of exchange data.

The observed increment is \(\Delta Y_i=\Delta X_i+\epsilon_i-\epsilon_{i-1}\). Noise enters adjacent returns with opposite signs: a positive pricing error at one observation raises one return and lowers the next. That dependence is why returns made from noisy prices do not behave like independent measurement errors. The broader [quadratic-variation guide](/en/learn/quadratic-variation-finance-explained) explains why sums of squared increments target path variation in the noise-free setting.

The target here is variance over a specified time window. It is not a bid–ask spread estimate, a forecast, or a realized execution cost. For a different transaction-price question, see the [Roll bid–ask spread estimator](/en/learn/roll-bid-ask-spread-autocovariance-estimator-explained). Keeping the target clear matters because several market microstructure procedures use high-frequency prices but estimate different quantities.

Why the all-tick realized variance becomes noise-dominated

For \(n\) fine-grid returns, define

\[ RV_{\mathrm{all}}=\sum_{i=1}^{n}(Y_{t_i}-Y_{t_{i-1}})^2. \]

Under the baseline model, the noise part of one observed return is \(\eta_i=\epsilon_i-\epsilon_{i-1}\). Its variance is \(\operatorname{Var}(\eta_i)=2\omega^2\). Ignoring cross-terms that have zero expectation under the independence assumptions, the expected fine-grid sum is approximately

\[ E[RV_{\mathrm{all}}]\approx IV_T+2n\omega^2. \]

The first term is the desired integrated variance. The second accumulates one noise contribution per fine return. As the sampling interval shrinks and \(n\) grows, the noise term can dominate. This is the central reason “sample as often as possible” is not a safe rule when measurement noise is ignored. Aït-Sahalia, Mykland, and Zhang analyze this finite optimal sampling issue in their study of sampling in the presence of microstructure noise.

The approximation is an expectation statement under a model. A particular sample also contains random cross-terms between efficient-price and noise increments, so its realized value need not equal the expectation. If noise variance changes intraday, errors are serially dependent, or sampling is irregular, the simple \(2n\omega^2\) expression need not describe the bias. In those cases a familiar formula can produce a precise-looking number while targeting the wrong adjustment.

Average shifted sparse grids instead of choosing one arbitrary grid

Choose an integer spacing \(K>1\). For each offset \(k=0,\ldots,K-1\), form a sparse-grid realized variance by taking every \(K\)-th observed price:

\[ RV_{K,k}=\sum_j\left(Y_{t_{k+jK}}-Y_{t_{k+(j-1)K}}\right)^2. \]

Each sparse increment spans a longer interval, but the grid still covers the same broad observation window. Average over available offsets:

\[ \overline{RV}_K=\frac{1}{K}\sum_{k=0}^{K-1}RV_{K,k}. \]

The offsets matter. A single sparse grid could begin just before or after a large move and give a result sensitive to that choice. Averaging shifted grids reduces this starting-point dependence. It does not create \(K\) independent datasets: the grids reuse observations and are statistically dependent.

Under independent, equal-variance observation noise, the difference between two noise terms at the endpoints of a sparse return still has variance \(2\omega^2\), because its endpoints are distinct observations. There are fewer sparse returns than fine returns, so the average sparse estimate carries less accumulated noise bias. Its signal contribution remains approximately the same integrated variance because the sparse increments span the window rather than only a fraction of it.

The figure accompanying this guide is conceptual: it contrasts a jagged observed path with coarser shifted views around a smoother latent path. It does not depict market data or a measured estimator output. For estimators built from high-low observations rather than this two-grid construction, see the [range-based volatility estimator guide](/en/learn/range-based-volatility-estimators-ohlc-explained).

<!-- learn:illustration --> <!-- Text-free concept: a jagged observed price path and shifted sparse sampling grids around a smoother latent path. Conceptual; no market data or estimator output. -->

A jagged observed price path and several wider, shifted sampling intervals arranged around a smoother latent path
Conceptual illustration of fine observations with microstructure noise and overlapping sparse grids. It is not market data or an estimator output.

Subtract the scaled fast variance and apply finite-sample normalization

Let \(\bar n\) be the average number of sparse returns across the offset grids, and set \(\lambda=\bar n/n\). Under the baseline model, the average sparse-grid noise bias is approximately \(2\bar n\omega^2\). The fine-grid estimate has noise bias approximately \(2n\omega^2\), so \(\lambda RV_{\mathrm{all}}\) contributes approximately \(2\bar n\omega^2\). Subtracting it cancels the leading noise term:

\[ TSRV_{\mathrm{raw}}=\overline{RV}_K-\lambda RV_{\mathrm{all}}. \]

This raw difference also subtracts a fraction \(\lambda\) of the integrated-variance signal, since both realized-variance terms contain signal. A common finite-sample normalization is therefore

\[ \widehat{IV}_{TSRV}=\frac{\overline{RV}_K-\lambda RV_{\mathrm{all}}}{1-\lambda}. \]

The denominator restores the signal coefficient under the simple expectation calculation. Conventions and endpoint handling vary across implementations, so the exact definition of \(n\), the set of offsets, and \(\bar n\) should be documented. This correction does not remove sampling uncertainty or repair a violated noise model. In the classic asymptotic setup, \(\lambda\) becomes small as the sample grows; the finite-sample factor can still matter in a short window.

The estimate can be negative when \(\overline{RV}_K<\lambda RV_{\mathrm{all}}\). A negative result is not negative physical variation; it means the finite-sample noise subtraction exceeded the sparse-grid estimate. Replacing it with zero may be convenient for a downstream display, but it changes the estimator and adds a truncation rule. Preserve and report the raw estimate during analysis, investigate the chosen scales and assumptions, and state any later truncation explicitly.

A six-return hypothetical calculation, step by step

Consider seven hypothetical observed log-price points, which produce six fine returns in basis points: \([1,1,-1,-1,1,1]\) bp. This deliberately small sequence is only arithmetic illustration; it is not exchange data. The fine-grid realized variance is

\[ RV_{\mathrm{all}}=1^2+1^2+(-1)^2+(-1)^2+1^2+1^2=6\,\mathrm{bp}^2. \]

Set \(K=2\). The first offset groups neighboring fine returns into sparse returns \([2,-2,2]\) bp, with squared sum \(4+4+4=12\,\mathrm{bp}^2\). The shifted offset yields \([0,0]\) bp, with squared sum \(0\). Averaging the two grid estimates gives \(\overline{RV}_K=(12+0)/2=6\,\mathrm{bp}^2\).

With six fine returns and these two offsets, the average number of sparse returns is \(\bar n=(3+2)/2=2.5\). Thus \(\lambda=2.5/6=5/12\), and

\[ \widehat{IV}_{TSRV}=\frac{6-(5/12)6}{1-5/12}=\frac{3.5}{7/12}=6\,\mathrm{bp}^2. \]

The variance estimate is \(6\,\mathrm{bp}^2=6\times10^{-8}\) in decimal-return-squared units. Its square root is about \(2.45\) bp over this hypothetical window; it is not annualized. The numerical equality with the fine-grid RV is a feature of this small constructed sequence, not an identity of TSRV. Different price paths can produce a higher, lower, or even negative estimate.

The arithmetic also makes the model’s limitation visible: because the true latent path and noise are not separately observed, this sample alone cannot show that \(6\,\mathrm{bp}^2\) equals the actual integrated variance. One can audit the sums, scale, and units, but not infer the hidden decomposition from the example. For another family of high-frequency noise corrections, realized-kernel methods are discussed in the empirical literature; those estimators have their own assumptions and implementation choices.

What the two scales estimate—and how to choose K

The scale \(K\) trades off the number of sparse returns against the averaging and noise correction. A very small \(K\) leaves many short increments and may retain considerable noise sensitivity. A very large \(K\) leaves few coarse returns, making the sparse estimate unstable and more exposed to finite-window effects. Offset averaging helps with grid alignment, but it cannot supply information absent from a short or poor-quality sample.

In the original asymptotic analysis, the preferred sparse spacing grows with sample size; under that paper’s notation and baseline conditions it is of order \(n^{2/3}\). This is a theoretical scaling result, not a universal instruction to use a particular number of seconds or observations. The optimum depends on the noise-to-signal ratio and assumptions; later work and practical implementations may use different tuning rules and endpoint conventions. Report the sampling interval, \(K\), offsets, session boundaries, and finite-sample normalization so another analyst can reproduce the estimate.

A defensible analysis can compare a prespecified set of plausible \(K\) values and show the sensitivity rather than selecting the setting that produces the most attractive result. If estimates shift materially, that instability is part of the finding. The scale should be selected using the data design and a documented rule, not after looking for the preferred volatility estimate. A useful comparison can include a coarser realized variance or a noise-robust alternative, but the measures should be labeled separately.

When classical TSRV assumptions break

The classical correction is derived for a specific noise structure. Bid–ask bounce and other microstructure effects can be serially dependent, vary through the trading session, depend on liquidity, or correlate with efficient-price increments. Hansen and Lunde document in their realized-variance study that realized variance and microstructure noise can interact in ways that include time dependence and correlation with efficient-price changes. Aït-Sahalia, Mykland, and Zhang later develop two-scale methods for dependent microstructure noise; that extension is a reason to distinguish the modified estimator from the original iid-noise formula, not to assume the original formula handles every dependence pattern.

Other practical issues include irregular observation times, asynchronous assets, stale quotes, price discreteness, jumps, bad prints, opening and closing effects, and missing observations. These can affect both the fine-grid correction and which returns appear on the sparse grids. Jumps may be part of the target quadratic variation or may need to be separated depending on the research question; TSRV alone does not decide that modeling choice. Filtering out observations also changes the grid and should be documented.

Noise robustness is not immunity to all data problems. Before treating an estimate as a measure of latent variance, define whether the input is transaction prices, quotes, or midpoint prices; state the return transform and clock; check whether observations are equally spaced; record cleaning and session rules; and assess whether the estimated value is sensitive to \(K\). If the noise process violates the baseline assumptions, use a method designed for that setting and explain its assumptions. Do not describe a negative estimate as negative volatility or silently turn it into a positive number.

Report the estimator without presenting it as a forecast

A transparent report names the target window, input price series, sampling interval, fine-return count, sparse spacing \(K\), offset set, endpoint rule, noise assumptions, normalization, units, and any truncation. It can show both the raw difference and normalized estimate, plus a sensitivity range over prespecified scales. These details make the calculation auditable and stop a variance estimate from being mistaken for a quote, a future forecast, or evidence of a profitable trading rule.

Integrated variance is measured over the observed window. Annualizing it requires a stated time convention and should not conceal intraday patterns, overnight intervals, or gaps in trading. Taking its square root changes the units to volatility over the window; it does not itself create a forecast horizon. A historical realized measure can be one input to a model, but predictive performance must be evaluated separately on data not used to choose the estimator settings.

The main contribution of TSRV is methodological: under an explicit additive-noise model, it uses two sampling scales to reduce the leading noise bias that makes naive high-frequency realized variance inconsistent. It does not identify the source of every price fluctuation, eliminate model risk, or say what volatility will be next. Keep those boundaries alongside the reported number.

Common questions

Q1Does TSRV use every observation?

It uses all observations in the fine-grid realized variance and reuses observations across shifted sparse grids. Each sparse grid individually samples less often. The method combines these pieces; it does not discard the fine observations or turn the grids into independent samples.

Q2Is two-scale realized volatility a volatility forecast?

No. The standard target is integrated variance over the observed window. Taking a square root expresses its magnitude in volatility units for that window, but forecasting a later period requires a separate model and out-of-sample evaluation.

Q3Should a negative TSRV estimate be changed to zero?

Not without stating the change. A negative finite-sample value can arise when the scaled fine-grid correction exceeds the sparse estimate. Clipping it changes the estimator and can affect averages and inference; report the raw value and any downstream rule.

Q4Does the original TSRV formula handle all microstructure noise?

No. Its classic derivation uses a baseline noise model, including independent observation errors. Time-varying or serially dependent noise, irregular sampling, and dependence between noise and efficient-price changes may require a different or modified estimator.

Sources and further reading

Report an issue

We’ll prepare an email with this article link. Mark receives the report only after you send it

Quick check

Read the guide? Check yourself with 3 questions

Question 1 / 3

Question 01

Why can naive realized variance rise when the sampling interval gets very short?

Choose an answer to see the explanation

Options glossary