टू-स्केल रियलाइज़्ड वोलैटिलिटी: माइक्रोस्ट्रक्चर शोर के साथ विचरण का अनुमान
जानें कि टू-स्केल रियलाइज़्ड वोलैटिलिटी, स्वतंत्र माइक्रोस्ट्रक्चर शोर के bias को घटाने के लिए सूक्ष्म और विरल price grids को कैसे जोड़ती है—काल्पनिक उदाहरण, मान्यताओं और सीमाओं सहित
इस गाइड मेंउच्च-आवृत्ति observations बढ़ने पर variance estimate खराब क्यों हो सकता है
संक्षिप्त सारांश
टू-स्केल रियलाइज़्ड वोलैटिलिटी (TSRV), सभी observations से बने fine-grid estimate को कई अधिक विरल और shifted sampling grids के औसत के साथ जोड़ती है। निर्दिष्ट additive और independent noise model में fine-grid estimate, sparse grids के noise bias का अनुमान लगाकर उसे घटाने में मदद करता है। लक्ष्य integrated variance है; यह noisy observations को सटीक कीमतों में नहीं बदलता और न ही अगले return का forecast करता है।
उच्च-आवृत्ति observations बढ़ने पर variance estimate खराब क्यों हो सकता है
रियलाइज़्ड variance अक्सर squared returns को जोड़कर निकाली जाती है। यदि observed log prices, efficient price को बिना error के दिखाते, तो अधिक observations से उसके path को बेहतर पकड़ा जा सकता था। लेकिन वास्तविक transaction prices और quotes में bid–ask bounce, price discreteness, latency और अन्य market-microstructure effects भी होते हैं। बहुत छोटे intervals में ये प्रभाव दो observations के बीच efficient-price move से बड़े हो सकते हैं।
इसलिए sampling frequency बढ़ने पर naive realized variance भी बढ़ सकती है, भले session की underlying integrated variance न बदली हो। यह अपने-आप में इस बात का प्रमाण नहीं कि asset की आर्थिक volatility बढ़ी है; noisy price differences को square करके बार-बार जोड़ने से ऐसा हो सकता है। सामान्य sum-of-squared-returns measure के लिए realized-volatility calculation guide देखें। यह लेख high-frequency data में bias के एक विशिष्ट स्रोत पर केंद्रित है।
Zhang, Mykland और Aït-Sahalia ने दो resolutions को अलग-अलग काम में लगाने के लिए TSRV प्रस्तुत किया (दो time scales पर मूल अध्ययन)। Sparse grid, noise-affected increments की संख्या घटाता है। अलग-अलग starting points वाली sparse grids का औसत किसी एक मनमाने start point पर निर्भरता कम करता है। Full fine grid, घटाए जाने वाले noise bias का scaled estimate देता है। नतीजा assumptions और finite-sample choices पर निर्भर है; यह uncertainty वाला estimator है, latent path का noise-free observation नहीं।
Efficient log price और observation noise अलग रखें
मानें \(X_t\) latent efficient log price है, जिसे इस window में local variance rate \(\sigma_t^2\) वाले continuous process से model किया गया है। लक्ष्य इसका integrated variance है:
\[ IV_T=\int_0^T\sigma_t^2\,dt. \]
Observation time \(t_i\) पर दर्ज log price मानें
\[ Y_{t_i}=X_{t_i}+\epsilon_{t_i}, \]
जहाँ \(\epsilon_{t_i}\) observation noise है। Classical derivation के baseline model में noise का mean zero, observation times के बीच independent, efficient-price process से independent और variance स्थिर \(\omega^2\) होता है। यह regular sampling grid और window में efficient price के लिए उचित continuous-process model भी मानता है। ये modeling conditions हैं, हर market dataset का सार्वभौमिक वर्णन नहीं।
Observed increment है \(\Delta Y_i=\Delta X_i+\epsilon_i-\epsilon_{i-1}\)। एक observation का pricing error पास-पास के दो returns में उलटे signs के साथ आता है: positive error एक return बढ़ाता और अगला घटाता है। इसलिए noisy prices से बने returns independent measurement errors जैसे व्यवहार नहीं करते। Noise-free setting में squared increments का sum path variation को क्यों target करता है, यह finance में quadratic variation guide समझाती है।
यहाँ लक्ष्य एक तय समय-window की variance है। यह bid–ask spread estimate, forecast या realized execution cost नहीं है। Transaction-price के अलग प्रश्न के लिए Roll bid–ask spread estimator देखें। लक्ष्य स्पष्ट रखना जरूरी है, क्योंकि कई microstructure methods high-frequency prices का उपयोग करते हैं, पर अलग-अलग quantities estimate करते हैं।
सभी observations की realized variance noise से क्यों प्रभावित होती है
Fine grid में \(n\) returns हों, तो all-observation realized variance है:
\[ RV_{\mathrm{all}}=\sum_{i=1}^{n}(Y_{t_i}-Y_{t_{i-1}})^2. \]
Baseline model में एक observed return का noise component \(\eta_i=\epsilon_i-\epsilon_{i-1}\) है और इसकी variance \(\operatorname{Var}(\eta_i)=2\omega^2\)। Independence assumptions के तहत जिन cross-terms का expectation zero है उन्हें छोड़ें, तो fine-grid sum का expected value लगभग
\[ E[RV_{\mathrm{all}}]\approx IV_T+2n\omega^2 \]
है।
पहला term target integrated variance है; दूसरा हर fine return से noise contribution जोड़ता है। Sampling interval घटने और \(n\) बढ़ने पर noise term हावी हो सकता है। इसलिए measurement noise को अनदेखा करते हुए “जितनी बार हो सके sample करें” सुरक्षित नियम नहीं। Aït-Sahalia, Mykland और Zhang ने microstructure noise के बीच सीमित optimal sampling frequency का अध्ययन किया।
यह approximation model के तहत expectation बताती है। किसी वास्तविक sample में efficient-price increments और noise increments के random cross-terms भी होते हैं, इसलिए realized value expectation से अलग हो सकती है। Intraday noise variance बदलती हो, errors serially dependent हों या sampling irregular हो, तो सरल \(2n\omega^2\) bias को ठीक से न दर्शाए। तब जानी-पहचानी formula भी सटीक दिखने वाला, पर गलत correction दे सकती है।
एक मनमानी grid चुनने के बजाय shifted sparse grids का औसत लें
Integer spacing \(K>1\) चुनें। हर offset \(k=0,\ldots,K-1\) के लिए हर \(K\)-वें observed price लेकर sparse-grid realized variance बनाएं:
\[ RV_{K,k}=\sum_j\left(Y_{t_{k+jK}}-Y_{t_{k+(j-1)K}}\right)^2. \]
हर sparse increment लंबा interval cover करता है, लेकिन grid लगभग वही observation window cover करता है। उपलब्ध offsets का औसत लें:
\[ \overline{RV}_K=\frac{1}{K}\sum_{k=0}^{K-1}RV_{K,k}. \]
Offsets मायने रखते हैं। एक sparse grid बड़ी price move से ठीक पहले या बाद शुरू हो सकती है, जिससे result उस मनमाने चुनाव के प्रति sensitive होगा। Shifted grids का औसत start-point dependence घटाता है। इससे \(K\) independent datasets नहीं बनते: grids उन्हीं observations को दोबारा उपयोग करते हैं और सांख्यिकीय रूप से dependent होते हैं।
Independent और equal-variance observation noise में sparse return के दोनों endpoints के noise values अलग observations से आते हैं, इसलिए अंतर की variance अब भी \(2\omega^2\) है। Sparse returns की संख्या fine returns से कम होती है, इसलिए sparse average में कुल noise bias कम जमा होता है। Sparse increments window को cover करते हैं, उसका केवल हिस्सा नहीं; इसलिए signal contribution लगभग वही integrated variance रहती है।
इस लेख का illustration अवधारणात्मक है: यह jagged observed path की तुलना smoother latent path के आसपास कई shifted sparse views से करता है। यह market data या किसी estimator का output नहीं है। High-low observations पर आधारित estimator के लिए OHLC range-based volatility estimators guide देखें।
<!-- learn:illustration -->

Scaled fine variance घटाएँ और finite-sample normalization लगाएँ
Shifted grids में sparse returns की औसत संख्या \(\bar n\) मानें और \(\lambda=\bar n/n\) रखें। Baseline model में average sparse-grid noise bias लगभग \(2\bar n\omega^2\) है। Fine-grid noise bias लगभग \(2n\omega^2\) है, इसलिए \(\lambda RV_{\mathrm{all}}\) का contribution लगभग \(2\bar n\omega^2\) होगा। इसे घटाने से leading noise term cancel होती है:
\[ TSRV_{\mathrm{raw}}=\overline{RV}_K-\lambda RV_{\mathrm{all}}. \]
यह raw difference integrated-variance signal का \(\lambda\) हिस्सा भी घटाता है, क्योंकि दोनों realized-variance terms में signal है। एक सामान्य finite-sample normalization है:
\[ \widehat{IV}_{TSRV}=\frac{\overline{RV}_K-\lambda RV_{\mathrm{all}}}{1-\lambda}. \]
सरल expectation calculation में denominator signal coefficient को वापस लाता है। Implementations में conventions और endpoint handling अलग हो सकते हैं; इसलिए \(n\) की परिभाषा, offset set और \(\bar n\) की गणना बताएं। यह correction sampling uncertainty नहीं मिटाती और गलत noise model भी ठीक नहीं करती। Classical asymptotic setup में sample बढ़ने पर \(\lambda\) छोटा होता है, पर छोटे window में finite-sample factor महत्वपूर्ण हो सकता है।
जब \(\overline{RV}_K<\lambda RV_{\mathrm{all}}\), estimate negative हो सकती है। इसका अर्थ physical variation negative होना नहीं; finite sample में noise subtraction sparse estimate से बड़ी हो गई। Downstream display में zero दिखाना सुविधाजनक हो सकता है, लेकिन उससे estimator बदलता है और truncation rule जुड़ता है। Analysis में raw value रखें, scales और assumptions जाँचें, और बाद में truncation करें तो स्पष्ट बताएं।
छह returns का काल्पनिक हिसाब, चरण-दर-चरण
मानें सात hypothetical observed log-price points से bp में छह fine returns मिले: \([1,1,-1,-1,1,1]\) bp। यह छोटी sequence केवल arithmetic दिखाने के लिए है, exchange data नहीं। Fine-grid realized variance:
\[ RV_{\mathrm{all}}=1^2+1^2+(-1)^2+(-1)^2+1^2+1^2=6\,\mathrm{bp}^2. \]
\(K=2\) चुनें। पहला offset पास-पास के fine returns जोड़कर sparse returns \([2,-2,2]\) bp बनाता है, जिनका squared sum \(4+4+4=12\,\mathrm{bp}^2\) है। Shifted offset \([0,0]\) bp देता है, squared sum शून्य। दोनों grid estimates का औसत \(\overline{RV}_K=(12+0)/2=6\,\mathrm{bp}^2\) है।
छह fine returns और दो offsets के साथ sparse returns की औसत संख्या \(\bar n=(3+2)/2=2.5\) है। इसलिए \(\lambda=2.5/6=5/12\), और
\[ \widehat{IV}_{TSRV}=\frac{6-(5/12)6}{1-5/12}=\frac{3.5}{7/12}=6\,\mathrm{bp}^2. \]
Variance estimate \(6\,\mathrm{bp}^2=6\times10^{-8}\) decimal-return-squared units में है। इस hypothetical window में इसका square root लगभग \(2.45\) bp है; annualized नहीं। इस बनाई गई छोटी sequence में fine-grid RV से संख्यात्मक समानता TSRV की identity नहीं। दूसरी price paths से estimate अधिक, कम या negative भी हो सकती है।
यह गणना model की सीमा भी दिखाती है: latent path और noise अलग से दिखाई नहीं देते, इसलिए यह sample साबित नहीं करता कि \(6\,\mathrm{bp}^2\) actual integrated variance है। Sum, scale और units audit किए जा सकते हैं, पर hidden decomposition उदाहरण से नहीं निकाली जा सकती। Literature में realized-kernel जैसे अन्य high-frequency noise corrections भी हैं; उनके अपने assumptions और implementation choices हैं।
दोनों scales क्या estimate करते हैं और K कैसे चुनें
\(K\), sparse returns की संख्या, averaging और noise correction के बीच trade-off बनाता है। बहुत छोटा \(K\) कई छोटे increments छोड़ता है और estimate noise-sensitive रह सकती है। बहुत बड़ा \(K\) कुछ coarse returns ही छोड़ता है, जिससे sparse estimate unstable और finite-window effects के प्रति अधिक exposed होती है। Offset averaging grid alignment में मदद करती है, लेकिन छोटे या खराब sample में अनुपस्थित information नहीं बनाती।
मूल asymptotic analysis में sample size के साथ पसंदीदा sparse spacing बढ़ती है; उस paper की notation और baseline conditions में इसका order \(n^{2/3}\) है (मूल विश्लेषण)। यह theoretical scaling result है, किसी तय seconds या observations की universal सलाह नहीं। Optimum noise-to-signal ratio और assumptions पर निर्भर करता है; बाद के research और practical implementations अलग tuning rules और endpoint conventions अपना सकते हैं। Reproduction के लिए sampling interval, \(K\), offsets, session boundaries और finite-sample normalization लिखें।
एक defensible analysis पहले से तय किए गए plausible \(K\) values की तुलना कर sensitivity दिखा सकती है, बजाय बाद में सबसे आकर्षक result देने वाली setting चुनने के। Estimates में बड़ा बदलाव instability का संकेत है और उसे भी result में बताएं। Scale data design और पहले से documented rule से चुनें। Coarser realized variance या noise-robust alternative से तुलना उपयोगी हो सकती है, पर measures अलग-अलग label करें।
Classical TSRV assumptions कब टूटती हैं
Classical correction एक खास noise structure के लिए निकाली गई है। Bid–ask bounce और दूसरे microstructure effects serially dependent, session के दौरान बदलते, liquidity पर निर्भर या efficient-price increments से correlated हो सकते हैं। Hansen और Lunde ने realized variance और microstructure noise पर अपने अध्ययन में time dependence और efficient-price changes के साथ correlation जैसी interactions दर्ज कीं। Aït-Sahalia, Mykland और Zhang ने बाद में dependent microstructure noise के लिए two-scale methods विकसित कीं। यह extension modified estimator और original iid-noise formula में भेद करने का कारण है, यह मानने का नहीं कि original formula हर dependence संभालती है।
अन्य समस्याओं में irregular observation times, asynchronous assets, stale quotes, price discreteness, jumps, bad prints, opening/closing effects और missing observations शामिल हैं। ये fine-grid correction और sparse grids में शामिल returns दोनों बदल सकते हैं। Research question के अनुसार jumps target quadratic variation का हिस्सा हो सकती हैं या अलग करनी पड़ सकती हैं; TSRV यह modeling choice नहीं करती। Observations filter करने से grid भी बदलती है, जिसे document करें।
Noise robustness का मतलब हर data problem से immunity नहीं। Estimate को latent variance कहने से पहले तय करें कि input transaction prices, quotes या midpoints हैं; return transform और clock बताएं; समान spacing जाँचें; cleaning, session rules और \(K\)-sensitivity दर्ज करें। Noise process baseline assumptions तोड़े तो उसी स्थिति के लिए बना method लें और उसकी assumptions बताएं। Negative estimate को negative volatility न कहें, और चुपचाप positive value में न बदलें।
Estimator को forecast बताए बिना report करें
Transparent report में target window, input price series, sampling interval, fine-return count, sparse spacing \(K\), offset set, endpoint rule, noise assumptions, normalization, units और कोई truncation शामिल हों। Raw difference और normalized estimate साथ दिखाए जा सकते हैं, साथ में पहले से तय scales की sensitivity range। इन विवरणों से calculation audit होती है और variance estimate को quote, future forecast या profitable trading rule का प्रमाण समझने की भूल घटती है।
Integrated variance observed window पर मापी जाती है। Annualize करने के लिए time convention बताएं और intraday patterns, overnight intervals या trading gaps छिपने न दें। Square root लेने से उसी window के लिए volatility units मिलती हैं; prediction horizon नहीं बनता। Historical realized measure किसी model का input हो सकती है, पर predictive performance को estimator settings चुनने में उपयोग न किए गए out-of-sample data पर अलग से जाँचें।
TSRV का योगदान methodological है: explicit additive-noise model में दो sampling scales का उपयोग naive high-frequency realized variance के leading noise bias को घटाता है। यह हर price movement का source नहीं बताता, model risk नहीं हटाता और अगली volatility नहीं बताता। Number के साथ ये boundaries भी दें।
आम सवाल
Q1क्या TSRV सभी observations का उपयोग करती है?
Fine-grid realized variance में सभी observations उपयोग होती हैं, और shifted sparse grids भी उन्हीं observations को दोबारा लेती हैं। हर sparse grid अलग से कम frequency पर sample करती है। विधि इन हिस्सों को जोड़ती है; fine observations को discard नहीं करती और grids को independent samples नहीं बनाती।
Q2क्या two-scale realized volatility, volatility forecast है?
नहीं। इसका सामान्य target observed window की integrated variance है। Square root उसे उसी window की volatility units में व्यक्त करती है; बाद की अवधि forecast करने के लिए अलग model और out-of-sample evaluation चाहिए।
Q3क्या negative TSRV estimate को zero कर देना चाहिए?
बिना बदलाव बताए नहीं। Scaled fine-grid correction sparse estimate से बड़ी होने पर finite-sample value negative हो सकती है। उसे clip करने से estimator बदलती है और averages/inference पर असर पड़ सकता है; raw value और बाद का rule बताएं।
Q4क्या original TSRV formula हर microstructure noise संभालती है?
नहीं। Classical derivation independent observation errors सहित baseline noise model मानती है। समय के साथ बदलता या serially dependent noise, irregular sampling, अथवा noise और efficient-price changes की dependence के लिए दूसरा या modified estimator चाहिए हो सकता है।
स्रोत और आगे पढ़ें
समस्या की रिपोर्ट करें
हम इस लेख का लिंक जोड़कर ईमेल तैयार करेंगे। भेजने के बाद ही Mark को आपकी रिपोर्ट मिलेगी
त्वरित जाँच
गाइड पढ़ने के बाद 3 सवालों से खुद को जाँचें
सवाल 01
Sampling interval बहुत छोटा होने पर naive realized variance क्यों बढ़ सकती है?
व्याख्या देखने के लिए एक उत्तर चुनें
विकल्प शब्दावली
Volatility calculated from price changes that occurred under a stated return, sampling-window, and annualization rule; different conventions can produce different values.
विस्तृत गाइड पढ़ेंQuadratic variationThe limit of sums of squared process increments over increasingly fine partitions, measuring accumulated second-order path variation and generating Itô corrections.
विस्तृत गाइड पढ़ेंIV rankThe current IV's position between a lookback-period low and high; one outlier can distort it, so it should not be read as a standalone signal.
विस्तृत गाइड पढ़ें