VPIN: फ़ॉर्मूला, volume buckets और Flash Crash पर बहस
जानें कि VPIN order-flow imbalance को summarize करने के लिए volume-time buckets और Bulk Volume Classification का उपयोग कैसे करता है, एक hypothetical calculation देखें, और 2010 Flash Crash से जुड़े विवादित evidence को समझें।
इस गाइड मेंVPIN क्या मापता है — और इसका नाम क्या बढ़ा-चढ़ाकर बता सकता है
संक्षिप्त सारांश
VPIN हाल के equal-volume buckets में buy–sell volume imbalance के absolute size का rolling summary है। इसके लेखकों ने इसे volume time में order-flow toxicity monitor करने के लिए प्रस्तावित किया, लेकिन इसके inputs और interpretation पर विवाद है। यह score सीधे observed probability नहीं है कि कोई trade informed था, न ही Flash Crash की proven warning या अपने-आप में trading signal।
VPIN क्या मापता है — और इसका नाम क्या बढ़ा-चढ़ाकर बता सकता है
VPIN का पूरा नाम volume-synchronized probability of informed trading है। Easley, López de Prado और O’Hara ने इसे order-flow toxicity से जुड़ा high-frequency measure बताया: ऐसी स्थिति जिसमें liquidity providers को adverse selection का सामना हो सकता है, क्योंकि incoming trades के informed side से आने की संभावना अधिक होती है। उनका method volume imbalance को trading intensity के साथ जोड़ता है और fixed wall-clock intervals के बजाय volume clock पर update होता है (Easley, López de Prado, and O’Hara, 2012)।
नाम को सावधानी से समझें। VPIN का calculated value 0.32 होने का अर्थ यह नहीं कि अगला trade informed trader से आने की calibrated 32% संभावना है। यह चुने हुए data, classification, bucket और window rules के तहत estimated imbalances का normalized average है। यह नहीं बताता कि कौन trader informed है, information-motivated trading और liquidity-motivated trading में फर्क नहीं करता, और यह साबित नहीं करता कि liquidity provider को वास्तव में नुकसान हुआ।
VPIN signed order-flow imbalance से भी अलग है। हर bucket का absolute imbalance पहले लिया जाता है और फिर average किया जाता है, इसलिए buy-heavy और sell-heavy दोनों buckets positive योगदान देते हैं। इससे हाल के volume intervals में flow कितना one-sided था, उसका सार मिलता है, लेकिन final score में imbalance की direction हट जाती है। Classic PIN model guide trade arrivals पर आधारित संबंधित measure समझाती है.
Volume-time buckets updates को trading activity से जोड़ते हैं
Wall-clock chart हर तय अंतराल पर observations लेता है, जैसे हर second या minute। Volume clock तब आगे बढ़ता है जब तय trading volume पूरा हो जाए। यदि bucket 1,000 contracts का है, तो वह हर 1,000 contracts trade होने पर बंद होता है—busy period में इसमें 20 seconds लग सकते हैं और quiet period में 12 minutes।
हर bucket का intended total traded volume एक जैसा होता है, लेकिन वह जितना wall-clock time कवर करता है, वह बदलता रहता है। फिर VPIN हाल के \(n\) volume buckets की rolling window से calculate होता है और नया bucket पूरा होने पर update होता है। इसलिए active periods में updates ज्यादा और quiet periods में कम होते हैं।
Volume time sampling convention है; यह data से market activity को हटाने का तरीका नहीं। Trading तेज़ होने पर VPIN window कितनी जल्दी आगे बढ़ती है, यह बदल सकता है; प्रति घंटे observations की संख्या भी बदलती है। इसलिए results instrument, session, bucket volume, bucket boundary पार करने वाले trades के treatment और volume clock के initialization पर निर्भर करते हैं। Reproducible analysis में इन choices को बताना चाहिए।
Bulk Volume Classification groups में trade direction का अनुमान लगाती है
VPIN को buyer-initiated और seller-initiated volume चाहिए। यदि trade-level direction labels उपलब्ध और विश्वसनीय हों, तो उनकी quality को ध्यान में रखते हुए उन्हें aggregate किया जा सकता है। जब भरोसेमंद labels उपलब्ध नहीं होते, original VPIN paper ने Bulk Volume Classification (BVC) सुझाई: time bar में price change के आधार पर उस bar के volume में buy और sell की संभावित हिस्सेदारी का अनुमान लगाना।
BVC को आम तौर पर इस तरह लिखा जाता है:
\[ \widehat V^{B}_t=V_t\Phi\!\left(\frac{\Delta P_t}{\widehat\sigma_{\Delta P}}\right), \qquad \widehat V^{S}_t=V_t-\widehat V^{B}_t, \]
जहाँ \(V_t\) bar का total volume है, \(\Delta P_t\) उसका price change है, \(\widehat\sigma_{\Delta P}\) price changes के scale का बताया गया estimate है, और \(\Phi\) standard normal cumulative distribution function है। Positive standardized price change ज्यादा volume buys को देता है; negative change ज्यादा sells को। Standardized change शून्य हो तो formula volume को बराबर बाँटता है।
उदाहरण के लिए, यदि hypothetical time bar में 1,000 contracts हैं और standardized price change \(0.5\) है, तो \(\Phi(0.5)\approx0.6915\)। BVC लगभग 691.5 contracts buy volume और 308.5 sell volume को assign करती है। Fractional values estimates हैं; ये देखे गए fractional contracts या ज्ञात trade intentions नहीं। Result bar interval, price measure, volatility-scale estimate और distribution assumption पर निर्भर है।
BVC price movement से volume को groups में classify करती है; यह हर trade की direction स्थापित नहीं करती। इस estimate के लिए इस्तेमाल time bars और बाद के equal-volume buckets अलग-अलग steps हैं। Study को दोनों document करने चाहिए। Andersen और Bondarenko की 2015 comparison ने उनके E-mini S&P 500 futures benchmark में quotes और trades से बनाई गई trade classification के विरुद्ध BVC को standard tick rule से कमजोर पाया (Andersen and Bondarenko, 2015)। Lee–Ready trade-direction classification guide एक अन्य method और उसकी सीमाएँ बताती है।
VPIN formula हर bucket का absolute imbalance average करता है
मान लें \(V\) प्रति bucket target volume है, और \(\widehat V^B_i\), \(\widehat V^S_i\), bucket \(i\) में estimated buy और sell volumes हैं। हर complete bucket के लिए \(\widehat V^B_i+\widehat V^S_i=V\)। एक सामान्य rolling form है:
\[ \mathrm{VPIN}_t= \frac{1}{nV} \sum_{i=t-n+1}^{t} \left|\widehat V^B_i-\widehat V^S_i\right|. \]
Numerator हाल के \(n\) buckets में हर bucket के buy–sell difference का absolute value जोड़ता है। Denominator \(nV\) उन buckets का total volume है। Complete buckets में यदि nonnegative buy/sell estimates का sum \(V\) है, तो score 0 और 1 के बीच रहता है; कुछ displays इसे 100 से गुणा करके percentage दिखाते हैं।
शून्य के पास value का अर्थ है कि उन buckets में estimated buy और sell volumes अपेक्षाकृत balanced थे। बड़ी value बताती है कि rolling volume window में estimated flow average में अधिक one-sided था। Score यह नहीं बताता कि one-sided flow buying था या selling, क्योंकि absolute value sign हटा देता है।
अलग implementations trade-direction method, bucket size, rolling-window length, sampling interval और price input में भिन्न हो सकते हैं। व्यवहार में ये choices statistic की definition का हिस्सा हैं। अलग settings से निकले values अपने-आप comparable नहीं होते।
चार buckets का hypothetical calculation दोहराया जा सकता है
मान लें चार complete buckets में हर एक में \(V=100\) contracts हैं। बताई गई एक trade-classification procedure लगाने के बाद estimated buys और sells ये हैं:
| Bucket | Estimated buys | Estimated sells | Absolute imbalance |
|---|---|---|---|
| 1 | 70 | 30 | \(|70-30|=40\) |
| 2 | 55 | 45 | \(|55-45|=10\) |
| 3 | 20 | 80 | \(|20-80|=60\) |
| 4 | 60 | 40 | \(|60-40|=20\) |
चार absolute imbalances का योग \(40+10+60+20=130\) contracts है। Total volume \(4\times100=400\) contracts है, इसलिए
\[ \mathrm{VPIN}=\frac{130}{400}=0.325=32.5\%. \]
उसी window में estimated buys 205 और sells 195 हैं। Signed net imbalance केवल \(205-195=10\) contracts, यानी total volume का \(10/400=2.5\%\), है। यह VPIN से अलग है, क्योंकि signed net में positive और negative bucket imbalances आंशिक रूप से cancel हो जाते हैं, जबकि VPIN उनके absolute sizes जोड़ता है।
यहाँ सभी bucket volumes और classifications hypothetical हैं। Arithmetic formula समझाता है; यह किसी वास्तविक asset, exchange, probability या market event का estimate नहीं है। यदि estimates BVC से आए, तो classification step दोहराने के लिए उसके price bars और volatility scale भी बताने होंगे।
<!-- learn:illustration --> <!-- बिना text का वैचारिक चित्र: समान volume वाले buckets में अनुमानित खरीद और बिक्री के अनुपात अलग हैं; rolling summary असंतुलन का परिमाण रखती है, लेकिन उसकी दिशा हटा देती है। यह market data, crash probability, forecast या trading signal नहीं है। -->

Flash Crash के evidence पर प्रकाशित बहस हुई है
Easley, López de Prado और O’Hara के 2012 study ने VPIN को order-flow-toxicity measure के रूप में पेश किया और short-term volatility से इसका संबंध बताया (Easley, López de Prado, and O’Hara, 2012)। उनके 2011 के पहले analysis ने 6 मई 2010 Flash Crash से पहले elevated readings का भी वर्णन किया (Easley, López de Prado, and O’Hara, 2011)। यही कारण है कि VPIN को संभावित market-stress monitor की तरह देखा गया, लेकिन इससे यह साबित नहीं होता कि वह crashes को भरोसेमंद ढंग से predict करती है।
Andersen और Bondarenko ने अपनी 2014 reconstruction में अलग conclusion निकाला। उन्होंने report किया कि VPIN short-horizon volatility की कमजोर predictor थी, Flash Crash के बाद ही maximum पर पहुँची, और इसका predictive content काफी हद तक trading intensity से जुड़ा था। उनका analysis कुछ खास trade-classification procedures और implementations को परखता है; यह नहीं बताता कि हर market में VPIN का हर version एक जैसा व्यवहार करता है (Andersen and Bondarenko, 2014)।
Easley, López de Prado और O’Hara ने 2014 में जवाब दिया। उन्होंने replication और interpretation से असहमति जताई; उनका तर्क था कि analysis ने उनकी methodology या conclusions को test नहीं किया और comparison में market microstructure मायने रखता है (Easley, López de Prado, and O’Hara, 2014)। यह exchange implementation और empirical interpretation पर असहमति दिखाता है, VPIN को early-warning signal के रूप में सर्वसम्मत validation नहीं।
बाद के assessment में Andersen और Bondarenko ने E-mini S&P 500 futures के trade classification को benchmark करने के लिए quote और trade data इस्तेमाल किया। उन्होंने बताया कि उनके benchmark में BVC के बजाय tick rule बेहतर था और तर्क दिया कि उनके setting में volatility-related classification errors VPIN की apparent volatility prediction को समझा सकते हैं (Andersen and Bondarenko, 2015)। यह measurement की गंभीर चिंता है, क्योंकि VPIN estimated buy/sell volume पर निर्भर है; फिर भी यह निर्दिष्ट data और procedures का empirical finding है।
Ke और Lin का 2017 paper maximum-likelihood estimation पर आधारित VPIN का वैकल्पिक formulation विकसित करता है। उनका तर्क है कि छोटे volume buckets और कम informed trades में original metric unstable हो सकती है; वे अपनी alternative method से अधिक consistent estimates report करते हैं (Ke and Lin, 2017)। यह contribution दिखाता है कि model construction मायने रखती है; यह तय नहीं करता कि VPIN का कोई version सभी assets, venues और market regimes में भरोसेमंद real-time warning है।
कुल मिलाकर, literature इस सरल दावे का समर्थन नहीं करता कि VPIN ने हर relevant episode को “predict” किया या “predict करने में विफल” रही। अलग classification choices, data, implementations और tests से researchers को अलग results मिले। यह असहमति खुद construction report करने और incremental performance को simple activity तथा volatility benchmarks के विरुद्ध जाँचने का कारण है।
Bucket size, classification और sampling score बदल सकते हैं
Bucket volume \(V\) तय करता है कि हर observation में कितना trading समेटा जाए। छोटे buckets updates ज्यादा बार देते हैं, लेकिन estimates को individual bursts और classification noise के प्रति अधिक sensitive बना सकते हैं। बड़े buckets local variation smooth करते हैं और update frequency घटाते हैं। सही scale market की सामान्य activity और research question पर निर्भर है; crash देखने के बाद चुनी गई value apparent warning performance को बढ़ा सकती है।
Rolling window \(n\) तय करता है कि कितना recent volume शामिल होगा। छोटा window तेज़ प्रतिक्रिया देता है, पर noisy हो सकता है; बड़ा window smooth और धीमा है। \(V\) और \(n\), दोनों के साथ equivalent wall-clock duration distribution भी report करें। बराबर संख्या के buckets calm और active periods में बराबर elapsed time नहीं होते।
BVC की अपनी choices हैं: time-bar length, trade या quote price input, volatility-estimation window, zero price changes का treatment और assumed distribution। Trade-level signs इस्तेमाल करने पर भी quote timing, midpoint comparisons, tick rules या venue-specific message conventions से classification errors मायने रख सकती हैं। 2015 study की critique उसके tested setting में volatility changes को systematic BVC classification errors से जोड़ती है। Order book आधारित order-flow imbalance displayed supply और demand में बदलाव मापती है; यह इस तरह classify किए गए trade volume से अलग measure है.
Initialization भी मायने रखता है। Rolling VPIN window किसी bucket boundary से शुरू होती है। Starting point shift करने से कौन-से trades साथ रखे गए, यह बदल सकता है और bucket imbalances बदल सकते हैं। Tests को इस sensitivity को quantify करना चाहिए; वे उस alignment को नहीं चुनें जो event से पहले सबसे तेज rise दिखाता है। Trading hours, tick size, contract design या venue mix में structural changes समय के साथ comparability को और घटा सकते हैं।
Warning role देने से पहले incremental evidence जाँचें
एक सावधान VPIN study को target event देखने से पहले data और algorithm specify करना चाहिए। Market और contract, session, trade/quote filters, price input, trade-side procedure या BVC settings, \(V\), \(n\), initialization rule, और partial या boundary-crossing buckets के treatment को बताएं। अलग activity levels में VPIN window कितना wall-clock time कवर करती है, यह भी दिखाएँ।
Predictive claim के लिए chronological out-of-sample design अपनाएँ और VPIN की तुलना आसानी से दिखने वाले baselines—जैसे volume, trade intensity, realized volatility और order imbalance—से करें। इन quantities को शामिल करने के बाद देखें कि VPIN अतिरिक्त information देती है या नहीं। False alarms, missed events, lead time, uncertainty और अलग thresholds पर results report करें; सिर्फ ऐसा retrospective chart नहीं जिसमें चुने हुए crash से पहले indicator बढ़ा हो। Event को देखकर चुना गया threshold advance warning साबित नहीं करता।
Flow toxicity के claim के लिए adverse selection या liquidity outcomes को स्वतंत्र रूप से मापें, जैसे subsequent price impact, spreads, depth या market-maker के realized losses, और बताएं कि इन्हें कैसे identify किया गया। VPIN और volatility का correlation अकेले toxicity interpretation validate नहीं करता। बड़ा score ज्यादा one-sided classified volume, BVC को प्रभावित करती ऊँची volatility या तेज़ चलती volume clock से आ सकता है।
VPIN को contested diagnostic मानें, trading signal नहीं
VPIN descriptive summary के रूप में उपयोगी हो सकती है, यदि शोधकर्ता equal-volume windows में estimated absolute order-flow imbalance track करना चाहते हों। Transparent construction और simpler measures के विरुद्ध टिकने वाले incremental contribution के साथ यह बड़े analysis का feature भी हो सकती है। इसका अर्थ सीमित है: यह हाल के volume की एक खास classification summarize करती है।
ऊँचा score अकेले यह साबित नहीं करता कि informed traders active हैं, liquidity crisis आने वाली है, prices किसी तय दिशा में चलेंगे या position घटाने-बढ़ाने से profit होगा। कम score सुरक्षित या liquid market का प्रमाणपत्र नहीं। किसी operational use के लिए independent validation, उचित controls, explicit decision rule, latency-aware data और realistic execution-cost analysis चाहिए।
VPIN को उसकी measurement assumptions और contrary evidence के साथ प्रस्तुत करें। जब पाठक volume clock, BVC या alternative classifier, rolling window और out-of-sample benchmarks देख सकें, तब वे आँक सकते हैं कि statistic क्या दिखाती है और inference कहाँ समाप्त होती है।
तुलना के लिए, Roll spread estimator return की autocovariance से spread का एक proxy निकालता है और अलग सवाल का जवाब देता है.
आम सवाल
Q1क्या VPIN की 0.32 reading का मतलब informed trading की 32% संभावना है—या मुझे trading रोक देनी चाहिए?
नहीं। सामान्य score अनुमानित absolute volume imbalance का normalized average है; यह किसी trade या trader के informed होने की calibrated probability नहीं है। यह direction भी नहीं बताता, आसन्न liquidity crisis साबित नहीं करता, और execution cost या सीमाओं को शामिल नहीं करता। इसलिए केवल ऊँची reading का अर्थ trading रोकना नहीं है; किसी भी decision rule के लिए अलग out-of-sample validation और risk controls चाहिए।
Q2क्या VPIN equal time intervals इस्तेमाल करती है?
नहीं। इसकी rolling observations equal traded-volume amounts से बनती हैं, इसलिए हर bucket की wall-clock duration market activity के साथ बदलती है। BVC पहले अलग time bars में signed volume estimate कर सकती है।
Q3क्या VPIN ने 6 मई 2010 के Flash Crash को predict किया था?
Published evidence contested है। Easley, López de Prado और O’Hara ने elevated readings को possible warning value का evidence बताया; Andersen और Bondarenko ने अपनी reconstruction में measure का peak event के बाद report किया और predictive content पर सवाल उठाया। बाद की exchange और assessment methods तथा interpretation पर सहमत नहीं हैं।
Q4क्या एक ही market के दो VPIN charts सीधे compare किए जा सकते हैं?
तभी जब data, price inputs, trade classification, volume-bucket size, rolling window, initialization और construction की दूसरी choices पर्याप्त रूप से aligned हों। अलग settings अलग scores और update speeds दे सकती हैं।
स्रोत और आगे पढ़ें
समस्या की रिपोर्ट करें
हम इस लेख का लिंक जोड़कर ईमेल तैयार करेंगे। भेजने के बाद ही Mark को आपकी रिपोर्ट मिलेगी
त्वरित जाँच
गाइड पढ़ने के बाद 3 सवालों से खुद को जाँचें
सवाल 01
VPIN अपने recent volume buckets में किसका average लेती है?
व्याख्या देखने के लिए एक उत्तर चुनें