Johansen Cointegration Test: Rank, Trace और Maximum-Eigenvalue
समझें कि Johansen test का cointegration rank क्या गिनता है, trace और maximum-eigenvalue tests कैसे अलग हैं, और lag तथा deterministic terms क्यों मायने रखते हैं।
इस गाइड मेंCointegration rank स्वतंत्र stationary संबंधों की संख्या है
संक्षिप्त सारांश
Johansen procedure बहुचर समय-श्रृंखला सिस्टम में long-run matrix के rank की जाँच करती है। rank \(r\) चुने हुए model के तहत परस्पर स्वतंत्र stationary combinations की संख्या है। इससे अपने-आप tradable portfolio की पहचान, causality का प्रमाण या अनुमानित संबंध के आगे बने रहने की गारंटी नहीं मिलती।
Cointegration rank स्वतंत्र stationary संबंधों की संख्या है
मान लें \(y_t\) में \(k\) variables हैं और बताए गए data-generating setup के तहत हर variable को \(I(1)\) माना गया है। Cointegration rank \(r\), उनके levels के linear combinations में परस्पर स्वतंत्र stationary संबंधों की संख्या है। \(r=0\) होने पर उस specification में ऐसा कोई combination नहीं मिला। \(0<r<k\) होने पर \(r\) स्वतंत्र long-run relations होते हैं और सामान्य \(I(1)\) system conditions में \(k-r\) common stochastic trends बने रहते हैं। rank 1 का अर्थ पूरे system में एक relation है; यह ज़रूरी नहीं कि किसी एक स्पष्ट asset pair का संबंध हो।
यह पूरे system का सवाल है। दो-step Engle–Granger procedure एक normalized long-run equation estimate करके उसके fitted residual को test करती है। Johansen की likelihood method vector autoregression में cointegrating space का dimension estimate और test करती है, इसलिए एक से अधिक स्वतंत्र relation भी दिखा सकती है। दोनों तरीके संबंधित लेकिन अलग सवालों के जवाब देते हैं; finite samples या अलग lag और deterministic specifications में नतीजे अलग हो सकते हैं।
VECM में long-run matrix का rank दिखाई देता है
यहाँ deterministic terms छोड़कर, levels में VAR को सरल vector error-correction form में लिखा जा सकता है:
\[ \Delta y_t = \Pi y_{t-1} + \sum_{j=1}^{p-1}\Gamma_j\Delta y_{t-j}+\varepsilon_t, \qquad \Pi=\alpha\beta^{\mathsf T}. \]
यहाँ \(p\), levels VAR का lag order है। \(\beta\) के columns long-run combinations बताते हैं; \(\alpha\) में adjustment loadings होते हैं, जो पिछले disequilibrium को मौजूदा बदलाव से जोड़ते हैं। जब \(\Pi\) का rank \(r\) होता है, तो उसे \(k\times r\) के \(\alpha\) और \(\beta\) matrices में factor किया जा सकता है। System में error-correction terms की संख्या केवल variables की संख्या नहीं, बल्कि matrix का rank तय करता है।
यदि सभी \(k\) series वास्तव में \(I(1)\) हैं, तो सामान्य cointegrated मामलों में \(r<k\)। Full-rank नतीजा संकेत देता है कि levels, fitted specification में शामिल deterministic terms के आसपास stationary हो सकते हैं; मूल integration-order premise को फिर जाँचें। rank zero यह साबित नहीं करता कि variables के बीच किसी भी तरह का संबंध नहीं है। इसका अर्थ है कि चुने गए model ने जाँचे गए rank पर stationary level combination नहीं पाया।

Trace और maximum-eigenvalue tests के alternatives अलग हैं
अनुमानित eigenvalues को \(1>\hat\lambda_1\ge\hat\lambda_2\ge\cdots\ge\hat\lambda_k\ge0\) के क्रम में रखें। ये likelihood procedure की reduced-rank समस्या से आते हैं; इन्हें variance explained के सामान्य प्रतिशत न समझें। Candidate rank \(r\) पर trace statistic है:
\[ \operatorname{LR}_{\mathrm{trace}}(r)=-T\sum_{i=r+1}^{k}\ln(1-\hat\lambda_i), \]
यह \(H_0:\operatorname{rank}(\Pi)\le r\) के विरुद्ध \(H_A:\operatorname{rank}(\Pi)>r\) को test करता है और बाकी eigenvalues के साक्ष्य जोड़ता है। Maximum-eigenvalue statistic है:
\[ \operatorname{LR}_{\max}(r,r+1)=-T\ln(1-\hat\lambda_{r+1}), \]
यह rank \(r\) की तुलना अगले rank \(r+1\) से करता है। इसलिए trace test पूछता है कि कुल relations की संख्या \(r\) से ज़्यादा है या नहीं, जबकि maximum-eigenvalue test पूछता है कि अगला relation जोड़ने के पक्ष में डेटा है या नहीं। Rank null के लिए किसी भी statistic का सामान्य chi-square reference distribution नहीं है।
काल्पनिक eigenvalue नतीजे से गणना देखें
तीन variables, 200 उपयोगी observations और क्रमशः \(0.16\), \(0.05\), \(0.01\) eigenvalues वाला system कल्पना करें। ये केवल समझाने के लिए बनाए गए मान हैं, market estimates नहीं। \(r=0\) पर trace statistic तीनों terms जोड़कर लगभग \(47.14\) है, जबकि maximum-eigenvalue statistic केवल पहले term से लगभग \(34.87\) है। \(r=1\) पर ये लगभग \(12.27\) और \(10.26\) हैं; \(r=2\) पर दोनों अंतिम eigenvalue का उपयोग करते हैं और लगभग \(2.01\) होते हैं।
यह गणना test statistics देती है, rank का निर्णय नहीं। निर्णय के लिए statistic की तुलना ठीक उसी deterministic case और test sequence के critical values या p-values से करें। उदाहरण के लिए, software table में “\(r=0\) reject” का अर्थ उस table की assumptions और चुने significance level के तहत rank zero के विरुद्ध evidence है; यह किसी खास asset spread के stationary होने की सीधी probability नहीं। Cutoff से बड़ा p-value बताता है कि उस rank null को reject नहीं किया गया, यह नहीं कि वह सच साबित हो गया।
Lag order और deterministic terms तय करते हैं कि कौन-सा system test हो रहा है
Rank की व्याख्या से पहले उचित VAR lag order चुनें। Levels VAR का order \(p\), VECM में \(p-1\) lagged differences के अनुरूप है। बहुत कम lags से residuals में serial correlation बच सकती है; बहुत अधिक lags degrees of freedom खर्च करके finite-sample inference को अस्थिर कर सकते हैं। Information criteria संभावित orders व्यवस्थित करने में मदद करते हैं, लेकिन residual diagnostics, sample size और आर्थिक timing भी मायने रखते हैं। Order चुनने का तरीका बताएं और उचित विकल्प जाँचें।
तय करें कि intercepts और trends cointegrating relation के बाहर होंगे, उसके भीतर सीमित होंगे, या maintained model में नहीं होंगे। इससे long-run behavior और rank-test distribution दोनों बदलते हैं। सबसे अनुकूल p-value ढूँढकर specification न चुनें; सवाल और डेटा के अनुसार चुनें। Critical values ठीक उसी deterministic case, variables की संख्या और implementation से मेल खाने चाहिए। Johansen rank statistics की distribution nonstandard है। MacKinnon, Haug और Michelis ने निर्दिष्ट Johansen-type likelihood-ratio tests के लिए response-surface distributions निकालीं, जिनमें exogenous \(I(1)\) variables की अनुमति देने वाला विस्तार भी शामिल है। उनके values हर rank-test setup पर लागू नहीं होते।
Deterministic term की जगह software का केवल सजावटी विकल्प नहीं है। आम parameterization में cointegrating relation के भीतर सीमित constant उस equilibrium combination का nonzero mean होने देता है; relation के बाहर unrestricted constant levels में अलग deterministic drift की अनुमति दे सकता है। Trend restrictions भी model द्वारा माने गए long-run path को बदलती हैं। सिर्फ “constant included” लिखना अधूरा है; constant कहाँ आता है और हर restriction क्या है, बताएं।
Estimated vectors को normalization और आर्थिक व्याख्या चाहिए
मुख्य estimate cointegrating space है। Rank-one model में \(\beta\) vector को किसी भी nonzero constant से गुणा करने पर यह नहीं बदलता कि कौन-सा level combination stationary है। इसलिए analyst normalization चुनता है, जैसे एक coefficient को 1 रखना। Higher rank में एक ही space को span करने वाले कई bases हो सकते हैं; software के दिखाए vectors अपने-आप अनोखे आर्थिक संबंध नहीं बन जाते। किसी खास vector की व्याख्या के लिए restrictions या अतिरिक्त theory चाहिए, और normalization भी बताना चाहिए।
Adjustment matrix \(\alpha\) बताती है कि fitted system में disequilibrium पर कौन-से variables प्रतिक्रिया करते हैं। Sign चुने हुए normalization पर निर्भर है; इसकी entries system के conditional coefficients हैं, स्वतंत्र causal effects नहीं। छोटा या शून्य loading weak-exogeneity hypothesis सुझा सकता है, पर निष्कर्ष के लिए maintained model के तहत संबंधित formal restrictions test करनी होंगी। केवल cointegration rank से यह नहीं पता चलता कि कौन-सा variable “lead” करता है, कौन-सा संबंध आर्थिक रूप से अर्थपूर्ण है, या shock भविष्य के prices को कैसे प्रभावित करेगा।
यह nonuniqueness केवल display का मामला नहीं, matrix factorization की विशेषता है। Nonsingular \(H\) के लिए \(\beta H\) और \(\alpha(H^{-1})^{\mathsf T}\) एक ही long-run matrix \(\Pi\) देते हैं। इसलिए estimated space identified हो सकता है, भले दिखाए गए अलग-अलग vectors unique न हों। किसी vector को tradable spread चुनने के लिए स्पष्ट normalization या economic restriction चाहिए; उसकी units और risk का अलग विश्लेषण भी करें।
Sequential decisions और model uncertainty rank के दावे को सीमित करते हैं
आम तौर पर candidate ranks को शून्य से क्रम में test करते हैं और पहले ऐसे null पर रुकते हैं जिसे reject न किया जाए। चुना हुआ rank, चयनित significance level और specification के तहत model choice है। Trace और maximum-eigenvalue नतीजे अलग हों तो मतभेद छिपाएँ नहीं; lag length, deterministic terms, कम test power, structural breaks और sample sensitivity जाँचें और दोनों नतीजे दें।
Likelihood-ratio critical values nonstandard हैं और system specification पर निर्भर करते हैं। Unit root के पास का व्यवहार, छोटा sample, outliers, regime changes, \(I(2)\) variables या exogenous integrated regressors साधारण \(I(1)\) setup को अनुपयुक्त बना सकते हैं। सामान्य rank table इन मामलों को अपने-आप शामिल नहीं करती। Johansen का मूल काम Gaussian VAR assumptions में likelihood inference विकसित करता है; बाद की response-surface गणनाएँ निर्दिष्ट test setups के लिए reference distributions देती हैं, हर data problem के लिए नहीं।
उदाहरण के लिए, यदि trace test \(r=0\) और \(r=1\) को reject करे, लेकिन \(r=2\) को reject न करे, तो सामान्य sequential procedure उस specification और significance level के तहत rank 2 चुनती है। इसका अर्थ यह नहीं कि rank 2 सही होने की probability 1 है। Maximum-eigenvalue sequence पास-पास के ranks के null hypotheses की तुलना करती है, इसलिए उसका निर्णय भी दिखाएँ। एक test की critical value दूसरे statistic पर लागू न करें।
चुना हुआ rank sample period पर भी निर्भर है। Economic relation, market structure या variable definitions बदलने पर शुरुआती और बाद के sample के estimates अलग हो सकते हैं। Sensitivity analysis अस्थिरता दिखा सकती है, लेकिन कई endpoints खोजना एक और multiple-testing समस्या बनाता है। पहले से तय windows बताएं और बाद में आज़माए गए splits को exploratory evidence मानें।
Cointegration rank अपने-आप pairs-trading signal नहीं है
Trading study में estimated \(\beta\) vector spread की एक संभावित परिभाषा हो सकता है, लेकिन rank entry और exit levels, position sizes या expected returns नहीं देता। कई cointegrating vectors हों तो किसी combination का चयन normalization, constraints और data selection पर निर्भर हो सकता है। Statistical relation estimate होने पर भी weights short positions, leverage या कम liquidity ला सकते हैं।
हर decision date पर उपलब्ध जानकारी से ही system estimate करें और selection व re-estimation की पूरी प्रक्रिया को समय-क्रम में test करें। Sample के बाहर संबंध स्थिर है या नहीं जाँचें; ज़रूरत के अनुसार dividends, contract rolls, currency conversion, financing, borrow availability, bid–ask spreads, market impact और execution timing शामिल करें। Rank null reject करना चुने model के बारे में evidence है, profitability claim नहीं। Engle–Granger test और error correction single-equation विकल्प समझाता है; pairs trading में cointegration बनाम correlation बताता है कि correlation अलग गुण क्यों है।
Rank test दोहराने लायक पर्याप्त जानकारी दें
\(k\) variables, transformations, observation frequency, sample dates और integration-order evidence बताएं। Levels VAR lag order, deterministic terms की जगह, effective sample size, trace और maximum-eigenvalue statistics, मेल खाते critical values या p-values, significance level और sequential rank decisions दें। बताएँ कि दोनों tests सहमत थे या नहीं, आगे कौन-सा rank लिया और sensitivity checks से नतीजा बदला या नहीं।
Vectors दिखाते समय normalization, restrictions और संबंधित adjustment loadings दें। Trading application में हर rebalance पर उपलब्ध डेटा, vector-selection rule, out-of-sample design, costs और failure handling भी बताएँ। Johansen का 1988 cointegration-vector analysis90041-3) cointegrating space और rank के likelihood approach को विकसित करता है। उनका 1991 Gaussian VAR लेख rank inference और vectors से जुड़ी hypotheses पर चर्चा करता है। MacKinnon, Haug और Michelis ने cointegration likelihood-ratio tests की response-surface distributions दी हैं। संबंधित मार्गदर्शिका stationarity और unit roots समझाती है।
आम सवाल
Q1क्या rank 1 का अर्थ है कि ठीक दो assets एक pair बनाते हैं?
नहीं। इसका अर्थ पूरे system में एक स्वतंत्र stationary combination है। उसमें कई variables हो सकते हैं और अर्थ निकालने के लिए समझने योग्य normalization चाहिए।
Q2क्या trace और maximum-eigenvalue tests को एक ही rank चुनना चाहिए?
नहीं। उनके alternatives अलग हैं, इसलिए नतीजे अलग हो सकते हैं। दोनों बताएँ और model, lag, deterministic terms तथा sample sensitivity जाँचें।
Q3क्या rank zero को reject करना साबित करता है कि spread लाभदायक होगा?
नहीं। यह चुने model के तहत no-cointegration null के विरुद्ध evidence है। इससे trading rules, sample के बाहर स्थिरता या execution costs के बाद net returns नहीं मिलते। ---
स्रोत और आगे पढ़ें
समस्या की रिपोर्ट करें
हम इस लेख का लिंक जोड़कर ईमेल तैयार करेंगे। भेजने के बाद ही Mark को आपकी रिपोर्ट मिलेगी
त्वरित जाँच
गाइड पढ़ने के बाद 3 सवालों से खुद को जाँचें
सवाल 01
I(1) माने गए तीन-variable system में cointegration rank 1 का क्या अर्थ है?
व्याख्या देखने के लिए एक उत्तर चुनें
विकल्प शब्दावली
The existence of a stationary linear combination among nonstationary series, implying a shared long-run equilibrium restriction under a specified model.
विस्तृत गाइड पढ़ेंMean reversionA model-dependent tendency for a variable to move back toward a fixed or changing reference; it does not automatically imply stationarity or tradability.
विस्तृत गाइड पढ़ेंएट द मनीवह स्थिति जिसमें विकल्प का स्ट्राइक मूल्य अंतर्निहित परिसंपत्ति के बाजार मूल्य के बहुत करीब हो। उसमें अभी महत्वपूर्ण आंतरिक मूल्य न हो, फिर भी बचे हुए समय और अनिश्चितता के कारण प्रीमियम हो सकता है।
विस्तृत गाइड पढ़ें