कमज़ोर उपकरणों की जाँच: प्रथम-चरण F और मजबूत अनुमान
जानें कि प्रथम-चरण F, आंशिक R-वर्ग, Stock–Yogo मानदंड, मजबूत निदान और Anderson–Rubin अनुमान कमज़ोर उपकरणों से जुड़े अलग-अलग प्रश्नों के उत्तर कैसे देते हैं।
इस गाइड मेंउपकरण की शक्ति क्यों मायने रखती है
संक्षिप्त सारांश
प्रथम-चरण आँकड़ा बताता है कि शामिल नियंत्रणों के बाद अपवर्जित उपकरण किसी अंतर्जात regressor को कितना समझाते हैं। यह उपकरण की वैधता सिद्ध नहीं करता; “F 10 से ऊपर” भी कोई सार्वभौमिक गारंटी नहीं है। उपयुक्त निदान अंतर्जात regressors की संख्या और error assumptions पर निर्भर करता है; कमज़ोर IV के प्रति मजबूत अनुमान भी चौड़ा या असीमित रह सकता है।
उपकरण की शक्ति क्यों मायने रखती है
Instrumental variables, अंतर्जात regressor \(X\) में उस बदलाव का उपयोग करते हैं जिसका अनुमान शामिल controls \(W\) को ध्यान में रखते हुए अपवर्जित instruments \(Z\) से लगाया जाता है।
यदि \(Z\), \(X\) के शेष बदलाव का बहुत कम हिस्सा समझाता है, तो इस स्रोत से structural effect के बारे में data में जानकारी सीमित होती है।
कमज़ोर relevance में finite sample 2SLS अनुमान OLS की ओर biased हो सकता है और उसका sampling distribution normal approximation से काफी अलग हो सकता है। Standard error सटीक दिखने पर भी conventional Wald interval की coverage खराब हो सकती है।
Staiger और Stock weak-instrument asymptotics से इन विफलताओं को समझाते हैं और nonstandard confidence sets की आवश्यकता बताते हैं (1997 का लेख)।
शक्ति identification का केवल एक हिस्सा है। मजबूत first stage यह नहीं दिखाता कि \(Z\), structural error से स्वतंत्र है या outcome को केवल \(X\) के माध्यम से प्रभावित करता है।
instrumental variables मार्गदर्शिका इन अलग assumptions और पहचाने जा रहे प्रभाव पर चर्चा करती है।
प्रथम चरण शर्तबद्ध उपकरण-विविधता अलग करता है
एक अंतर्जात regressor और \(q\) अपवर्जित instruments के लिए first stage लिखें:
\[ X_t = W_t'\delta + Z_t'\pi + v_t. \]
यहाँ \(W_t\) में शामिल exogenous controls होते हैं, अक्सर intercept सहित; \(Z_t\) में अपवर्जित instruments होते हैं। मानक first-stage F test संयुक्त प्रतिबंध \(H_0:\pi=0\) जाँचता है: \(W_t\) के बाद अपवर्जित instruments कोई अतिरिक्त explanatory power नहीं जोड़ते।
यह conditional relevance का प्रश्न है, exclusion या independence का test नहीं। बताएँ कि कौन-से controls शामिल हैं, कितने अपवर्जित instruments को संयुक्त रूप से जाँचा गया, sample और covariance assumptions क्या हैं, और कौन-सा statistic लिया गया।
कई अपवर्जित instruments हों तो किसी एक coefficient का t statistic संयुक्त test का विकल्प नहीं है।
आंशिक R-वर्ग और पारंपरिक F आँकड़ा
आंशिक \(R^2\) बताता है कि \(W\) हटाने के बाद \(X\) में बची variation का कितना भाग residualized अपवर्जित instruments समझाते हैं।
मानें \(R^2_p\) partial \(R^2\), \(n\) sample size, \(k\) शामिल regressors की संख्या (यदि intercept है तो उसे भी गिनें), और \(q\) अपवर्जित instruments की संख्या है।
पारंपरिक homoskedastic linear-regression गणना में incremental F statistic है:
\[ F = \frac{R^2_p/q}{(1-R^2_p)/(n-k-q)}. \]
अंश हर अपवर्जित-instrument restriction पर fit में सुधार मापता है; हर residual degree of freedom इस्तेमाल कर हर denominator residual variance का अनुमान देता है। यह सूत्र हर robust statistic की सार्वभौमिक पहचान नहीं है।
Robust covariance estimator test की गणना और reference distribution बदल देता है।
Partial \(R^2\) और F संबंधित, लेकिन अलग वर्णनात्मक प्रश्नों के उत्तर देते हैं: partial \(R^2\) scale-free fit measure है, जबकि F sample size और restrictions की संख्या भी दर्शाता है।
बहुत बड़े sample में छोटा partial \(R^2\) भी बड़ा F दे सकता है; इससे हर finite-sample IV approximation भरोसेमंद नहीं हो जाती।
F का 10 से ऊपर होना सार्वभौमिक threshold क्यों नहीं
प्रचलित 10 एक rule of thumb है, pass/fail theorem नहीं। Staiger और Stock इसे एक खास weak-instrument framework में व्यावहारिक मार्गदर्शन के रूप में चर्चा करते हैं; यह हर sample, estimator, instruments की संख्या या error process के लिए valid inference की गारंटी नहीं देता।
Stock और Yogo instrument weakness को relative 2SLS bias या Wald-test size distortion की सीमाओं से परिभाषित करते हैं और चुने हुए मानदंडों के critical values देते हैं। इसलिए critical value criterion और model dimensions पर निर्भर है।
पारंपरिक first-stage और Cragg–Donald परिणाम homoskedastic तथा independent errors मानते हैं।
उन tables का कोई आँकड़ा बिना बदलाव के हर robust setting में नहीं लगाया जा सकता (लेखक का अध्याय पृष्ठ)।
10 से बड़ा मान exclusion, independence या causal interpretation सिद्ध नहीं करता। 10 से कम होना भी नहीं सिद्ध करता कि उपकरण बेकार है या कोई खास estimate अमान्य है।
Statistic को बताई गई assumptions के तहत diagnostic समझें, फिर design के अनुकूल inference चुनें।
Hansen J मार्गदर्शिका बताती है कि overidentification test strength assessment का विकल्प क्यों नहीं है।
<!-- learn:illustration -->

कई अंतर्जात regressors के लिए संयुक्त शक्ति का आकलन ज़रूरी है
एक से अधिक endogenous regressor होने पर अलग-अलग first-stage F कमज़ोर पहचान वाले किसी linear combination को चूक सकते हैं।
हर regressor अलग से predicted दिख सकता है, जबकि instruments से आया बदलाव सभी structural coefficients को सटीकता से estimate करने के लिए आवश्यक combinations को span न करे।
Homoskedastic linear IV setup में Cragg–Donald minimum-eigenvalue statistic इस joint identification information का सार देता है। उसके Stock–Yogo critical values निर्दिष्ट bias या size-distortion criteria को लक्षित करते हैं।
Assumptions मायने रखते हैं: मनमाने heteroskedasticity, serial dependence या clustering के बाद परिचित critical values अपने-आप valid नहीं होते।
अपवर्जित instruments की संख्या, अंतर्जात regressors की संख्या और first-stage coefficient matrix का rank—सभी मायने रखते हैं। पूरा setup बताएँ और उसके लिए बनाया test लें; व्यक्तिगत first-stage F का औसत लेकर joint strength न निकालें।
Robust errors के लिए design-matched diagnostics चाहिए
Financial observations heteroskedastic, serially dependent या firm, venue अथवा policy unit के अनुसार clustered हो सकते हैं। Second-stage standard errors cluster करने भर से classical first-stage F या Cragg–Donald calibration robust नहीं हो जाते।
एक endogenous regressor के लिए Montiel Olea और Pflueger, बताई गई conditions में heteroskedasticity, autocorrelation और clustering की अनुमति देने वाला effective-F weakness test विकसित करते हैं।
यह covariance information से conventional first-stage F को scale करता है और फिर परिणाम की तुलना criterion-specific critical values से करता है (2013 का लेख)।
Effective F हर software द्वारा दिखाए गए robust Wald F के समान मात्रा नहीं है।
Kleibergen–Paap rk Wald F को robust covariance estimators के साथ अक्सर report किया जाता है, लेकिन Stock–Yogo critical values अलग conventional statistics और assumptions के लिए निकाले गए थे। इन्हें एक ही calibrated scale की तरह compare न करें।
Estimator, statistic, covariance estimator, clustering level या dependence treatment, और critical-value framework report करें।
Newey–West या cluster-robust covariance अपनी assumptions के अंतर्गत sampling uncertainty संभालता है; अकेले weak identification ठीक नहीं करता।
कमज़ोर relevance में Anderson–Rubin inference वैध रह सकता है
एक endogenous regressor वाले model में candidate coefficient \(\beta_0\) के लिए outcome को \(Y-X\beta_0\) से adjust करें और included controls के बाद जाँचें कि अपवर्जित instruments उसे संयुक्त रूप से समझाते हैं या नहीं।
Null और instrument exogeneity के तहत यह \(\beta=\beta_0\) का Anderson–Rubin test है।
Test estimated first-stage coefficient से भाग नहीं देता, इसलिए इसकी validity उस coefficient के बड़ा होने पर निर्भर नहीं होनी चाहिए।
मूल Anderson–Rubin construction model की distributional assumptions का उपयोग करता है; applied work में covariance estimator और reference distribution भी sampling design से मेल खाने चाहिए। Test को robust कह देने से कम clusters या serial dependence को अनदेखा नहीं किया जा सकता।
Candidate values पर test को invert करने से confidence set मिलता है। जानकारी कमज़ोर हो तो set चौड़ा, अलग-अलग हिस्सों में बँटा या असीमित हो सकता है; यह software failure नहीं, सीमित identifying information का निष्कर्ष है। Anderson–Rubin tests की power भी कम हो सकती है।
GMM overidentification मार्गदर्शिका अलग प्रश्न का उत्तर देती है; इसे weak-IV-robust coefficient inference न समझें।
मूल लेख Anderson और Rubin (1949) है।
आंशिक F की गणना का उदाहरण
एक काल्पनिक first stage में \(n=200\) observations, intercept सहित \(k=3\) included regressors, \(q=2\) excluded instruments और partial \(R^2_p=0.04\) मानें।
पारंपरिक homoskedastic formula में residual degrees of freedom \(200-3-2=195\) हैं।
\[ F = \frac{0.04/2}{(1-0.04)/195} = \frac{0.02}{0.96/195} = 4.0625. \]
यह गणना बताती है कि दिए गए inputs और assumptions के तहत conventional joint statistic 4.0625 है। यह अपने-आप invalidity सिद्ध नहीं करती, Stock–Yogo के criterion-specific निष्कर्ष तय नहीं करती, या heteroskedasticity अथवा clustering के तहत वही नतीजा साबित नहीं करती।
Robust analysis के लिए अपना design-matched statistic और calibration चाहिए।
यदि वही दो instruments fuzzy regression-discontinuity design में उपयोग हों, तो first-stage discontinuity उस design की relevance analysis का हिस्सा है।
ऊपर की F गणना RDD assumptions, bandwidth चुनाव या उस design के उपयुक्त inference का विकल्प नहीं है।
Diagnostic और identification तर्क अलग-अलग report करें
Endogenous regressors, excluded instruments, included controls, sample, first-stage coefficients, partial \(R^2\), और exact strength statistic बताएँ।
कई endogenous regressors के लिए joint statistic और assumptions बताएँ; robust settings में केवल “F = …” लिखने के बजाय covariance और critical-value method का नाम दें।
यदि statistic weak identification का संकेत दे तो target coefficient के लिए weak-IV-robust test या confidence set report करें। बताएँ कि set bounded है या नहीं और उसका आकार substantive conclusion को कैसे बदलता है।
सिर्फ सारांश देना आसान हो इसलिए बेकार interval को conventional Wald interval से न बदलें।
अंततः institutional और economic evidence से instrument की independence और exclusion paths का बचाव करें। Relevance diagnostics, balance checks, placebo outcomes और overidentification test समस्याएँ दिखा सकते हैं, पर हर assumption सिद्ध नहीं कर सकते।
Difference-in-differences design अलग identification assumptions उपयोग करता है; अधिक मजबूत first stage उसे इस design का विकल्प नहीं बनाता।
आम सवाल
Q1क्या first-stage F का 10 से अधिक होना instrument की validity सिद्ध करता है?
नहीं। यह किसी खास calibration के अंतर्गत relevance diagnostic भर है। यह independence या exclusion स्थापित नहीं कर सकता।
Q2क्या partial R-squared और first-stage F statistic एक ही हैं?
नहीं। Partial R-squared incremental fit बताता है। Conventional F sample size, restrictions की संख्या और residual degrees of freedom पर भी निर्भर है।
Q3क्या robust F की तुलना सीधे Stock–Yogo critical values से कर सकता हूँ?
केवल तब जब statistic और assumptions calibration से मेल खाते हों। Robust statistics को अक्सर अलग critical values और interpretation चाहिए।
Q4यदि instrument कमज़ोर दिखे तो क्या report करें?
Design-matched diagnostic और weak-IV-robust test या confidence set दें; साथ में बताएँ कि set चौड़ा, disconnected या unbounded है या नहीं।
स्रोत और आगे पढ़ें
समस्या की रिपोर्ट करें
हम इस लेख का लिंक जोड़कर ईमेल तैयार करेंगे। भेजने के बाद ही Mark को आपकी रिपोर्ट मिलेगी
त्वरित जाँच
गाइड पढ़ने के बाद 3 सवालों से खुद को जाँचें
सवाल 01
Conventional first-stage F test क्या जाँचता है?
व्याख्या देखने के लिए एक उत्तर चुनें
विकल्प शब्दावली
Correlation between a regressor and the regression disturbance, causing OLS to mix a target effect with omitted pathways, simultaneity, or measurement error.
विस्तृत गाइड पढ़ेंRegression discontinuity designA design using a treatment-probability jump at a known running-variable cutoff plus potential-outcome continuity to identify a local effect.
विस्तृत गाइड पढ़ेंDifference-in-differencesA quasi-experimental method subtracting the comparison group's contemporaneous change from the treated group's change to construct an untreated counterfactual.
विस्तृत गाइड पढ़ें