विकल्प और फ्यूचर्स की सभी गाइड
कई ideas test करने पर एक threshold क्यों fail होता है16 min read

Finance में Multiple Testing और False Discovery Rate

Multiple comparisons, family-wise error, false discovery rate, Bonferroni, Holm, Benjamini–Hochberg, dependence, q-values, factor mining, backtest selection और research governance समझें

Mark द्वारा तैयार · नीचे प्राथमिक स्रोत

सीधा जवाब

Multiple testing तब होता है जब कई hypotheses, strategies, factors, assets, horizons या specifications को साथ search किया जाता है। तब individually valid p-values भी chance winners बना सकते हैं। Family-wise error rate कम-से-कम एक false rejection की probability control करता है, जबकि false discovery rate stated assumptions के अंतर्गत rejections में expected false share control करता है। समस्या केवल final table नहीं, पूरी research family है

अधिक tests अधिक chance winners बनाते हैं

यदि m independent true nulls को level α पर test किया जाए, तो कम-से-कम एक false rejection की probability 1−(1−α)^m है

α=0.05 और m=20 पर यह probability लगभग 64% है, भले हर individual test का nominal false-rejection rate 5% हो

Dependence exact calculation बदलती है, लेकिन यह दिखावा करने का कारण नहीं कि केवल reported winner test हुआ था

Family error का सवाल तय करती है

एक research claim के लिए विचार किए गए सभी factors, transformations, universes, lags, horizons, filters, model choices और reruns test family में आ सकते हैं

Results देखने के बाद family define करने से multiplicity कम दिखती है और related trials को अलग करके correction को harmless बनाया जा सकता है

Defensible family decision process और selection opportunity follow करती है, केवल paper या dashboard में बची rows को नहीं

FWER और FDR अलग losses control करते हैं

Family-wise error rate family के भीतर कम-से-कम एक false rejection करने की probability है

False discovery rate E[V/max(R,1)] है, जहाँ V false rejections की संख्या और R total rejections की संख्या है

जब कोई false claim costly हो तो FWER stricter है; अनेक candidates screen करते समय controlled false share स्वीकार हो तो FDR अधिक powerful हो सकता है

Procedures की अपनी assumptions होती हैं

Bonferroni हर hypothesis को α/m पर test करता है और independence के बिना FWER control करता है, हालांकि conservative हो सकता है

Holm का step-down method भी FWER control करता है और basic Bonferroni rejection rule से कभी कम powerful नहीं होता

Benjamini–Hochberg p-values को order करके independence और certain positive dependence में FDR control करता है; दूसरी dependence structures के लिए justified methods चाहिए

Adjusted values posterior truth probabilities नहीं हैं

Adjusted p-value चुनी procedure के अंतर्गत वह smallest family error level है जिस पर hypothesis reject होती

Q-value को आम तौर पर result को discovery कहने के लिए estimated minimum FDR threshold की तरह interpret किया जाता है, method और assumptions के अधीन

इनमें से कोई भी अपने आप selected strategy के false होने की probability नहीं है; उसके लिए अलग model और conditioning statement चाहिए

Finance एक बड़ी adaptive family छिपाता है

Factor mining, technical rules, option signals, alternative data, portfolio constraints और cost assumptions हजारों correlated trials बना सकते हैं

Published factors past searches से selected होते हैं, इसलिए isolated t-statistic near two से नए factor का मूल्यांकन accumulated discovery pressure ignore करता है

Shared assets और periods tests को dependent बनाते हैं, जबकि regime change selection correction को future economic performance की guarantee बनने से रोकता है

Governance को search record बचाना चाहिए

हर hypothesis, code version, universe, outcome, transformation, delay, cost और stopping choice को stable research identifier के साथ record करें

Discovery, confirmation और live monitoring अलग रखें; confirmatory rules freeze करें और genuinely untouched data या prospective evaluation उपयोग करें

Raw और adjusted evidence, family definition, dependence assumptions, effect sizes, costs, stability, rejected ideas और live degradation report करें

आम सवाल

Multiple testing problem क्या है?

कई hypotheses test और select करते समय winners को अकेले test जैसा judge करने से false discoveries बढ़ने की समस्या

FWER और FDR में क्या अंतर है?

FWER कम-से-कम एक false rejection की probability control करता है; FDR सभी rejections में expected false proportion control करता है

Benjamini–Hochberg कैसे काम करता है?

यह p-values को order करता है, rank-dependent thresholds से compare करता है और stated assumptions के अंतर्गत largest qualifying rank तक reject करता है

क्या FDR control trading strategy को reliable बनाता है?

नहीं। यह defined test family में selection error address करता है, regime change, costs, leakage, model risk या future profitability नहीं

स्रोत और आगे पढ़ें

संबंधित गाइड