विकल्प और फ्यूचर्स की सभी गाइड
Overlapping market variables को कुछ axes में compress करें18 मिनट पढ़ें

Principal Component Analysis (PCA) Explained

Centering, scaling, covariance, component scores और loadings, explained variance, component selection और finance में PCA की सीमाएं समझें

Mark द्वारा तैयार · नीचे प्राथमिक स्रोत

सीधा जवाब

Principal component analysis (PCA) correlated variables को छोटे set of orthogonal axes में transform करता है जो sample variance को जितना संभव हो preserve करते हैं। First component सबसे अधिक variance capture करता है और हर अगला component पहले axes के orthogonal रहते हुए बची हुई variance में सबसे अधिक capture करता है। PCA risk summarize और predictors बनाने में मदद कर सकता है, लेकिन economic factors अपने आप discover नहीं करता। Results units, sample dates, covariance estimation और component selection पर निर्भर करते हैं, इसलिए scores, loadings, explained variance और out-of-sample stability को एक ही analysis में report करना चाहिए।

Data को deliberate baseline पर रखें

PCA centered data पर apply होता है, यानी हर variable का sample mean घटाने के बाद। वरना average levels variation की intended direction पर हावी हो सकते हैं।

Calculation से पहले meaningful transformations चुनें। Nonstationary price levels में common trend useful shared risk signal जैसा दिख सकता है।

Missing values, outliers, frequency और sample window covariance बदलते हैं। PCA से पहले data design बाद में दिए गए नाम से अधिक महत्वपूर्ण है।

Covariance और correlation अलग questions के उत्तर देते हैं

Covariance-matrix PCA original units preserve करता है, इसलिए अधिक volatility वाला variable leading directions पर अधिक प्रभाव डाल सकता है।

Correlation-matrix PCA पहले हर variable को unit variance तक scale करता है। यह unit effects सीमित करता है, लेकिन magnitude के economically meaningful differences हटा सकता है।

Objective choice तय करता है। Similar-unit yield changes और mixed indicators के collection को जरूरी नहीं कि same matrix चाहिए।

Components क्रम से variance maximize करते हैं

First principal component वह unit direction है जिसके projected observations में possible sample variance सबसे अधिक है।

हर अगला component पहले चुने गए सभी components के orthogonal रहते हुए remaining projected variance maximize करता है।

Orthogonality fitted sample में component scores को uncorrelated बनाती है। यह साबित नहीं करती कि underlying economic shocks independent हैं।

Loadings और scores अलग काम करते हैं

Loadings वे weights हैं जो original variables को component में combine करते हैं। वे बताते हैं कि कौन-से variables किसी axis की दिशा में या उसके विरुद्ध हैं।

Scores हर observation को उस axis पर project करने से मिले coordinates हैं। Loadings direction define करते हैं; scores उसके along time में move करते हैं।

Eigenvector को minus one से multiply करने पर axis नहीं बदलती। अकेला sign flip economic relation reverse होने का evidence नहीं है।

Explained variance अधूरा objective है

Component का eigenvalue उसकी sample variance के बराबर है। उसे सभी eigenvalues के योग से divide करने पर explained-variance ratio मिलता है।

Scree plot और cumulative explained variance dimension चुनने की शुरुआत हैं। 80% जैसे fixed threshold हर जगह optimal नहीं होते।

Low-variance direction में predictive information या binding hedge relation फिर भी हो सकती है। Explained variance return predictability के बराबर नहीं है।

Dimension को fitted sample के बाहर चुनें

बहुत कम components structure छोड़ देते हैं, जबकि बहुत अधिक components persistent common movement के साथ sample noise भी preserve करते हैं।

Intended use के अनुसार reconstruction error, task-specific cross-validation, eigenvalue gaps और bootstrap stability compare करें।

Rolling samples में loadings और explained variance फिर estimate करें। Historical forecast में full-sample axes इस्तेमाल करने से future information आती है।

Finance common movement summarize करने के लिए PCA इस्तेमाल करता है

Yield-curve changes level, slope और curvature-like axes में compress हो सकते हैं, जबकि equity returns broad और group-specific directions दिखा सकते हैं।

Portfolios में leading axes बता सकती हैं कि nominal holdings की बड़ी संख्या के बावजूद risk concentrated है या नहीं।

Statistical axis को “growth” या “liquidity” कहने के लिए external variables, economic reasoning और दूसरे sample में validation चाहिए।

Result को reproducible बनाने वाले choices report करें

Variables, transformations, currencies, units, centering, scaling, missing-value treatment, outliers और estimation window disclose करें।

Matrix, eigenvalues, explained variance, loadings, scoring rule और dimension criterion बताएं ताकि दूसरा analyst result rebuild कर सके।

Out-of-sample performance और rolling stability दिखाएं। PCA चुने हुए data का estimated coordinate system है, reality का fixed map नहीं।

आम सवाल

क्या PCA और factor analysis एक ही हैं?

नहीं। PCA total variance को re-express करता है, जबकि common-factor model common covariance को variable-specific variance से अलग करता है।

क्या हर variable को हमेशा standardize करना चाहिए?

नहीं। Standardization unit differences को control करता है, लेकिन volatility magnitude के potentially meaningful differences भी हटा देता है।

क्या loading sign flip का अर्थ structure बदल गया है?

जरूरी नहीं। Eigenvector और उसका negative same axis define करते हैं, इसलिए windows compare करने से पहले signs align करने चाहिए।

क्या high-explained-variance component अच्छा trading signal है?

नहीं। Explained variance contemporaneous movement मापता है; predictive power के लिए अलग out-of-sample test चाहिए।

स्रोत और आगे पढ़ें

संबंधित गाइड