Survival Data में Censoring और Truncation
Right, left और interval censoring, truncation observed sample को कैसे बदलता है, independent censoring क्यों जरूरी है और IPCW bias कब घटा सकता है समझें
सीधा जवाब
Censoring का अर्थ है unit observe हुई लेकिन exact event time केवल partly known है: last follow-up के बाद, first observation से पहले या किसी interval के भीतर। Truncation अलग है—unit dataset में तभी enter करती है जब उसका event time sampling rule satisfy करे, इसलिए कुछ units कभी observe नहीं होतीं। Standard survival estimators को future event time से censoring independent चाहिए, संभव हो तो measured history पर condition करने के बाद। Informative dropout में inverse-probability-of-censoring weighting तभी मदद कर सकती है जब censoring model, covariates और positivity credible हों
Incomplete time भी information है
यदि performing loan data cutoff तक पहुँचता है, तो default time unknown है लेकिन observed duration से बड़ा होना known है। उसे discard करने से valid exposure information खोती है
Survival methods observed duration और event के बारे में known information दोनों encode करते हैं। Censored record missing row या forever confirmed nonevent नहीं है
हर censoring reason अलग record करें। Administrative study end, customer disappearance, data failure और contract exit अलग assumptions imply करते हैं
Censoring की कई directions हैं
Right censoring कहती है कि event last known time के बाद होता है। Follow-up default, churn या closure से पहले end होने पर यह common है
Left censoring कहती है event detection time से पहले हो चुका था। Interval censoring event को दो inspections के बीच locate करती है, exact timestamp के बिना
Interval को midpoint से replace करना invented precision है और distribution bias कर सकता है। Turnbull-type estimators interval information directly retain करते हैं
Truncation तय करता है कौन sample में दिखाई देता है
Left truncation में unit तभी observe होती है जब entry तक survive करे। Earlier failures incomplete time के साथ present नहीं, absent होती हैं
Coverage शुरू होने पर alive funds को ही include करने वाला fund database survivorship selection बनाता है। Ordinary risk sets durable funds को overrepresent करती हैं
Delayed-entry methods unit को entry पर ही risk set में add करती हैं। Validity के लिए फिर भी entry और future event time covariates पर condition करने के बाद suitably independent होने चाहिए
Independent censoring hidden bridge है
Kaplan–Meier observed past information से related censoring tolerate कर सकता है, यदि analysis structure पर condition करने के बाद censoring future event time के बारे में extra information न बताए
एक fixed date पर administrative censoring उस dropout से अधिक plausible है जो deteriorating credit quality या आने वाले platform exit से caused हो
Similar censoring percentages assumption prove नहीं करतीं। Mechanism समझने के लिए reasons, timing, covariate histories और loss से पहले event rates compare करें
Naive remedies नया bias बना सकते हैं
Complete-case analysis केवल observed events रखकर short durations oversample करती है। Censoring को event मानना outcome definition बदल देता है
हर censored unit को full horizon तक event-free मानना survival overstate करता है। Last observation carried forward भी यही structural problem रखता है
Refinancing जैसे competing events automatically censoring नहीं हैं। यदि वे default रोकते हैं, तो competing-risk estimand या clearly defined hypothetical analysis में शामिल करें
IPCW observed units को reweight करता है
Inverse probability of censoring weighting हर unit के observable बने रहने की chance estimate करके disappeared histories जैसी observed histories को upweight करता है
इसके लिए conditional independent censoring explain करने वाले measured predictors और relevant histories में positive observation probability चाहिए
Near-zero probabilities extreme weights और unstable estimates बनाती हैं। Weight distributions, effective sample size, truncation rules और censoring risk के across balance दिखाएँ
Financial datasets में कई observation filters होते हैं
Loans origination के बाद enter और sale, refinance, charge-off या vendor coverage change के जरिए exit कर सकते हैं; हर event का अर्थ अलग है
Options databases start date से पहले expired listings omit कर सकती हैं, जबकि hedge-fund databases केवल successful histories backfill कर सकती हैं। ये selection और truncation issues हैं
Source population से analysis risk sets तक flow diagram बनाएं। बड़ा dataset selective observation process को correct नहीं करता
Sensitivity missing mechanism को vary करे
Censoring reason, calendar window, follow-up cap, delayed-entry rule, weight model और weight truncation threshold बदलकर estimates repeat करें
Censored और retained units की observed characteristics तथा prior outcomes compare करें। Negative controls या linked data hidden departures दिखा सकते हैं
Report करें कौन-सी assumptions target identify करती हैं, कौन-से diagnostics testable हैं और plausible informative-censoring scenarios में conclusions कैसे बदलते हैं
आम सवाल
क्या censored unit उस unit जैसी है जिसे event नहीं हुआ?
नहीं। वह केवल last observed time तक event-free है; उसके बाद क्या हुआ unknown है
Left truncation और left censoring में क्या अंतर है?
Left-truncated units entry तक survive करें तभी दिखाई देती हैं; left-censored units observe होती हैं लेकिन उनका event exact known time से पहले हुआ था
क्या administrative censoring bias नहीं करती?
अक्सर अधिक defensible है, लेकिन calendar entry, changing cohorts या event-related study exit इसे informative बना सकते हैं
क्या IPCW unmeasured dropout reasons ठीक कर सकती है?
अपने आप नहीं। यह observed history से censoring explain होने और nonzero follow-up probability पर निर्भर करती है