Bank Statistics में Chi-Square Test: CAIIB ABM गाइड
CAIIB ABM की सांख्यिकी में हर सवाल औसत या correlation coefficient निकालने के लिए नहीं होता। एक बार-बार आने वाली श्रेणी बिल्कुल अलग सवाल पूछती है: जो पैटर्न आप देख रहे हैं, क्या वह वास्तविक है, या यह सिर्फ संयोग हो सकता है? यही सवाल chi-square test in bank statistics हल करता है। यह वह मानक टूल है जिसका इस्तेमाल बैंक यह जांचने के लिए करते हैं कि क्या दो categorical factors — जैसे loan channel और default status — वाकई एक-दूसरे से जुड़े हैं, या क्या complaints, returns या defaults का देखा गया distribution उससे मेल खाता है जिसकी अपेक्षा थी। यह लेख टेस्ट के दोनों प्रमुख रूपों, degrees-of-freedom के नियम, और उन गलतियों से गुजरता है जो CAIIB candidates के अंक काटती हैं।
📊 Chi-Square Test in Bank Statistics क्या है?
Chi-square (χ²) टेस्ट एक non-parametric सांख्यिकीय टेस्ट है जो continuous measurements के बजाय categorical (count) डेटा पर लागू होता है। यह "loan का औसत आकार क्या है" पूछने के बजाय यह पूछता है कि "क्या categories में गिनतियों का पैटर्न उससे मेल खाता है जिसकी अपेक्षा तब होगी जब कोई वास्तविक संबंध न हो, या जब कोई ज्ञात distribution लागू हो।" चूंकि बैंक डेटा categorical variables से भरा होता है — product type, region, channel, complaint category, default status — chi-square MIS review और audit sampling में लगातार सामने आता है, एक बार जब raw आंकड़े उस तरह व्यवस्थित हो जाएं जैसा definition of statistics वाले chapter में बताया गया है।
यह टेस्ट observed frequencies (जो वास्तव में हुआ, आपके डेटा से) की तुलना expected frequencies (जिसकी थ्योरी या null hypothesis भविष्यवाणी करती है) से करता है। observed और expected counts के बीच जितना बड़ा अंतर होगा, chi-square statistic उतना ही बड़ा होगा, और यह संभावना उतनी ही कम होगी कि यह अंतर केवल संयोग से है। इन counts के इस्तेमाल के लायक बनने से पहले, इन्हें सार्थक categories में बांटना जरूरी है — वही कौशल जो classification and tabulation of banking data chapter में बनाया गया है, जो असल में वह बुनियाद है जिस पर हर chi-square problem टिकी होती है।
🧮 टेस्ट के दो रूप: Independence और Homogeneity
CAIIB ABM दो नज़दीकी रूप से जुड़े applications की परीक्षा लेता है, और exam आपसे इनमें फर्क बताने की उम्मीद रखता है। test of independence यह जांचता है कि क्या एक ही group पर मापे गए दो categorical variables आपस में जुड़े हैं — उदाहरण के लिए, क्या default status एक ही portfolio के borrowers में loan disbursement channel (branch बनाम digital) से independent है। test of homogeneity यह जांचता है कि क्या किसी variable का distribution अलग-अलग groups में एक जैसा है — उदाहरण के लिए, क्या complaint categories चार regional zones में एक जैसी distributed हैं। दोनों मामलों में गणित एक जैसा ही है; सिर्फ hypothesis के शब्द और डेटा का layout अलग होता है।
एक तीसरा, आसान रूप है goodness-of-fit test, जो एक observed distribution की तुलना एक single theoretical या historical distribution से करता है — उदाहरण के लिए, यह जांचना कि इस quarter के cheque-return reasons पिछले साल के proportions से अभी भी मेल खाते हैं या नहीं। ये तीनों रूप एक contingency या frequency table का इस्तेमाल करते हैं, row और column totals से (या theoretical proportions से) expected counts निकालते हैं, और observed व expected के बीच squared, standardised अंतरों को जोड़ते हैं।
💡 Exam Tip: अगर सवाल आपको ONE variable और जांचने के लिए theoretical proportions का सेट देता है, तो यह goodness-of-fit है। अगर यह आपको एक grid में cross-tabulated TWO variables देता है, तो यह test of independence है (या homogeneity, यह sampling किस तरह की गई थी उस पर निर्भर करता है)।

📐 Contingency Table पढ़ना और सही टेस्ट चुनना
नीचे दी गई table यह बताती है कि एक CAIIB candidate को यह कैसे तय करना चाहिए कि कौन-सा chi-square variant लागू होता है, उस तरह के scenario का इस्तेमाल करके जो bank statistics के सवालों में आता है।
| Test Type | यह क्या जांचता है | विशिष्ट Bank उदाहरण | Minimum Expected-Frequency नियम पूरा? |
|---|---|---|---|
| Goodness-of-fit | क्या एक observed distribution किसी expected/theoretical distribution से मेल खाता है | क्या इस महीने के cheque-return reasons पिछले पांच साल के औसत split से मेल खाते हैं | ✅ आमतौर पर, अगर categories सही तरीके से pool की गई हों |
| Test of independence | क्या एक ही sample पर मापे गए दो categorical variables जुड़े हैं | क्या loan default status disbursement channel (branch बनाम digital) से independent है | ✅ अगर हर cell का expected count ≥ 5 हो |
| Test of homogeneity | क्या दो या अधिक अलग populations किसी variable का एक जैसा distribution रखती हैं | क्या अलग से sample किए गए regional zones में complaint categories अलग हैं | अगर किसी cell < 5 हो तो pooling या Yates' correction चाहिए |
Contingency table में किसी भी cell के लिए expected frequency की गणना (row total × column total) ÷ grand total के रूप में की जाती है। यही एक फॉर्मूला लगभग हर CAIIB chi-square numerical में सबसे ज्यादा काम आता है, इसलिए exam के दबाव में इसे derive करने के बजाय इसे रट लेना बेहतर है।

🎯 Degrees of Freedom और Critical Value का फैसला
r rows और c columns वाली contingency table के लिए degrees of freedom (df) होता है (r − 1) × (c − 1); k categories वाले simple goodness-of-fit test के लिए df सीधे k − 1 होता है। इसके बाद calculated chi-square statistic की तुलना chosen significance level (आमतौर पर 5%) और संबंधित df पर chi-square distribution table से मिले critical value से की जाती है — यही critical-value logic syllabus में कहीं और estimation और confidence-interval से जुड़े सवालों में भी इस्तेमाल होती है।
अगर calculated χ² critical value से ज्यादा है (या समान रूप से, अगर p-value significance level से कम है), तो आप independence या good fit के null hypothesis को reject कर देते हैं — पैटर्न सांख्यिकीय रूप से significant है और संयोग होने की संभावना कम है। अगर calculated value critical value से कम है, तो आप null hypothesis को reject नहीं करते; इस बात का पर्याप्त सबूत नहीं है कि दोनों variables जुड़े हैं, या distribution वाकई बदला है।
Null hypothesis को reject न करना उसे सही साबित करने के बराबर नहीं है — इसका मतलब सिर्फ यह है कि sample इतना सबूत नहीं दे पाया कि किसी वास्तविक association का निष्कर्ष निकाला जा सके।

⚠️ CAIIB Candidates कहां गलती करते हैं
सबसे आम numerical गलती यह भूल जाना है कि chi-square के लिए frequencies (counts) चाहिए, percentages या averages कभी नहीं — counts की जगह proportions डाल देने से एक बेमतलब statistic मिलता है। दूसरी आम गलती तब होती है जब expected cell frequencies बहुत छोटी हों (आमतौर पर 5 से कम) फिर भी टेस्ट लगा दिया जाए; standard उपाय है adjacent categories को pool करना या 2×2 table के लिए Yates' continuity correction लगाना, और examiners यह जांचना पसंद करते हैं कि क्या आपको इस exception के होने का पता भी है।
तीसरी गलती chi-square को correlation के साथ confuse करना है। Chi-square यह बताता है कि क्या दो categorical variables जुड़े हैं; यह दो continuous variables के बीच संबंध की ताकत या दिशा नहीं बताता, जो regression बताता है। अगर कोई सवाल आपको continuous डेटा देता है — जैसे loan amount बनाम tenure — तो आप correlation and regression के दायरे में हैं, chi-square के नहीं।
⚠️ Common Mistake: degrees of freedom और फैसला (reject / fail to reject) बताए बिना सिर्फ chi-square value लिख देने पर ज्यादा से ज्यादा आंशिक अंक मिलते हैं — निष्कर्ष ही वह चीज़ है जिसे examiner असल में जांच रहा होता है।
🏦 बैंक असल में इस टेस्ट का इस्तेमाल कहां करते हैं
exam हॉल से आगे, internal audit और MIS teams नियमित रूप से chi-square checks चलाती हैं। कोई fraud-monitoring desk यह टेस्ट कर सकता है कि क्या flagged transactions उस समय से independent हैं जब वे process हुए, यह देखने के लिए कि पैटर्न वास्तविक है या संयोग। सर्विस performance को benchmark करने वाली एक quality-review team — जो balanced scorecard for bank performance जैसे measurement में बनी होती है — goodness-of-fit test का इस्तेमाल यह जांचने के लिए कर सकती है कि क्या इस quarter का customer-satisfaction category split अभी भी baseline से मेल खाता है, ताकि noise के बजाय वास्तविक बदलाव पकड़ा जा सके। सशक्त corporate governance in banks के तहत बोर्ड को रिपोर्ट करने वाली audit और risk committees अब किसी finding को escalate करने से पहले ऐसे सांख्यिकीय समर्थन की उम्मीद करती हैं, न कि सिर्फ descriptive complaint count की। इससे domain judgement की जगह नहीं ली जा सकती — छोटे, कमजोर sample वाले dataset में सांख्यिकीय रूप से significant result को भी sanity check चाहिए, वही सावधानी जो किसी भी RBI-published banking statistics release को पढ़ते समय निष्कर्ष निकालने से पहले बरतनी चाहिए। यह याद रखना भी जरूरी है कि हर banking गणना probabilistic नहीं होती — CAIIB BFM में forward rate agreement जैसा एक deterministic instrument फॉर्मूले से price होता है, जिसमें कोई hypothesis test शामिल नहीं होता — इसलिए यह जानना कि कब सवाल chi-square मांगता है और कब एक fixed फॉर्मूला, यह भी उस कौशल का हिस्सा है जिसे परखा जा रहा है।
🧠 Practice MCQs: Chi-Square Test in Bank Statistics
Q1. एक बैंक यह जांचना चाहता है कि क्या एक ही portfolio के borrowers में loan default status disbursement channel (branch बनाम digital) से जुड़ा है। कौन-सा टेस्ट लागू होता है? (a) Goodness-of-fit test (b) Test of independence (c) t-test (d) F-test
Answer: (b) — एक ही sample पर मापे गए दो categorical variables का association जांचना chi-square test of independence है।
Q2. 4 rows और 3 columns वाली chi-square contingency table में degrees of freedom है: (a) 12 (b) 7 (c) 6 (d) 3
Answer: (c) — Contingency table के लिए degrees of freedom (r − 1) × (c − 1) = (4 − 1) × (3 − 1) = 6 होता है।
Q3. Contingency table में किसी cell की expected frequency की गणना इस तरह की जाती है: (a) Row total minus column total (b) (Row total × Column total) ÷ Grand total (c) Grand total ÷ number of cells (d) Row total ÷ Column total
Answer: (b) — किसी भी cell की expected frequency उसके row total को column total से गुणा करके, grand total से भाग देने पर मिलती है।
Q4. अगर 2×2 contingency table में कई cells की expected frequencies 5 से कम हैं, तो सही तरीका है: (a) नियम को नज़रअंदाज़ करके आगे बढ़ना (b) Yates' continuity correction लगाना या categories को pool करना (c) correlation test पर switch करना (d) sample size को अपने आप दोगुना कर देना
Answer: (b) — जब expected cell frequencies बहुत छोटी हों, तो Yates' correction (2×2 tables के लिए) या adjacent categories की pooling chi-square approximation को valid बनाए रखती है।
Q5. अगर calculated chi-square value chosen significance level पर critical value से कम है, तो सही निष्कर्ष है: (a) Null hypothesis को reject करना (b) Null hypothesis को reject न करना (c) टेस्ट invalid है (d) अलग df के साथ फिर से calculate करना
Answer: (b) — जब calculated statistic critical value से कम रह जाता है, तो independence या good fit के null hypothesis को reject करने के लिए पर्याप्त सबूत नहीं होता।
100+ MCQs के साथ chapter-wise mock tests चाहिए? मुफ़्त practice शुरू करें →
❓ अक्सर पूछे जाने वाले सवाल
Banking statistics में chi-square test किसलिए इस्तेमाल होता है?
यह जांचता है कि categorical डेटा — categories में बंटी loans, complaints, defaults या returns की गिनतियां — कोई वास्तविक association या expected pattern से वास्तविक बदलाव दिखाती हैं, न कि ऐसा अंतर जो आसानी से संयोग की वजह से हो सकता है।
Goodness-of-fit और test of independence में क्या फर्क है?
Goodness-of-fit एक observed distribution की तुलना एक single expected या theoretical distribution से करता है। Test of independence यह जांचता है कि एक ही sample पर cross-tabulated दो categorical variables एक-दूसरे से जुड़े हैं या नहीं।
अगर किसी cell की expected frequency 5 से कम हो तो क्या होता है?
Standard chi-square approximation अविश्वसनीय हो जाता है। 2×2 table में सामान्य उपाय Yates' continuity correction है; बड़ी table में, adjacent categories को तब तक pool किया जाता है जब तक expected frequencies पर्याप्त न हो जाएं।
क्या chi-square test parametric है या non-parametric?
यह non-parametric है — यह categorical डेटा की frequency counts पर काम करता है और t-test या z-test जैसे टेस्टों के विपरीत यह नहीं मानता कि underlying population normal distribution का पालन करती है।
Chi-square CAIIB ABM के उन topics में से एक है जो कागज़ पर डरावना दिखता है लेकिन असल में एक फॉर्मूला, एक degrees-of-freedom नियम, और अंत में एक साफ फैसले पर आकर सिमट जाता है। ABM blog archive से कुछ contingency tables पर observed-बनाम-expected logic तब तक practice करें जब तक यह पैटर्न अपने आप न आने लगे, फिर CAIIB course के question bank पर खुद को timed सवालों से परखें।
Quick quiz on this topic
5 exam-style questions from our free test bank — check yourself before you move on.
Practice this topic
मुफ़्त मॉक टेस्ट दें, चैप्टर PDF डाउनलोड करें या वीडियो क्लास देखें — सब iibf.store पर मुफ़्त है।
पढ़ना जारी रखें