I’m working on expected frequencies for a chi-square test of independence. Here are my observed counts (rows: Female, Male; columns: A, B, C):
– Female: A=28, B=30, C=12 (row total 70)
– Male: A=14, B=24, C=12 (row total 50)
– Column totals: A=42, B=54, C=24 (grand total 120)
My attempt is E(cell) = (row total × column total) / grand total. That gives:
– E(Female,A) = 70×42/120 = 24.5
– E(Female,B) = 70×54/120 = 31.5
– E(Female,C) = 70×24/120 = 14
– E(Male,A) = 50×42/120 = 17.5
– E(Male,B) = 50×54/120 = 22.5
– E(Male,C) = 50×24/120 = 10
Is this the right way to compute expected frequencies here (using the sample’s row and column totals)? Also, should I keep the decimals when running the test, or round to whole numbers? If I round, the totals no longer match exactly, which makes me uncertain.
















3 Responses
Yes-your formula and numbers are correct. For a chi-square test of independence (or homogeneity), the expected count in each cell is (row total × column total) / grand total, using the sample’s marginal totals. Your six expected values match that rule, and a quick check shows the expected row and column totals return to 70, 50 and 42, 54, 24 respectively, as they should. Do not round the expected counts; the test uses the fractional values, and rounding can change the chi-square statistic and break the margins. Simple example: for Female–A, E = 70×42/120 = 24.5, and its contribution to the test statistic is (28 − 24.5)² / 24.5 = 12.25 / 24.5 ≈ 0.50. Degrees of freedom here are (2 − 1)(3 − 1) = 2, and all your expected counts are at least 10, so the chi-square approximation is fine.
You’re on the right track. For a chi-square test of independence, the expected count in each cell is indeed (row total × column total) / grand total, using the sample’s row and column totals (the marginal proportions estimated from your data). The numbers you computed match that formula, and they add back up to the correct row and column totals, which is a good quick check.
For running the test, I’d keep the decimal expected counts as-is. The chi-square formula uses those expected values directly, and rounding can nudge the test statistic a bit and mess up the margins. If you do round for a table in a report, I’d only round for display and still compute with the unrounded values (some folks try to “fix” the last cell so totals match after rounding, but that quietly changes the test, so I tend to avoid it). Your expected counts are all above the common “at least 5 per cell” rule of thumb-though I’ve also seen 10 mentioned, so I always blink twice at values right at 10; part of me wonders about a correction like Yates’s, but that’s really for 2×2 tables, so I don’t think it applies here. If you want a friendly walkthrough, Khan Academy has a nice explanation of expected counts and the chi-square test of independence: https://www.khanacademy.org/math/statistics-probability/inference-categorical-data/chi-square-tests/v/chi-square-test-association-introduction
Nice! Using E = (row total × column total)/grand total is the classic independence trick, so your 24.5, 31.5, 14, 17.5, 22.5, 10 look right; I’d keep the decimals for the test, though I think rounding to integers usually won’t change much even if the margins wobble a bit.