196. Chi-Square Goodness of Fit
Compares an observed categorical distribution to a theoretical categorical distribution.
- : observed frequency in category
- : expected frequency in category
- : number of categories
Example
Your professor gave a hard test. He said 45% got A’s, 20% B’s, 15% C’s, 10% D’s and 10% F’s. You don’t believe him. You take a simple random sample of 60 students and find 20 got A’s, 15 B’s, 8 O’s, 12 D’s and 5 F’s. At , test.
| Category | A | B | C | D | F | Total |
| Expected % () | 45% | 20% | 15% | 10% | 10% | 100% |
| Observed # () | 20 | 15 | 8 | 12 | 5 | 60 |
| Expected # | 60 | |||||
- Hypothesis
- Test Statistic & P-Value
- Decision
Since (equivalently ), we fail to reject .
- Conclusion
At the significance level there is not enough evidence to conclude that the professor’s claimed grade distribution differs from the true one.
Example
Complaints by Quarter (testing for a uniform distribution)
| Q1 | Q2 | Q3 | Q4 | ||
| 43 | 52 | 54 | 40 | 189 |
Under all four quarters are equally likely, so for every .
| Q1 | Q2 | Q3 | Q4 | ||
| 43 | 52 | 54 | 40 | 189 | |
| 189 |
| Q1 | Q2 | Q3 | Q4 | ||
| 43 | 52 | 54 | 40 | 189 | |
| 189 |
, and . Since we fail to reject (): the complaints are consistent with a uniform distribution across quarters.
chi2_got.py
from scipy import stats
f_obs = np.array([43, 52, 54, 40])
f_exp = np.array([47, 47, 47, 47])
stats.chisquare(f_obs=f_obs, f_exp=f_exp)