Lesson 3.B.4 - Chi-Square Test for Homogeneity

Key Question: Should underperforming large schools be broken into smaller schools?

Content: Expected Counts | Chi-Square Statistic | Chi-Square Test for Homogeneity

Alignment: CED Topics 3.14-3.15

Video

Student Items

Handout: pdf, doc

Mastery Check: link

Supplemental/Calculator Videos: link

Teacher Items

Handout Key: pdf, doc

Mastery Check Key: link

Slide Deck: pdf, ppt

Lesson Feedback: link

Course Resources

Resources for teaching our AP® Statistics curriculum.

  • Lesson Flow - timing and flow of class, using our lesson materials
  • Pacing Guide - pacing our units, with daily or block schedules
  • CED Alignment Guide - aligning our lessons to the AP® Statistics Course and Exam Description

Teaching Resources

Resources for teaching with Skew The Script.

Lesson Notes

Lesson-specific insights from the creators of this lesson.

GIF

Education policymakers and researchers often like to study “turnaround schools,” or schools that show high year-to-year growth on their state accountability ratings. When searching state databases for such schools, one clear pattern emerges: turnaround schools tend to be small schools. This finding seems to lend evidence to ideas from the small school movement: that small schools are less bureaucratic, faster to build consensus, and easier to reform. However, as students discover in this lesson, there may be more to this story than meets the eye.

Learning Targets
  • Determine expected counts for a two-way table
  • Calculate and interpret the chi-square statistic for a two-way table
  • Conduct a chi-square test for homogeneity

Before proceeding: Familiarize yourself with the lesson materials linked above (e.g. handout, handout key, slides, video). Then, for additional background and teaching tips from the lesson creators, check out the sections below.


  • As students proceed through the lesson, they discover a surprising fact: small schools aren’t just more likely to show big growth. They’re also more likely to show big declines. To avoid giving away this “punchline,” it can be helpful to avoid having students look closely at the percent of schools in each size category that show big growth, moderate change, and big declines until they reach the Discussion Question. At that point, students will experience the full surprise and be tasked with explaining why it occurs.
  • Although the chi-square test is new, the central idea is familiar: compare what actually happened (observed counts) with what we would expect to happen if the null hypothesis were true (expected counts). Returning to this observed-versus-expected framing throughout the lesson can help students see the chi-square test as an extension of earlier inference, rather than an entirely new concept.
  • Consider lingering on Question 4 in the Handout before introducing the formula for the chi-square statistic. Students can already identify where observed counts differ from expected counts. However, they may not be able to determine if a difference is substantial enough to be meaningful or significant. This motivates the purpose of the chi-square statistic before students encounter its calculation.

First, download this lesson's Handout Key and read through its Discussion Question section. Then, check out our model discussion norms and the additional background notes below.

  • Students may initially attribute the higher rate of improvement among small schools to characteristics of small schools themselves. Once they notice that small schools also have the highest rate of large declines, encourage them to reconsider that explanation and ask what features of smaller schools could produce more extreme outcomes in either direction.
  • If students need a bridge, connect the discussion to earlier work with sampling distributions: smaller samples tend to produce greater sample-to-sample variability. Because fewer students contribute to a small school’s overall rating, differences in student performance from one year to the next can produce larger swings in the school-level measure.
  • Advocates for small schools have long argued that smaller schools are easier to reform, as they provide easier settings for collaboration and consensus-building. For example, a 1989 New York Times op-ed by Deborah Meier, an early advocate of the modern small schools movement, argued that “Small schools offer opportunities to solve every one of these critical issues…staff can meet to discuss issues and differences without complex governance structures; understanding the budget does not require an advanced degree in accounting. Looking in on colleagues and, sharing ideas, becomes possible.”
  • The pattern surfaced by this lesson – that smaller schools have more variability in their performance due to their smaller sample size of students – is also discussed in several popular books, including Thinking, Fast and Slow and How Not to Be Wrong.
  • The context provides a useful opportunity to revisit the principle that correlation does not imply causation. The data show an association between school size and changes in school performance, but they do not establish a causal link. In fact, the higher rates of both big improvements and big declines among small schools suggest that sampling variability may be the primary reason for the pattern, rather than any other inherent traits of small schools.
  • Although chi-square looks very new, the underlying reasoning is familiar: describe what we would expect under the null hypothesis (expected counts), measure how far the observed data (observed counts) depart from that expectation, and determine whether such a discrepancy could reasonably occur by chance. Emphasizing this common structure can help students understand chi-square as another application of statistical inference, rather than a totally new procedure to memorize.
  • Expected counts connect directly to students’ earlier work with independent events in lesson 2.A.5. Under the null hypothesis, knowing a school’s size does not affect the probability of its performance falling into a certain category. In probability notation: P(Big improvement in ratings | school is small) = P(Big improvement in ratings). This mirrors the definition of independence: P(A | B) = P(A). Making this connection can help students see expected counts as a familiar concept, rather than as another calculation to memorize.
  • One advantage of the chi-square test is that it allows us to ask a single overall question about whether the distributions differ across the groups, rather than conducting many separate pairwise tests (e.g. conducting many two-sample z-tests for the difference between population proportions). A significant result tells us that there is evidence of a difference somewhere in the table, but it does not tell us exactly which groups or categories account for that difference. In more advanced statistical work, researchers may conduct follow-up comparisons while adjusting for the increased risk of Type I error that comes from conducting multiple tests.
  • For a more detailed look at how the p-value for a chi-square test is calculated from a chi-square distribution, check out this supplemental video.

Student Supports

Lesson-specific resources to support all learners.

  • It can be helpful to ask students, “What would it mean if the chi-square statistic were zero?” A value of zero shows that every observed count is exactly equal to its expected count. From there, students can reason that increasingly large chi-square values represent increasingly large discrepancies between expected and observed counts, meaning that the observed counts are unusual compared to what is expected under the null.
  • When calculating χ², students may benefit from thinking of each cell as making its own contribution to the overall statistic. A cell in which observed and expected counts are close contributes little; a cell with a large discrepancy contributes more. This framing can help students interpret the statistic as a summary of the discrepancies across the entire table, rather than as one long calculation.
  • Vocabulary used in the context of the lesson may include words that are unfamiliar or have several meanings. In particular, the following mathematical terms may need clarification or a definition provided:
    • Observed count
    • Expected count
    • Homogeneity
  • In addition, the following contextual terms may need clarification or a definition provided:
    • Turnaround school
    • School accountability rating
    • School reform
  • The Greek letter chi (χ) is unfamiliar to many students. It can be helpful to provide the following guidance on pronouncing it: “The chi-square test is pronounced ‘k-eye’ square. Not ‘ch-eye’ square, as in a chai latte. And not ‘chee’ square, as in Tai Chi. But ‘k-eye,’ as in ‘k’ followed by ‘eye’ like ‘eyeball.’”