Lesson 3.B.3 - Hypothesis Test for Two Proportions

Key Question: What causes racial wage gaps?

Content: Two-Sample z-Test for the Difference between Population Proportions

Alignment: CED Topics 3.12-3.13

Video

Student Items

Handout: pdf, doc

Mastery Check: link

Calculator Videos: link

Teacher Items

Handout Key: pdf, doc

Mastery Check Key: link

Slide Deck: pdf, ppt

Course Resources

Resources for teaching our AP® Statistics curriculum.

  • Lesson Flow - timing and flow of class, using our lesson materials
  • Pacing Guide - pacing our units, with daily or block schedules
  • CED Alignment Guide - aligning our lessons to the AP® Statistics Course and Exam Description

Teaching Resources

Resources for teaching with Skew The Script.

Lesson Notes

Lesson-specific insights from the creators of this lesson.

GIF

This lesson returns to a context covered in the previous lesson: an experiment in which researchers sent fake, identical resumés to employers. The resumés only differed in one aspect: the applicant’s name. Half of the resumés had a commonly-white name and half had a commonly-black name (as determined by birth certificate records). In the prior lesson, students created a confidence interval to analyze the difference in callback rates between the name groups. In this lesson, they’ll conduct a hypothesis test to see if the difference was statistically significant.

Learning Targets
  • Check conditions for a two-sample z-test for the difference between population proportions
  • Conduct a two-sample z-test for the difference between population proportions
  • Draw connections between confidence intervals and hypothesis tests

Before proceeding: Familiarize yourself with the lesson materials linked above (e.g. handout, handout key, slides, video). Then, for additional background and teaching tips from the lesson creators, check out the sections below.


  • Although this lesson analyzes the same study as the prior lesson (Lesson 3.B.2), the materials are designed without presuming that students have seen the prior lesson. So, instructors can use this lesson even if they haven’t used the prior lesson. However, instructors who have used the prior lesson with their students can save time in class by keeping the discussion of the study design as a brief resurfacing, since this material was also covered in the prior lesson.
  • Because this study is a well-designed experiment with random assignment to treatment, we can conclude the difference in callback rates is caused by the names. However, the underlying mechanism behind the name advantage is less clear. Intentional racial discrimination is one possible mechanism. However, it’s also possible that the differences in callback rates could arise from unintentional biases. For example, commonly-white names may tend to be more familiar for the people making hiring decisions, and they may unconsciously gravitate to such names. Framing these considerations with students is important for interpreting the study and provides a helpful review of the concept of study generalizability.
  • This lesson closely parallels the previous one, providing a natural opportunity to compare and contrast confidence intervals and hypothesis tests. Reinforce that confidence intervals estimate a range of plausible values for a parameter, whereas hypothesis tests focus on evaluating the feasibility of one particular parameter value.
  • This lesson introduces the combined proportion \( (\hat{p}_c) \) used to calculate the standard error for a two-sample hypothesis test. It can be helpful to emphasize that combining reflects the assumption made under the null hypothesis: that there is truly no difference between the population proportions.

First, download this lesson's Handout Key and read through its Discussion Question section. Then, check out our model discussion norms and the additional background notes below.

  • This question provides an excellent opportunity to loop back to the topic of generalizability. The resumé study measures a specific context: callbacks for resumés sent to employers in sales, administrative support, clerical services, and customer services. So, its results may not fully generalize to other parts of the hiring process (e.g. interviews) or other industries (e.g. academia, sports management, finance, etc.).
  • Another challenge to this study’s generalizability comes from the fact that these resumés were sent in response to newspaper ads for jobs. This is an uncommon channel for finding a job in today’s world. More study is needed of the dynamics of more modern job application channels.
  • This lesson is based on a landmark audit study (free working paper version here) by economists Marianne Bertrand and Sendhil Mullainathan. Researchers sent nearly 5,000 fictitious resumés to job advertisements in Chicago and Boston, randomly assigning each resumé either a commonly-white or commonly-black first name. The jobs represented a variety of occupations that required different levels of experience and education, strengthening the study's relevance across multiple hiring settings. The researchers used birth certificate records to identify the names that were most exclusively associated with each racial group. Multiple male and female names were used within each group, helping ensure that the results were not driven by any single name.
  • The Bertrand and Mullainathan study became one of the best-known examples of an audit study. Since its publication, researchers have conducted numerous studies using the same basic audit study design in a variety of hiring contexts. Together, these studies illustrate how statistical evidence accumulates through repeated investigation.
  • Technically, because this lesson describes an experiment with random assignment to treatment, the sampling distribution should actually be referred to as the “randomization distribution.” Rather than describing all possible random samples (sampling distribution), randomization distributions describe all possible random assignments to treatment. However, the sampling distribution and randomization distribution are mathematically equivalent, and the distinction between them is not important in AP Statistics.
  • Generally, the following relationships tend to hold true:
    • If interval includes null → fail to reject null in test.
    • If interval does not include null → reject null in test.
  • However, although the relationships above are true most of the time, they are not strictly true for every case. For two-sample procedures, the standard error for intervals and tests are calculated using slightly different methods. So, in some edge cases, the relationships won’t hold. Nonetheless, these relationships provide good rules of thumb and conceptual connections for students.
  • Although experiments unlock the ability to make causal inferences, they can also be expensive and difficult to implement. Because of the cost, researchers often have to consider tradeoffs between causal inference and generalizability. For instance, it’s relatively easy and inexpensive to gather observational data about racial wage gaps in the workforce. However, this observational data doesn’t allow for causal conclusions. By contrast, the resumé experiment allows for causal conclusions. However, its results can’t be generalized to unstudied parts of the hiring process (e.g. interviews) or to other industries that weren’t sent resumés during the experiment.

Student Supports

Lesson-specific resources to support all learners.

  • Because both sample proportions are relatively small, it can be helpful to express them as percentages (rather than decimals) before comparing them. Seeing that 10.1% is noticeably larger than 6.7% often makes the direction and size of the observed difference easier to recognize before calculating the difference in proportions.
  • Reinforce that the combined sample proportion \( (\hat{p}_c) \) is used when checking the large counts condition for a two-sample hypothesis test. This is the only inference procedure in the AP Statistics course that uses the pooled (combined) proportion, making it a useful distinction to emphasize.
  • The calculation of the combined sample proportion can feel counterintuitive because students have long been taught that fractions cannot be combined by simply adding numerators and denominators. It can be helpful to emphasize that we are not combining two proportions. Instead, we are combining the successes from both samples and the observations from both samples to form one pooled sample before calculating a single proportion. Under the null hypothesis, this pooling is appropriate because the null hypothesis assumes both populations have the same true proportion.
  • Vocabulary used in the context of the lesson may include words that are unfamiliar or have several meanings. In particular, the following mathematical terms may need clarification or a definition provided:
    • Test statistic
    • p-value
    • Difference of proportions
  • In addition, the following contextual terms may need clarification or a definition provided:
    • Callback rate
    • Resumé