Lesson 4.7 - Simulation Test for Two Proportions
Key Question: What causes racial wage gaps?
Content: One vs Two Samples | Simulation Test for Two Proportions
Video
Course Resources
Resources for teaching our High School Statistics curriculum.
- Lesson Flow - timing and flow of class, using our lesson materials
- Pacing Guide - pacing our units, with daily or block schedules
- Alignment Guide - aligning our lessons to national and state standards for high school statistics
- Classroom Routines - a guidebook of classroom routines embedded within our lessons
Teaching Resources
Resources for teaching with Skew The Script.
- Discussion Norms - our model discussion norms for the classroom
- Letter to Parents - letter to share with parents about our nonpartisan approach
- Teaching Math on Civic Topics - tips for teaching math lessons that cover civic topics
Lesson Notes
Lesson-specific insights from the creators of this lesson.
In this lesson, students explore an experiment in which researchers sent fake, identical resumés to employers. These resumés only differed in one aspect: the applicant’s name. Half of the resumés had a commonly-white name and half had a commonly-black name (as determined by birth certificate records). Students use the study data to find the difference in callback rates and conduct a hypothesis test to see if the difference is statistically significant.
- Distinguish between one-sample and two-sample scenarios
- Conduct a simulation-based hypothesis test for the difference between two population proportions
Now that students have thoroughly explored hypothesis tests for one proportion (Lessons 4.3 - 4.6), this lesson extends their learning to hypothesis tests for two proportions. Specifically, students distinguish a two-sample inference scenario (comparing proportions measured among two samples) from a one-sample inference scenario (comparing a proportion from one sample against a claim about that proportion). Then, they utilize computer simulation – a practice they first used in Lesson 4.3 – to evaluate whether differences between pairs of proportions are statistically significant. Learning these methods for analyzing two-sample scenarios gives students the opportunity to thoroughly analyze more real-world experiments, which are often designed to compare two samples or treatments.
Before proceeding: Familiarize yourself with the lesson materials linked above (e.g. handout, handout key, slides, video). Then, for additional background and teaching tips from the lesson creators, check out the sections below.
- Because the study explored by this lesson is a well-designed experiment with random assignment to treatment, we can conclude that the difference in callback rates is caused by the names. However, the underlying mechanism behind the name advantage is less clear. Intentional racial discrimination is one possible mechanism. However, it’s also possible that the differences in callback rates could arise from unintentional biases. For example, commonly-white names may tend to be more familiar for the people making hiring decisions, and they may unconsciously gravitate to such names. Framing these considerations with students is important for interpreting the study and provides a helpful review of the concept of study generalizability.
- The experiment analyzed in this lesson is part of a class of studies called resumé audit studies, which have been utilized in a variety of hiring contexts to test for potential biases. One example is this resumé audit study of male and female names, which we analyzed in Lesson 2.4. Taking a moment to look at the similarities and differences between these studies can provide a powerful connection for students.
First, download this lesson's slide deck and handout key to see the prompt and sample responses for the Lesson Starter. Then, check out the additional background notes below.
Instructional routine: Ten-Minute Talk. This application of a ten-minute talk may feel a bit different than previous uses, as students are considering the topic from multiple perspectives rather than diving deeply into it. It may even be helpful to encourage students to use a T-chart (as they do for the Same or Different routine) to group contexts according to “small difference” or “large difference.” The key feature of the routine is to provide space for students to jot down their thoughts before discussing with a partner and engaging in whole group discussion. You can find more background on implementing a Ten-Minute talk here.
Purpose & Background: The goal of this lesson starter is to prepare students for considering the difference between two proportions that they’ll analyze throughout the lesson: 10.1% and 6.7%. Although these proportions may not seem very different at face value, students will begin to see how the difference may be considered quite large in certain contexts. In particular, the difference between 10.1% and 6.7% may be considered large when applied to contexts with large populations or with high-stakes repercussions. Later in the lesson, students discover that these percentages are part of the results of the resumé audit experiment. During the lesson, students work to determine whether the difference between the proportions is statistically significant, practically important, or both.
First, download this lesson's handout key and read through its Discussion Question section. Then, check out our model discussion norms and the additional background notes below.
- A 3.4 percentage point difference may initially appear modest. Encouraging students to consider the overall callback rate provides additional context. Because relatively few applications received callbacks overall, even a difference of a few percentage points represents a substantial relative advantage. As the original study noted: “a white name yields as many more callbacks as an additional eight years of experience” (pg. 3). Framing numerical differences in terms of equivalent differences among other variables – such as years experience – can be a helpful way to gauge practical importance.
- Multiple viewpoints about the broader causes of racial wage gaps may emerge during discussion. The statistical evidence from this experiment supports a causal conclusion about the effect of the assigned names in this particular hiring context. Broader questions about labor market outcomes may extend beyond the scope of this single study. These limits to the study’s generalizability provide a useful opportunity to connect back to concepts of generalizability from Unit 2 (Study Design).
- This lesson is based on a landmark audit study (free working paper version here) by economists Marianne Bertrand and Sendhil Mullainathan. The researchers sent nearly 5,000 fictitious resumés to job advertisements in Chicago and Boston, randomly assigning each resumé either a commonly-white or commonly-black first name. The jobs represented a variety of occupations that required different levels of experience and education, strengthening the study's relevance across multiple hiring settings. The researchers used birth certificate records to identify the names that were most exclusively associated with each racial group. Multiple male and female names were used within each group, helping ensure that the results were not driven by any single name.
- The Bertrand and Mullainathan study became one of the best-known examples of an audit study. Since its publication, researchers have conducted numerous studies using the same basic audit study design in a variety of hiring contexts. Together, these studies illustrate how statistical evidence accumulates through repeated investigation.
- Although experiments unlock the ability to make causal inferences, they can also be expensive and difficult to implement. Because of the cost, researchers often have to consider tradeoffs between causal inference and generalizability. For instance, it’s relatively easy and inexpensive to gather observational data about racial wage gaps in the workforce. However, this observational data doesn’t allow for causal conclusions. By contrast, the resumé experiment allows for causal conclusions. However, its results can’t be generalized to unstudied parts of the hiring process (e.g. interviews) or to other industries that weren’t sent resumés during the experiment.
- To generate the types of computer simulations we use in the lesson, our preferred tool is stapplet. For simulations comparing two proportions, utilize the One Categorical Variable, Multiple Groups applet. Input the observed data. Then, scroll down to the inference section and choose simulation.
Student Supports
Lesson-specific resources to support all learners.
- Because both sample proportions are relatively small, it can be helpful to express them as percentages (rather than decimals) before comparing them. Seeing that 10.1% is noticeably larger than 6.7% often makes the direction and size of the observed difference easier to recognize (compared to 0.101 and 0.067) before calculating the difference in proportions.
- Continuing to emphasize the “in a world where” language is especially helpful as we enter two-sample inference, as it provides students with intuitive framing for interpreting simulations. For example: “In a world where there is no difference between these proportions in the population, how unusual was the difference we observed in our samples?”
- Vocabulary used in the context of the lesson may include words that are unfamiliar or have several meanings. In particular, the following mathematical terms may need clarification or a definition provided:
- Difference of proportions
- Two-sample inference
- In addition, the following contextual terms may need clarification or a definition provided:
- Callback rate
- Resumé