Lesson 3.B.1 - Sampling Distribution for Two Proportions
Key Question: Will AI take peoples' jobs?
Content: One vs Two Samples | Sampling Distribution for Two Proportions & Conditions
Alignment: CED Topic 3.9
Video
Course Resources
Resources for teaching our AP® Statistics curriculum.
- Lesson Flow - timing and flow of class, using our lesson materials
- Pacing Guide - pacing our units, with daily or block schedules
- CED Alignment Guide - aligning our lessons to the AP® Statistics Course and Exam Description
Teaching Resources
Resources for teaching with Skew The Script.
- Discussion Norms - our model discussion norms for the classroom
- Letter to Parents - letter to share with parents about our nonpartisan approach
- Teaching Math on Civic Topics - tips for teaching math lessons that cover civic topics
Lesson Notes
Lesson-specific insights from the creators of this lesson.
Traditionally, when someone has an idea for a new software, website, or app, they hire coders to help them build it. However, instead, it may now be possible to use AI agents to translate new ideas into functioning code. How often does this happen? Will AI agents really start replacing programmers and other professionals – both now and in the long run? In this lesson, students explore these questions by analyzing trends in the unemployment rate among computer science graduates.
- Distinguish between one-sample and two-sample scenarios
- Find the sampling distribution for the difference between two population proportions and use it to evaluate claims
- Check the conditions for sampling two proportions and describe the purpose of each condition
Before proceeding: Familiarize yourself with the lesson materials linked above (e.g. handout, handout key, slides, video). Then, for additional background and teaching tips from the lesson creators, check out the sections below.
- The fast spread of generative AI models has also led to the fast spread of opinions, predictions, and perspectives about them. Ultimately, the outcomes driven by emerging technologies are difficult to predict – especially outcomes related to employment and the marketplace. At the outset of the lesson, students might share strong opinions for / against the merits of AI and about if / when AI will impact employment rates. Encouraging these initial perspectives will help build interest in the lesson. However, we recommend quickly moving into the main part of the lesson, in which students use real data sets to empirically explore employment and AI. After this data-based analysis, the Discussion Question presents an excellent opportunity to then re-open the broader conversation about AI and the future of employment – now informed by the data analysis undertaken in the lesson.
- Throughout this lesson, it’s helpful to emphasize that the parameter of interest isn’t a proportion. Rather, it’s the difference between two proportions p1 - p2. Therefore, a parameter value of zero (p1 - p2 = 0) indicates no difference between population proportions. In other words, values close to zero indicate similarity between proportions. Continually pointing to the difference as the key quantity of interest will help students internalize this important framework for the rest of the unit.
- It’s also helpful to emphasize that the order of subtraction should remain consistent throughout the analysis. The interpretation of the observed difference, the expected difference, and the final conclusion all depend on using the same ordering for both the sample statistics and the population parameters.
First, download this lesson's Handout Key and read through its Discussion Question section. Then, check out our model discussion norms and the additional background notes below.
- Encourage students to consider both the evidence presented in the lesson and the limitations of that evidence. The data we analyze in the lesson describe changes in unemployment among recent computer science graduates, but they do not establish AI as the sole cause of those changes. The data also only come from recent years and, therefore, cannot establish a long-term trend on their own. These discussions will surface helpful themes from earlier in the course – e.g. generalizability, causal inference, and study design – that shape the scope of potential conclusions.
- Ultimately, it can be helpful to acknowledge that no one can predict the future – especially when it comes to the complex interactions between technology and the market. Therefore, this discussion will necessarily be at least partly speculative. Recognizing this fact will help students be curious and empathetic about one another’s views, and it will help them get comfortable with uncertainty – an important trait for any statistician.
- The data used in this lesson come from the American Community Survey (ACS), a stratified random sample of the US population conducted by the US Census Bureau each year. Because the ACS gathers detailed demographic, economic, and social data and uses rigorous sampling methods, it is widely used for research and public policy.
- Out of curiosity, we asked an AI chatbot (ChatGPT) what it thought about this lesson. After reviewing the lesson slides, here was its reply: “The central theme is timely and compelling because it uses students’ real concerns about AI and employment to motivate two-proportion inference, while the ATM comparison adds useful nuance and prevents the lesson from settling for an overly simple conclusion.” Stop, we’re blushing. (but also, see the unfortunate potential downsides of AI’s tendency towards flattery)
- The standard deviation of the sampling distribution for the difference between two proportions is introduced without derivation. The emphasis in this lesson is on interpreting the model and understanding when it is appropriate to apply it, rather than proving the formula itself. However, this optional video is available for students and instructors that would like to see its informal derivation.
- This lesson’s analysis uses the sample unemployment rate in 2013 (p̂1 = 0.056) as the presumed true unemployment rate in both 2013 and 2024. A more precise analysis would utilize the pooled or combined sample unemployment rate from both years \( (\hat{p}_c = \frac{13+121}{233+867} = \frac{134}{1100} \approx 12.2\%) \) as the presumed true unemployment rate for both years. For now, we wanted to provide a simpler method, so that students wouldn’t get bogged down in these calculations and lose sight of the overall intuition. Instead, students will encounter the pooled or combined proportion when they learn about the two-sample z-test for the difference between population proportions in Lesson 3.B.3. We believe that is a more appropriate place to introduce this method, after students have developed intuition for the overall approach. That said, with either analysis, students will arrive at similar results and find convincing evidence that the true unemployment rate among recent computer science graduates is higher in 2024.
Student Supports
Lesson-specific resources to support all learners.
- Students are already familiar with the inference conditions for one proportion. This lesson provides a natural opportunity to emphasize that each condition must now be checked separately for both samples. Drawing parallels between the conditions for sampling one proportion and those for sampling two proportions can help students internalize the newer set of conditions.
- Encourage students to clearly identify Population 1 and Population 2 at the outset of the problem. A helpful support for this can be modifying the standard notation of p1 - p2 to utilize contextual subscripts. For example, if finding the difference between the proportions of students who passed the AP Exam in New York and in Texas, they could notate the difference as pNY - pTX. As long as students clearly define their variables, contextual subscripts are allowed on the AP Exam.
- Reinforce that a difference of zero represents “no difference” between the population proportions. This interpretation helps students understand why the sampling distribution is centered at zero when comparing two populations that are assumed to be equal.
- Vocabulary used in the context of the lesson may include words that are unfamiliar or have several meanings. In particular, the following mathematical terms may need clarification or a definition provided:
- Sampling distribution
- Two-sample inference
- Difference in proportions
- Independent samples
- In addition, the following contextual terms may need clarification or a definition provided:
- Unemployment rate
- AI agent
- Bank teller