Lesson 4.A.2 - Interval for a Mean
Key Question: Can you make a livable wage on social media?
Content: t-Distribution | One-Sample t-Interval for a Population Mean
Alignment: CED Topics 4.2-4.3
Video
Course Resources
Resources for teaching our AP® Statistics curriculum.
- Lesson Flow - timing and flow of class, using our lesson materials
- Pacing Guide - pacing our units, with daily or block schedules
- CED Alignment Guide - aligning our lessons to the AP® Statistics Course and Exam Description
Teaching Resources
Resources for teaching with Skew The Script.
- Discussion Norms - our model discussion norms for the classroom
- Letter to Parents - letter to share with parents about our nonpartisan approach
- Teaching Math on Civic Topics - tips for teaching math lessons that cover civic topics
Lesson Notes
Lesson-specific insights from the creators of this lesson.
As students scroll through their social media feeds, they inevitably encounter content from individuals who are able to successfully make a living on social media. And a number of these individuals are not just making a living – they’re quite rich. So, this raises a question for students: Should they drop out of school to pursue fame and fortune online? In this lesson, we use real data from YouTube creators to explore the feasibility of this career path.
- Describe the t-distribution, its purpose, and its degrees of freedom
- Calculate and interpret t critical values
- Construct and interpret a one-sample t-interval for a population mean
Before proceeding: Familiarize yourself with the lesson materials linked above (e.g. handout, handout key, slides, video). Then, for additional background and teaching tips from the lesson creators, check out the sections below.
- Although teenagers tend to know more about social media than their teachers, we’ve found that students tend to be quite open – even excited – to seriously explore the topic of social media with data, the principles of statistics, and the guidance of their instructors. However, that openness can quickly disappear if students feel that they’re being “given a lecturing” about the downsides of social media or the infeasibility of social media careers. Instead, we recommend letting the data drive the discussion, so students can draw their own conclusions about both the typical earnings and the risks of becoming a content creator.
- It’s helpful to point out the ways in which this lesson extends students’ previous work with confidence intervals for proportions. The overall structure of “point estimate ± margin of error” is unchanged, along with the basic framework for performing inference procedures (the 5C Method). Making these connections helps reinforce prior learning and makes the newer content easier to retain.
- This lesson provides a natural opportunity to distinguish sampling variability (which confidence intervals account for) from sampling bias (which they do not). Taking a moment after the Discussion Question to discuss the difference between sampling variability and sampling bias will help students internalize the distinction for more inference lessons to come.
First, download this lesson's Handout Key and read through its Discussion Question section. Then, check out our model discussion norms and the additional background notes below.
- In directly estimating the true mean wage among YouTubers, we used the approximate number of accounts (6 million) that have 1,000+ subscribers (the threshold at which YouTube starts paying creators). It may be helpful to note that, if we had also included accounts that are trying to get monetized but haven’t been able to reach 1,000 subscribers, the average yearly pay per aspiring creator would be even lower than the $1,389 estimate we find in the lesson.
- One helpful takeaway to emphasize for students is that, although confidence intervals can capture sampling variability, they cannot capture bias. The random condition is how we check for bias. Although our sample appeared to satisfy the random condition, it did not. We only randomly sampled among the population of creators who make videos about how much they make on YouTube, rather than randomly sampling from the whole population of all YouTubers. This resulted in undercoverage bias. So, when checking the random condition, students must not only make sure the sample was random – but also that the sampling frame covers the whole population of interest.
- As a closing note to the discussion, consider connecting the bias in our sampling method to the biased sampling process that we unknowingly perform as we scroll through online content. According to data from the tech firm Pex, in 2019 (the same year as the lesson’s YouTube video sample) the top performing 0.77% of videos on YouTube captured a whopping 82.83% of total views on the platform. A disproportionately high share of the videos that appear on our feeds are videos that have already gone viral. So, when we scroll, we scroll mainly through viral content. But the videos on our feed are not a representative sample of all the videos on the platform. In reality, a lot more non-viral videos exist that we never see. So, it may seem like everyone is going viral, when that’s really not the case. Our feeds, much like our sample from the lesson, provide an unrepresentative view of how common it is to go viral (and earn money) online.
- The data for this lesson was collected in 2019. The sampled videos showed snippets of each YouTuber’s earnings. When projecting those snippets into a yearly salary, we took the earnings shown and scaled them across a full year of time. For example, if a YouTuber showed their earnings from the past month, we multiplied that figure by 12 to estimate their earnings in a year.
- Some students may point out that the sample distribution of YouTuber salaries is strongly right-skewed. While skew does pull the sample mean above the sample median, it does not explain why the confidence interval misses the true population mean so dramatically. If the population is similarly skewed, the population mean would also exceed the population median. Skew alone can’t explain the inaccuracy of our interval. As noted in the Discussion Question notes of the handout key, the primary issue is sampling bias, not skewness.
- Consider sharing the story of William Sealy Gosset, the Guinness brewer who developed the t-distribution to make reliable decisions from small samples. Because Guinness did not allow employees to publish under their own names, Gosset published under the pseudonym “Student,” giving rise to the name Student’s t-distribution. This historical context provides a memorable introduction to the purpose of the t-distribution, while highlighting how statistical methods have been developed to solve real-world problems.
- Degrees of freedom measure how much independent information is available to estimate population values. One way to build intuition is to imagine creating a data set of five values, with only one requirement: when we take the average of those values, the result has to be 10. You can choose the first four data values however you like. But once the first four values are chosen, the fifth value is already determined. For example, imagine we choose the first four values as the following: 1, 8, 11, and 14. In order to make the average 10, the next value has to be 16. It cannot be any other value. So, although there are five observations, only four could be chosen independently, giving n - 1 degrees of freedom.
- As the sample size (n) increases, the sample standard deviation becomes a more precise estimate of the population standard deviation. Because there is less additional uncertainty from estimating σ with s, the t-distribution becomes increasingly similar to the normal distribution. In fact, if we take the limit as n approaches infinity, the t-distribution becomes equivalent to the standard normal distribution.
- The bias in the lesson’s sample doesn’t come from a failure to sample randomly. The sample was, in fact, performed randomly. Instead, the bias comes from failing to sample from the correct population. The population of interest is all YouTubers. But the sample is performed only among the population of YouTubers who make videos about their earnings. The population from which a sample is drawn is known as the “sampling frame.” When the sampling frame doesn’t match the whole population of interest, the result is undercoverage bias. Although the term “sampling frame” is outside the scope of AP Statistics, it can be a useful term to introduce to students as they continue discussing sampling throughout the course.
Student Supports
Lesson-specific resources to support all learners.
- Encourage students to make a sketch of the t-curve when calculating t critical values for a given confidence level and sample size. The drawing can help students clarify their thinking about areas under the curve, which leads to fewer mistakes for the inputs they need to put into their calculators.
- In the main lesson example, the sample size was large enough to invoke the Central Limit Theorem. However, in the practice exercises and on assessments, students will need to perform t-procedures for smaller sample sizes (n < 30). When the Central Limit Theorem cannot be invoked, encourage students to draw a dotplot of the sample data. As long as the graph does not show strong skewness or outliers, they can justify an assumption that the population may be approximately normal, allowing them to proceed with their inference procedure.
- Vocabulary used in the context of the lesson may include words that are unfamiliar or have several meanings. In particular, the following mathematical terms may need clarification or a definition provided:
- Degrees of freedom
- Standard error
- t-distribution
- Critical value
- Undercoverage bias
- In addition, the following contextual terms may need clarification or a definition provided:
- Livable wage
- Monetized
- Content creator
- Ad revenue