54 FRQs across all 9 units • Scoring guides • Sample responses
Scoring: E = Essentially correct (full credit) • P = Partially correct (half credit) • I = Incorrect (no credit)
A teacher recorded the number of absences for students in two sections of AP Statistics during a semester:
| Section A (n=28) | Section B (n=32) | |
|---|---|---|
| Mean | 3.5 | 5.8 |
| Median | 3.0 | 4.0 |
| SD | 2.1 | 4.3 |
| Min | 0 | 0 |
| Q1 | 2 | 2 |
| Q3 | 5 | 8 |
| Max | 9 | 19 |
(a) Compare the distributions of absences for the two sections. Include shape, center, spread, and unusual features.
(b) The teacher wants to report a "typical" number of absences for each section. Should the teacher use the mean or the median for Section B? Justify your answer.
The scores on a standardized math test are approximately Normally distributed with a mean of 500 and a standard deviation of 100.
(a) What proportion of students score above 650? Show your work.
(b) A student scored at the 90th percentile. What was their approximate score? Show your work.
(c) A different test has a mean of 24 and SD of 4. A student scored 30 on this test and 680 on the first test. On which test did the student perform better relative to other test-takers? Justify using z-scores.
A researcher studies the relationship between hours of sleep (x) and reaction time in milliseconds (y) for 20 adults. The LSRL is: y-hat = 450 - 15.2x, with r = -0.87 and r² = 0.757.
(a) Interpret the slope in context.
(b) Interpret r² in context.
(c) One subject slept 7 hours and had a reaction time of 360 ms. Calculate and interpret the residual.
(d) Would it be appropriate to use this model to predict reaction time for someone who slept 2 hours? Explain.
A student fits a linear model to data relating the age of a used car (in years) to its selling price (in thousands of dollars). The LSRL is: price-hat = 28.5 - 2.4(age). The residual plot shows a clear curved (U-shaped) pattern.
(a) Based on the residual plot, is a linear model appropriate for these data? Explain.
(b) The student observes that r² = 0.89. Does this high r² value mean the linear model is a good fit? Explain.
A physical therapist wants to compare the effectiveness of two stretching routines (Routine A and Routine B) on improving flexibility. She has 40 volunteers.
(a) Describe a completely randomized design for this experiment. Be specific about how you would assign subjects and what you would measure.
(b) Explain why random assignment is important in this experiment.
(c) Describe how she could redesign this as a matched pairs experiment. What is one advantage?
A city council wants to know whether residents support building a new community center. They place a survey at the front desk of the current community center and collect 200 responses. 78% support the project.
(a) Identify a source of bias in this sampling method and explain the likely direction of the bias.
(b) Describe a better sampling method the city council could use to get a more representative result.
A carnival game costs $2 to play. A player spins a wheel with these outcomes:
| Prize | $0 | $1 | $5 | $20 |
|---|---|---|---|---|
| Probability | 0.50 | 0.30 | 0.15 | 0.05 |
(a) Find the expected value of the net profit per game. Show your work.
(b) Find the standard deviation of the net profit.
(c) A player plays 50 independent games. Find the expected total net profit and the standard deviation of the total.
A manufacturer claims that 90% of its products pass quality inspection. A store receives a shipment of 20 items and inspects all of them.
(a) Explain why this situation can be modeled with a binomial distribution. Identify n and p.
(b) What is the probability that exactly 18 of the 20 items pass inspection?
(c) If only 15 of the 20 items pass, should the store doubt the manufacturer's claim? Justify using probability.
A large university reports that the mean GPA of all students is 3.1 with a standard deviation of 0.6. A random sample of 36 students is selected.
(a) Describe the sampling distribution of the sample mean GPA. Include shape, mean, and standard deviation. Justify any conditions.
(b) What is the probability that the sample mean GPA is greater than 3.3?
(c) Would your answer to part (a) change if the sample size were 10 instead of 36? Explain.
Suppose 35% of all adults in a city exercise regularly. A random sample of 150 adults is selected.
(a) Check the conditions for the sampling distribution of p-hat to be approximately Normal.
(b) Find the mean and standard deviation of the sampling distribution of p-hat.
(c) What is the probability that more than 40% of the sample exercises regularly?
A polling organization claims 60% of voters support a new park. A reporter suspects the proportion is different and surveys 150 random voters, finding 80 supporters.
(a) Conduct a significance test at α = 0.05 using the full 4-step inference procedure.
(b) Based on your test, would a 95% confidence interval contain 0.60? Explain without calculating.
In a random sample of 500 high school seniors, 215 say they plan to attend a 4-year university.
(a) Construct a 95% confidence interval for the proportion of all seniors who plan to attend a 4-year university. Follow the full inference procedure.
(b) Interpret the interval in context.
(c) A guidance counselor claims that a majority of seniors plan to attend a 4-year university. Does your interval support this claim? Explain.
A cereal company claims each box contains 16 oz. A consumer group suspects the boxes are underfilled. They randomly select 25 boxes and find x-bar = 15.7 oz with s = 0.8 oz.
(a) Conduct a significance test at α = 0.05. Use the full 4-step procedure.
(b) Based on your conclusion, what type of error (Type I or Type II) could you have made? Describe it in context.
A teacher wants to compare test performance between two class sections. Section 1 (n=30): x-bar=78.2, s=10.5. Section 2 (n=28): x-bar=82.8, s=9.3.
(a) Is there convincing evidence of a difference in mean test scores? Conduct a two-sample t-test at α = 0.05.
(b) A 95% confidence interval for μ1 - μ2 is (-9.88, 0.68). Interpret this interval and explain how it's consistent with your test result.
400 adults were surveyed about news source preference and age group:
| TV | Online | Total | ||
|---|---|---|---|---|
| 18-39 | 30 | 110 | 10 | 150 |
| 40-64 | 55 | 70 | 25 | 150 |
| 65+ | 65 | 20 | 15 | 100 |
| Total | 150 | 200 | 50 | 400 |
(a) State appropriate hypotheses.
(b) Calculate the expected count for "18-39, TV." Show the formula.
(c) χ² = 82.44, p-value ≈ 0. State your conclusion at α = 0.05.
(d) Which cells contribute most to χ²? What does this reveal about the association?
A die manufacturer claims their die is fair. You roll it 300 times and get: 1 (58), 2 (45), 3 (52), 4 (49), 5 (51), 6 (45).
(a) State hypotheses for a goodness-of-fit test.
(b) What is the expected count for each face? Check the Large Counts condition.
(c) Calculate the chi-square statistic and state the degrees of freedom.
A biologist studies temperature (°F) and cricket chirps per minute at 15 locations. Computer output:
| Predictor | Coef | SE Coef | T | P |
|---|---|---|---|---|
| Constant | -0.31 | 3.10 | -0.10 | 0.922 |
| Temperature | 0.212 | 0.0437 | 4.85 | 0.0003 |
S = 3.58 R-sq = 64.4%
(a) Write the LSRL equation. Define variables.
(b) Interpret the slope in context.
(c) Interpret r² in context.
(d) Is there convincing evidence of a linear relationship? Conduct a test at α = 0.05 using the output.
(e) Construct a 95% CI for the true slope. (t* = 2.160, df = 13)
A student fits a regression to predict house price from square footage (n = 40). The scatterplot looks roughly linear. The residual plot shows a slight fan shape (spread increases with x). A histogram of residuals is approximately symmetric and bell-shaped.
(a) Check each of the five LINER conditions using the information provided. For each, state whether it is met, not met, or cannot be determined.
(b) The "Equal variance" condition appears violated. What impact might this have on inference for the slope?
A histogram of 200 students' commute times (in minutes) to school shows a unimodal shape with a long tail to the right. The mean is 22 minutes, the median is 18 minutes, the standard deviation is 12 minutes, and the IQR is 14 minutes. One student has a commute of 65 minutes.
(a) Describe the distribution of commute times. Address all four components (shape, outliers, center, spread).
(b) Is the student with a 65-minute commute an outlier by the 1.5 × IQR rule? Show your work. (Q1 = 12, Q3 = 26)
(c) Would you recommend the mean or median to describe a typical commute time? Justify.
Battery life of a certain phone model is approximately Normally distributed with mean 11.5 hours and standard deviation 1.8 hours.
(a) What proportion of phones have a battery life between 10 and 14 hours?
(b) The company wants to advertise a guaranteed minimum battery life such that 90% of phones exceed it. What value should they use?
Side-by-side boxplots show the test scores for two classes. Class A: min=55, Q1=68, med=76, Q3=84, max=95. Class B: min=40, Q1=60, med=72, Q3=88, max=100.
(a) Compare the two distributions in context (center, spread, shape).
(b) Which class performed more consistently? Justify using a specific measure of spread.
A set of test scores has mean 72 and standard deviation 8. The teacher applies a curve: new score = 1.1(old score) + 5.
(a) Find the new mean and new standard deviation.
(b) A student originally scored 80. What is their new score, and did their z-score change? Explain.
A study of 30 cities examines the relationship between average temperature (°F) and monthly electricity usage (kWh). The LSRL is: usage-hat = 1450 - 12.3(temperature). r = -0.78, r² = 0.608.
(a) Identify the explanatory and response variables.
(b) Interpret the slope and y-intercept in context. Is the y-intercept meaningful?
(c) Interpret r² in context.
(d) Predict usage for a city with avg temp 70°F. Then explain why predicting for a city with avg temp 120°F would be problematic.
A scatterplot of 20 data points shows a moderate positive linear relationship with r = 0.55. One point has a very high x-value and falls far below the regression line. When this point is removed, r increases to 0.88.
(a) Is this point influential? Explain how you know.
(b) Does this point have high leverage? Explain.
(c) Should the researcher remove this point from the analysis? Discuss considerations.
Three studies examine the relationship between sleep and academic performance:
Study 1: A random sample of 500 college students is surveyed about sleep hours and GPA. More sleep is associated with higher GPA.
Study 2: 80 volunteers are randomly assigned to either a "sleep coaching" program or a control group. After 8 weeks, the coached group has significantly higher GPAs.
Study 3: Researchers randomly select 200 students from a university and randomly assign them to sleep coaching or control. The coached group shows higher GPAs.
(a) For each study, state whether we can generalize to a larger population and whether we can conclude causation. Justify each.
A farmer wants to test whether a new fertilizer increases crop yield compared to the standard fertilizer. He has 30 identical plots of land available. He suspects that plots on the east side of the farm receive more sunlight than those on the west side.
(a) Describe a completely randomized design.
(b) Describe a randomized block design that accounts for the sunlight difference. Explain why blocking improves the experiment.
(c) What is the purpose of having a control group (standard fertilizer) rather than comparing the new fertilizer to no fertilizer at all?
A school surveyed 400 students about transportation and grade level:
| Car | Bus | Walk | Total | |
|---|---|---|---|---|
| Underclass (9-10) | 40 | 120 | 40 | 200 |
| Upperclass (11-12) | 100 | 60 | 40 | 200 |
| Total | 140 | 180 | 80 | 400 |
(a) What is P(Car | Upperclass)?
(b) What is P(Upperclass | Car)?
(c) Are the events "Car" and "Upperclass" independent? Justify using probabilities.
A coffee shop sells small coffees (mean profit $1.50, SD $0.40) and pastries (mean profit $2.00, SD $0.60). Assume the profit from a coffee sale and a pastry sale are independent.
(a) A customer buys one coffee and one pastry. Find the mean and SD of the total profit.
(b) Another customer buys 2 coffees (no pastry). Find the mean and SD of the total profit from both coffees.
(c) Explain why SD(coffee + coffee) ≠ 2 × SD(coffee).
Wait times at a DMV are strongly right-skewed with μ = 45 minutes and σ = 20 minutes.
(a) Can we find P(individual wait > 60 min) using the Normal distribution? Explain.
(b) For a random sample of n = 64, describe the sampling distribution of x-bar. Justify the use of the Normal model.
(c) What is P(x-bar > 50) for a sample of 64?
Two students each collect 100 random samples from the same population (true mean μ = 50). Student A uses samples of size 10; Student B uses samples of size 100. Both plot the sampling distributions of x-bar.
(a) What will be the same about the two sampling distributions?
(b) What will be different? Be specific.
(c) Which student's estimates will be more useful for estimating μ? Why?
A company tests two website designs. Design A: 120 of 500 visitors made a purchase. Design B: 155 of 500 visitors made a purchase.
(a) Is there convincing evidence that Design B has a higher purchase rate? Conduct a full significance test at α = 0.05.
(b) Construct a 95% confidence interval for p_B - p_A and interpret it.
A pharmaceutical company tests whether a new drug lowers cholesterol more than the current standard. H0: The new drug is no better than the standard. Ha: The new drug is better.
(a) Describe a Type I error in this context. What would be a consequence?
(b) Describe a Type II error in this context. What would be a consequence?
(c) The company uses α = 0.01 instead of 0.05. How does this affect the probability of each type of error?
(d) Name two ways to increase the power of this test.
A trainer measures the resting heart rate of 12 participants before and after an 8-week exercise program. The mean difference (before - after) is d-bar = 5.2 bpm with s_d = 6.8 bpm.
(a) Is there convincing evidence that the program reduces resting heart rate? Conduct a matched pairs t-test at α = 0.05.
(b) Construct a 95% confidence interval for the true mean difference and interpret it.
For each scenario, identify the appropriate inference procedure and justify your choice.
(a) A nutritionist randomly selects 40 adults and measures their daily calcium intake. She wants to estimate the mean intake for all adults.
(b) A researcher compares test scores between students who took an online course (n=35) and students who took an in-person course (n=40). The groups are independent.
(c) A dentist measures cavity count for 25 patients before and after switching to a new toothpaste.
(d) A pollster wants to know whether the proportion of voters supporting a candidate exceeds 50%, based on a random sample of 600 voters where 324 expressed support.
Three hospitals tracked patient satisfaction (Satisfied, Neutral, Dissatisfied):
| Satisfied | Neutral | Dissatisfied | Total | |
|---|---|---|---|---|
| Hospital A | 80 | 30 | 10 | 120 |
| Hospital B | 75 | 45 | 30 | 150 |
| Hospital C | 55 | 40 | 35 | 130 |
| Total | 210 | 115 | 75 | 400 |
(a) Is this a test for independence or homogeneity? Explain.
(b) State the hypotheses.
(c) Calculate the expected count for Hospital A, Satisfied.
(d) The test gives χ² = 22.1 with df = 4 and p-value < 0.001. State your conclusion and identify which hospital appears most different.
A genetics model predicts offspring in a 9:3:3:1 ratio for four phenotypes. A researcher observes 160 offspring: 85, 35, 26, 14.
(a) State the hypotheses.
(b) Calculate the expected counts for each phenotype.
(c) Verify the conditions for the chi-square test.
(d) Calculate the chi-square statistic and state the degrees of freedom.
A study examines the relationship between study hours per week (x) and exam score (y) for n = 20 students. Computer output:
| Predictor | Coef | SE Coef | T | P |
|---|---|---|---|---|
| Constant | 52.3 | 5.8 | 9.02 | <0.001 |
| StudyHours | 3.45 | 0.82 | 4.21 | 0.0005 |
S = 8.4 R-sq = 49.6%
(a) Write the LSRL and interpret the slope in context.
(b) Conduct a full 4-step test for the significance of the slope at α = 0.05.
(c) Construct a 95% CI for β. (t* = 2.101 for df = 18)
(d) Interpret S = 8.4 in context.
A survey of 600 adults asked about exercise frequency and smoking status:
| Smoker | Non-Smoker | Total | |
|---|---|---|---|
| Exercises Regularly | 40 | 260 | 300 |
| Does Not Exercise | 80 | 220 | 300 |
| Total | 120 | 480 | 600 |
(a) Calculate the conditional distribution of smoking status for each exercise group.
(b) Is there an association between exercise and smoking? Justify using the conditional distributions.
A study finds a strong positive correlation (r = 0.82) between the number of firefighters sent to a fire and the amount of damage caused by the fire.
(a) A reporter writes: "Sending more firefighters causes more damage." Explain why this conclusion is not valid.
(b) Identify a likely lurking variable and explain how it accounts for the observed association.
(c) What type of study would be needed to establish a causal relationship? Is such a study practical here?
A radio host asks listeners to call in and vote on whether the minimum wage should be raised. 73% of 2,400 callers say yes.
(a) Identify the sampling method and explain why the result may not represent all adults.
(b) Describe two specific sources of bias and the likely direction of each.
A study finds that children who eat breakfast daily have higher test scores than those who skip breakfast.
(a) Can we conclude that eating breakfast causes higher test scores? Explain.
(b) Identify two potential confounding variables and explain how each could account for the association.
(c) Describe an experiment that could help determine whether breakfast actually causes better performance.
A basketball player makes 80% of her free throws. She shoots until she misses.
(a) What is the probability that her first miss comes on the 5th shot?
(b) What is the expected number of shots until her first miss?
(c) What is the probability she makes at least 3 shots before her first miss?
At a college, 60% of students have a job, 45% are in a club, and 20% have both a job and are in a club.
(a) Find P(job OR club).
(b) Find P(club | job).
(c) Are "has a job" and "is in a club" independent events? Justify mathematically.
A university claims that 70% of its graduates find employment within 6 months. A newspaper surveys a random sample of 200 recent graduates and finds that 128 (64%) found employment within 6 months.
(a) Identify the population, sample, parameter, and statistic in this context.
(b) Does the sample result prove that the university's claim is false? Explain using the concept of sampling variability.
A population has proportion p = 0.40.
(a) For n = 25, check whether the sampling distribution of p-hat is approximately Normal.
(b) For n = 100, find the mean and SD of the sampling distribution of p-hat and check the Normal condition.
(c) How many times larger must the sample be to cut the standard deviation of p-hat in half?
A test of H0: p = 0.50 vs Ha: p > 0.50 yields p-value = 0.03.
(a) Interpret the p-value in context (assume the test is about whether a majority of residents support a policy).
(b) At α = 0.05, state the conclusion in context.
(c) A student says "the p-value of 0.03 means there is a 3% chance the null hypothesis is true." Correct this interpretation.
A researcher wants to estimate the proportion of adults who support a new law with a margin of error no more than 3 percentage points at the 95% confidence level.
(a) If no prior estimate of p is available, what sample size is needed? Show your work.
(b) If a pilot study suggests p ≈ 0.30, what sample size is needed? How does this compare to part (a)?
A random sample of 36 packages from an assembly line has a mean weight of 16.2 oz with s = 0.9 oz. The target weight is 16.0 oz.
(a) Construct a 95% confidence interval for the true mean weight. Follow the inference procedure.
(b) Based on your interval, is there evidence the mean weight differs from the target of 16.0 oz?
(c) If the sample size were increased to 144 (same x-bar and s), how would the interval change?
A large study (n = 10,000) tests whether a new tutoring program raises SAT scores. H0: μ = 1060 (national avg). The sample mean is x-bar = 1063 with s = 200. The t-test gives t = 1.50, p = 0.067 at α = 0.05. But a 95% CI is (1059.1, 1066.9).
(a) State the conclusion of the hypothesis test at α = 0.05.
(b) Even if the result had been significant, would a 3-point increase on the SAT be practically meaningful? Discuss.
(c) Explain why a very large sample size can lead to statistically significant results that are not practically important.
Study A: A random sample of 500 employees at one company is classified by department AND job satisfaction level.
Study B: Random samples of 150 employees are taken from each of three different companies. Each employee is classified by job satisfaction level.
(a) Which study uses a chi-square test for independence? Which uses homogeneity? Explain.
(b) Write appropriate null and alternative hypotheses for each study.
A candy company claims their bags contain 30% red, 25% blue, 20% green, 15% yellow, and 10% orange candies. A student counts 200 candies: 70 red, 42 blue, 38 green, 28 yellow, 22 orange.
(a) Conduct a full chi-square goodness of fit test at α = 0.05. Include hypotheses, conditions, test statistic, and conclusion.
Computer output for a regression of height (inches) on shoe size for n = 30 adults:
| Predictor | Coef | SE Coef | T | P |
|---|---|---|---|---|
| Constant | 49.8 | 4.2 | 11.86 | <0.001 |
| ShoeSize | 1.65 | 0.45 | 3.67 | 0.001 |
S = 2.95 R-sq = 32.5%
(a) Identify the values of b, SE(b), t, and p from the output.
(b) Interpret S = 2.95 in context.
(c) A friend says "the model only explains 32.5% of the variability, so it's useless." Respond to this claim.
A researcher studies the relationship between advertising spending (thousands of $) and sales revenue (thousands of $) for n = 18 stores. A 95% confidence interval for the slope is (0.8, 4.2).
(a) Interpret this confidence interval in context.
(b) Based on this interval, would a two-sided test of H0: β = 0 at α = 0.05 reject the null? Explain without performing the test.
(c) Would a test of H0: β = 5 at α = 0.05 reject? Explain.
A biologist studies the relationship between the age (x, in days) and weight (y, in grams) of a growing organism. A scatterplot of y vs x shows a clear curved pattern. The residual plot for a linear model confirms a nonlinear relationship. The biologist then takes ln(y) and fits a linear model to ln(y) vs x. The new residual plot shows random scatter.
(a) Why was the original linear model inappropriate?
(b) The regression equation for the transformed data is: ln(y-hat) = 1.2 + 0.045x. What type of model does this suggest for the original (untransformed) data?
(c) Use the model to predict the weight of a 30-day-old organism.