Crash Course Statistics probability, confidence, and p-values lesson bundle cover for grades 9–12

P-Values Without the Misconceptions: A High School Classroom Routine

P-values are difficult for students for a simple reason: the number is easy to calculate or look up, but the interpretation is easy to overstate.

The goal of a high-school lesson should not be “memorize p < .05.” Students need a routine for deciding what a result does—and does not—justify.

Start with what a p-value does not mean

The American Statistical Association's Statement on Statistical Significance and P-Values was written in part because p-values were being interpreted too mechanically. Three classroom-safe corrections are especially important:

  • A p-value does not tell us the probability that the null hypothesis is true.
  • A p-value does not measure how large or important an effect is.
  • A scientific, business, or policy decision should not rest only on whether a result crosses one threshold.

This gives students a better question than “Is it significant?” Ask: What does this result add to the evidence, and what else do we need before making a decision?

Use a five-step classroom routine

  1. Identify the claim. What is the study or headline asking us to believe?
  2. Identify the comparison. What groups, conditions, or predictions are being compared?
  3. Interpret the p-value carefully. If the model assumptions hold, how surprising would results this extreme be under the null model?
  4. Check effect size and practical importance. Is the observed difference large enough to matter in the real setting?
  5. Check the study design and uncertainty. Was the sample appropriate? Was the study large enough? Were many analyses tried? Is the result consistent with other evidence?

Statistical significance is not practical significance

Large samples can detect small differences. The National Center for Education Statistics explains that interpreting statistical significance in large-scale assessments requires attention to both the magnitude of a difference and the uncertainty around estimates. See the NCES guide to statistical significance and sample size.

Give students a simple example: a study with tens of thousands of people finds that an intervention changes an average score by 0.2 points on a 100-point scale and reports p < .001. The result may be statistically detectable while still being too small to matter for a classroom decision.

Then reverse the situation: a small pilot study finds a noticeable difference but p = .09. Students should not automatically conclude “there is no effect.” A small sample may have low statistical power and may be unable to distinguish a real signal from noise reliably.

Teach Type I and Type II errors as decision risks

Students often learn Type I and Type II errors as vocabulary pairs and then forget them. Make the ideas concrete:

  • Type I error: acting as if an effect is real when it is not.
  • Type II error: missing an effect that is actually present.

The right balance depends on the consequence. A screening test, a classroom pilot, a court decision, and a medical trial do not all have the same cost for false alarms and missed signals.

Connect p-values to study design

A p-value cannot repair a weak design. Observational associations do not automatically establish cause and effect; the U.S. National Library of Medicine's health-statistics guide emphasizes that study design matters when making causal claims.

Before students accept a causal headline, ask whether participants were randomly assigned, whether important variables were controlled, and whether another explanation could produce the pattern.

Show why selective analysis creates false confidence

If researchers try many analyses and report only the successful-looking result, the published evidence can become misleading. The National Academies' work on reproducibility and replicability discusses practices such as p-hacking and selective reporting as threats to reliable inference.

For students, the practical rule is straightforward: transparency matters. Ask what was planned, what was measured, what was excluded, how many comparisons were tried, and whether the analysis can be reproduced or replicated.

A worked headline check

Headline: “New study proves a study app raises grades.”

Study information: 38 volunteers used the app for two weeks. Their average quiz score was 3 points higher than the comparison group's. The reported p-value was .08.

A weak response says: “p is bigger than .05, so the app does not work.”

A stronger response says: “This study does not provide strong statistical evidence under the chosen threshold, but it also does not prove the app has no effect. The sample is small, so the study may have limited power. I would want a larger, well-designed study and information about the size and consistency of the effect before making a strong claim.”

That is the kind of reasoning students can use with news headlines, school data, science claims, and later coursework.

Use the routine after the Crash Course p-value sequence

The Crash Course Statistics #12–#23 lesson set contains the probability, confidence-interval, p-value, and statistical-power sequence. In the Complete Full Curriculum, Week 6 adds the student-facing Should We Believe This Claim? lesson so students practice judging realistic claims without difficult computation.

Teachers who want to preview the overall course structure can use the FREE Educator Planning Guide, and the FREE #1 What Is Statistics lesson shows the episode-level classroom format.

The sentence students should remember

A p-value is one piece of evidence. It is not a verdict. A responsible conclusion also considers the research question, study design, sample, effect size, uncertainty, alternative explanations, and practical consequences.

Independent resource notice: K12 Movie Guides is not affiliated with, endorsed by, sponsored by, or authorized by Crash Course, Complexly, YouTube, the American Statistical Association, NCES, the National Academies, or the National Library of Medicine. External organizations are linked for teacher reference and public educational context.

Back to blog