Back to Tara
Tara

Sample unit plan

Scatter Plots, Association, and Correlation

Chapter 6 — a 13-day unit plan built by Tara after a teacher shared her available time and chose to include project and enrichment.

Sections 6.1 · 6.2 · 6.3 · 6.413 days · 60-minute periods
Tara

How this unit came to be

A teacher told Tara she had 13 sixty-minute periods to teach Chapter 6. Tara doesn't simply start planning. She checks the available time against the recommended range for the chapter, evaluates whether the pacing is realistic, and — when there's room to go deeper — offers the teacher a choice before building anything.

Tara determined that 13 days falls squarely within the recommended range for Chapter 6 and offered the teacher a choice: Include project and enrichment or Proceed as requested.

Tara's suggestion

You've got some wonderful flexibility built into this unit plan, which gives you a real opportunity to deepen student understanding beyond the mechanics! Rather than filling that time with additional practice problems, consider having students investigate a dataset that genuinely interests them—whether that's analyzing the relationship between social media usage and sleep, exploring sports statistics, or examining environmental data—where they'll grapple with why correlation matters and how to interpret what they find. This kind of authentic exploration helps students see scatter plots and correlation coefficients not as abstract procedures, but as powerful tools for understanding real patterns in the world around them. Your students will leave this chapter with both solid technical skills and a deeper appreciation for the story data can tell.

The teacher chose Include project and enrichment. What follows is the complete unit plan Tara created.

Pacing judgment — Level 3: Within Range

13 days (780 minutes / 13 hours) falls squarely within the recommended range for Chapter 6 (5–13 hours). This pacing allows for meaningful instruction on all four core concepts, technology skill development, enrichment tasks, and a substantial investigation that runs throughout the unit.

Unit Overview

Chapter 6: Scatter Plots, Association, and Correlation

Standard Introductory Statistics · 13 days × 60 minutes = 780 minutes

Big ideas: Relationships between two quantitative variables can be visualized, measured, and modeled. A correlation coefficient quantifies the strength and direction of a linear relationship, but correlation does not imply causation. Regression models allow us to make predictions, but all models have limitations — understanding residuals and model fit helps us know when a prediction is trustworthy.

Essential questions

  • What does a scatterplot reveal about the relationship between two variables that summary statistics alone cannot?
  • How do we distinguish between a strong association and a weak one, and what does the correlation coefficient actually measure?
  • When is a linear model appropriate, and when does it fail to capture the real pattern in the data?
  • What does a residual tell us about how well a model fits a particular observation?
  • Why is "correlation does not imply causation" one of the most important ideas in statistics?
DayFocusGuiding QuestionWhat Students Will DoResourcesNotes
Day 1Introduction to scatterplots and associationWhat does a scatterplot show that a table of numbers cannot?Warm-up: interpret a real-world scatterplot (e.g., study hours vs. test scores). Mini-lesson on creating scatterplots and describing direction, form, and strength. Guided practice: students create a scatterplot from a dataset and describe the association in writing. Exit ticket: sketch and describe a scatterplot for a new variable pair.Toolkit: 6.1 Slides; 6.1 Guided NotesEmphasize precise language: direction (positive/negative), form (linear/curved), strength (weak/moderate/strong). Do not introduce correlation coefficient yet.
Day 2Describing association and recognizing limitationsHow do we know whether an association is meaningful, and what can a scatterplot NOT tell us?Mini-lesson on outliers, clusters, and when linear form breaks down. Enrichment Task 1 (Urban Infrastructure and Public Safety): students analyze a pre-made scatterplot and write a brief analysis for a planning committee, including a discussion of correlation vs. causation and limitations.Toolkit: 6.1 Slides; 6.1 Guided NotesUse Task 1 to anchor the correlation-vs-causation discussion. This is a critical misconception to address early.
Day 3Launching the Education vs. Income investigationWhat does the relationship between education level and income look like, and how strong is it?Students receive the Education vs. Income dataset (50 states). They create a scatterplot, describe the association, and make a prediction about the relationship before calculating any statistics. This is the first phase of the ongoing project spine.Toolkit: 6.1 Slides; 6.1 Guided Notes; Dataset: Education vs. Income (50 states)This investigation runs through the unit. Students will return to it after each new concept. Frame it as a running investigation, not a one-time task.
Day 4Introduction to the correlation coefficientHow do we measure the strength and direction of a linear relationship with a single number?Warm-up: compare two scatterplots with different strengths; estimate which has a stronger correlation. Mini-lesson on r: what it measures, its range (−1 to 1), interpretation of sign and magnitude. Guided practice: interpret given correlation coefficients in context. Exit ticket: match correlation coefficients to scatterplots.Toolkit: 6.2 Slides; 6.2 Guided Notes; Video: Scatter Plots, Association, and CorrelationThe video covers both scatterplots and correlation; use it to reinforce the connection between visual and numerical measures.
Day 5Calculating correlation and extending the investigationHow do we calculate r, and what does it tell us about the Education vs. Income relationship?Mini-lesson: formula for r (conceptual understanding, not hand calculation). Introduction to Google Sheets for calculating correlation. Guided practice: students use Google Sheets to calculate r for the Education vs. Income dataset. They interpret the result and compare it to their earlier visual prediction.Toolkit: 6.2 Slides; 6.2 Guided Notes; Technology Resource: Calculating the Correlation Coefficient and Best Fit Regression Line Using Google SheetsTeach the Google Sheets skill here, as students need it immediately for the investigation. This is the point where technology becomes essential.
Day 6Interpreting correlation in context and discussing causationWhat does a strong correlation between two variables really mean, and what does it NOT mean?Enrichment Task 1 (Wildlife Observation and Environmental Change): students evaluate two correlations (staffing vs. birds, rainfall vs. birds) and discuss which is more meaningful and why. Class discussion on spurious correlation, confounding variables, and the limits of observational data.Toolkit: 6.2 Slides; 6.2 Guided NotesThis day consolidates the critical idea that correlation ≠ causation. Use Task 1 to make the discussion concrete.
Day 7Introduction to linear regressionHow do we use the relationship between two variables to make predictions?Warm-up: given a scatterplot with a line drawn through it, estimate the y-value for a given x-value. Mini-lesson on the line of best fit, the regression equation (y = a + bx), and interpretation of slope and y-intercept in context. Guided practice: interpret regression output and make predictions. Exit ticket: predict a value and explain what the slope means.Toolkit: 6.3 Slides; 6.3 Guided Notes; Video: Linear RegressionThe video covers regression output and interpretation. Use it to show how the regression line relates to the scatterplot.
Day 8Fitting a regression model and extending the investigationHow do we fit a regression line to the Education vs. Income data, and what does the equation tell us?Mini-lesson: using Google Sheets or TI-84 to fit a regression model. Students fit the regression line to the Education vs. Income dataset, write the equation, and interpret slope and y-intercept in context. They make a prediction for a state with ~30% bachelor's degree rate and compare to a real state.Toolkit: 6.3 Slides; 6.3 Guided Notes; Technology Resource: Calculating the Correlation Coefficient and Best Fit Regression Line Using Google SheetsStudents apply regression immediately to their ongoing investigation. This is where the project becomes predictive.
Day 9Understanding r² and model strengthWhat proportion of the variation in the response variable is explained by the explanatory variable?Mini-lesson on r² (coefficient of determination): what it measures, how to interpret it, and what it does NOT tell us. Guided practice: interpret r² values in context and discuss what factors might explain the remaining variation. Exit ticket: explain what r² = 0.81 means in a specific context.Toolkit: 6.3 Slides; 6.3 Guided Notesr² is often misunderstood as "percentage of data points on the line." Clarify that it measures explained variance. Use the Education vs. Income data to illustrate.
Day 10Introduction to residuals and model fitHow do we know whether a linear model is appropriate, and how well does it fit each individual observation?Warm-up: given a prediction and an actual value, calculate the residual. Mini-lesson on residuals: what they measure, how to calculate them, and how to interpret them in context. Guided practice: calculate residuals for several observations and discuss what they reveal about model fit. Exit ticket: interpret a residual in context.Toolkit: 6.4 Slides; 6.4 Guided NotesResiduals are often confusing. Use concrete examples: if the model predicts 85 and the actual is 89, the residual is +4 (the model underpredicted).
Day 11Evaluating residuals and model appropriatenessWhen is a linear model appropriate, and when does it fail?Mini-lesson: examining residual plots to assess linearity and constant variance. Enrichment Task 1 (Predicting Influenza Cases): students evaluate whether a linear model is appropriate by examining residuals and comparing actual vs. predicted values. Discussion: when does a linear model break down?Toolkit: 6.4 Slides; 6.4 Guided NotesThe flu data is curved, not linear — a strong example of when linear regression fails. Use it to show that a high r² does not guarantee a good model.
Day 12Completing the investigation and synthesisHow do all four concepts—scatterplots, correlation, regression, and residuals—work together to help us understand and model relationships?Students return to the Education vs. Income dataset. They calculate residuals for 3 selected states (one above the line, one below, one near the line), interpret what the residuals reveal, and write a brief reflection on what the model captures and what it leaves out. They suggest two additional explanatory variables that might improve the model.Toolkit: 6.1–6.4 Slides and Guided Notes (review as needed); Dataset: Education vs. Income (50 states)This is the consolidation day for the ongoing investigation. Students synthesize all four concepts. Allow time for writing and reflection.
Day 13Capstone analysis and reflectionHow do we apply all the tools from this unit to investigate a real-world question?Students complete a brief capstone analysis: select a new dataset (or use a provided alternative such as Inactivity vs. Obesity), create a scatterplot, calculate r, fit a regression line, interpret slope and r², make a prediction, calculate residuals, and write a short report interpreting findings in context and acknowledging limitations.Toolkit: 6.1–6.4 Slides and Guided Notes (as reference); Datasets: Inactivity vs. Obesity or Graduation Rates vs. Teen Pregnancy (student choice or teacher-selected)This day is reserved for completion, synthesis, and reflection. Students should have most of the analysis done; use this time for writing, discussion, and consolidation.

Key Learning Outcomes

By the end of this unit, students can:

  1. Create and interpret scatterplots to visualize the relationship between two quantitative variables and describe the direction, form, and strength of association using precise language.
  2. Calculate and interpret the correlation coefficient (r) in context, understanding that it measures the strength and direction of a linear relationship and that r ≠ causation.
  3. Fit a linear regression model using technology, interpret the slope and y-intercept in context, and use the model to make predictions.
  4. Interpret r² (coefficient of determination) as the proportion of variation in the response variable explained by the explanatory variable, and recognize what the remaining variation might represent.
  5. Calculate and interpret residuals to evaluate how well a linear model fits individual observations and to assess whether a linear model is appropriate for a dataset.
  6. Recognize the limitations of linear models and explain why correlation does not imply causation, why outliers matter, and when a linear model may be misleading.

Enrichment tasks (project spine)

6.1 — Urban Infrastructure and Public Safety (Day 2)

Students analyze a pre-made scatterplot and write a brief analysis for a planning committee, including a discussion of correlation vs. causation and limitations.

6.1 — Workforce Activity and Cognitive Performance (Days 1–2)

Students select an appropriate visual representation, describe the relationship, compare to expectations, and note features affecting interpretation.

6.2 — Wildlife Observation and Environmental Change (Day 6)

Students evaluate two correlations (staffing vs. birds, rainfall vs. birds) and discuss which is more meaningful and why — spurious correlation and confounding variables.

6.2 — Employee Training and Workplace Performance (Days 4–6)

Students analyze training hours vs. performance scores, calculate and interpret r, and discuss limitations.

6.3 — Housing Affordability and Homelessness (Days 7–8)

Students fit their own regression model using technology, interpret slope and intercept, make predictions, and discuss reliability.

6.4 — Predicting Influenza Cases (Day 11)

Students evaluate whether a linear model is appropriate by examining residuals. The flu data is curved — a strong example of when linear regression fails.

6.4 — Urban Design and Public Space Planning (Days 10–11)

Students fit a regression model, predict a value, then evaluate a new observation as a residual and discuss what it suggests about applying the model in a different context.

Common misconceptions & how to address them

"A strong correlation means one variable causes the other."

Address on Day 2 with concrete examples (ice cream sales and drowning deaths; shoe size and reading ability). Return on Day 6 with Enrichment Task 1 (Wildlife Observation). Reinforce throughout: every time a correlation is discussed, ask "Does this mean one causes the other? Why or why not?"

"The correlation coefficient tells us the percentage of data points on the line."

Clarify on Day 4: r measures the strength of the linear relationship, not the percentage of points fitting perfectly. Use r² on Day 9 to discuss explained variance.

"A residual is just the difference between two numbers."

Help students understand that residuals reveal how well the model fits each observation. Use concrete examples: "The model predicted 85, but the actual value was 89. The residual is +4, which means the model underpredicted by 4 units."

"A linear regression model is always appropriate."

Show on Day 11 with Enrichment Task 1 (Predicting Influenza Cases) that the flu data is curved, not linear. A high r² does not guarantee a good model.

Education vs. Income Investigation — Ongoing Spine

This investigation integrates all four core concepts throughout the unit. Students encounter each concept when they need it for the investigation and apply it immediately, creating a coherent narrative of how statistical tools work together.

  • Day 3: Launch — create scatterplot, describe association, make predictions
  • Day 5: Calculate correlation using Google Sheets, interpret r
  • Day 8: Fit regression line, interpret slope and y-intercept, make predictions
  • Day 12: Calculate residuals, interpret what they reveal, write synthesis reflection

Capstone Dataset Analysis — Days 12–13

Students demonstrate independent mastery by selecting a dataset (Inactivity vs. Obesity or Graduation Rates vs. Teen Pregnancy), conducting a complete analysis, and communicating findings in a brief professional report that acknowledges limitations and contextual factors.

Videos

Scatter Plots, Association, and Correlation (Day 4) · Linear Regression (Day 7) — both available in the Simpler Math video library.

Guided Notes & Slides

Sections 6.1–6.4 slides and guided notes available in the Teacher Toolkit. Use on Days 1–3 (6.1), 4–6 (6.2), 7–9 (6.3), and 10–12 (6.4).

Datasets

Education vs. Income (50 states) — ongoing spine (Days 3, 5, 8, 12). Inactivity vs. Obesity — capstone alternative. Graduation Rates vs. Teen Pregnancy — outlier focus (Day 6 or standalone).

Technology Resource

Calculating the Correlation Coefficient and Best Fit Regression Line Using Google Sheets — teach on Day 5 when students first need to calculate correlation.

Vetted Practice

Use the Tara worksheet generator for additional practice on scatterplot interpretation (Days 1–2), correlation coefficients (Days 4–6), regression equations (Days 7–9), and residuals (Days 10–11).

Formative

  • Exit tickets on Days 1, 4, 7, and 10 — each targeting the key concept introduced that day.
  • Warm-up activities on Days 1, 4, 7, and 10.
  • Investigation checkpoints on Days 3, 5, 8, and 12 — monitor scatterplot creation, correlation calculation, regression fitting, and residual interpretation.

Summative — Capstone Analysis (Days 12–13)

Students complete a brief written report that includes:

  1. A scatterplot with appropriate labels and a fitted regression line.
  2. Calculated correlation coefficient (r) with interpretation in context.
  3. Regression equation with interpretation of slope and y-intercept.
  4. At least one prediction with explanation of whether it is trustworthy.
  5. Calculation and interpretation of at least one residual.
  6. A reflection on what the model captures, what it leaves out, and what limitations exist.
  7. Discussion of whether the relationship is causal or merely correlational, with evidence.

Concepts that deserve careful time

  • Correlation vs. causation (Days 2, 6): Return to it repeatedly. Use concrete, memorable examples. Do not rush past this.
  • Interpretation of r² (Day 9): Students often confuse r² with the percentage of data points on the line. Clarify that r² measures explained variance.
  • Residuals and model fit (Days 10–11): Use concrete numerical examples before asking for interpretation. Show how residual plots reveal whether linearity is appropriate.
  • Technology skill (Day 5): Allocate sufficient time for students to become comfortable with Google Sheets. Do not assume all students will pick it up immediately.

If you have less time

  • Combine Days 1–2: introduce scatterplots and immediately address correlation vs. causation with Enrichment Task 1.
  • Combine Days 4–5: introduce r, then immediately teach Google Sheets and calculate correlation.
  • Shorten Days 10–11: teach residuals on Day 10, then use Day 11 for practice and one enrichment task.
  • Do not skip: the correlation-vs-causation discussion, the ongoing investigation, or residuals and model fit.

Final note on the capstone

The capstone analysis is not a test or a separate project. It is the culmination of the ongoing investigation plus a brief independent analysis. By Day 12, students should have most of the work done on the Education vs. Income dataset. Day 13 is reserved for completing that analysis, writing the reflection, and possibly conducting a brief analysis of a second dataset. The focus should be on synthesis, communication, and reflection — not on rushing through new content.

Ready to plan your own unit? Tara can build a complete unit plan for any chapter in the Statistics curriculum.

Try Tara's Planning Mode