← All Data Visualization with AI modules

Module 4 · Project 2

Association, Regression & Causal Boundaries

Stage 1 of 7 · Continuous probability models

Your goal right now

Understand why probability for a continuous variable is represented by area, not curve height.

A distribution describes how a random variable’s values are spread out or clustered. Discrete variables have separate possible values, like the outcome of a die roll. Continuous variables can take any value within an interval.

Reaction time, distance, and voltage are often modeled as continuous: there are infinitely many possible values between two measurements. A recorded measurement may be rounded, even when the underlying quantity is continuous.

For a continuous probability distribution, probabilities are represented by areas under the probability density function, rather than by the height of the density curve at a particular point.

A probability density function (PDF) describes these models. The area under its curve over a range gives the probability of that range. The total area is \(1\); a single exact value has no width, so \(P(X=c)=0\).

Each shaded region is a probability. The normal and exponential tails extend beyond the displayed window.

Continuous models support statistical inference and hypothesis tests, from laboratory measurements to financial data. Before using one, check that its assumptions fit your data.

QuantegyAI feedback does not submit to Canvas.

Statistics focuscontinuous random variables, density curves, CDFs, normal models, z-scores, percentiles, correlation, covariance, confounding, simple linear regression, residuals, leverage, heteroscedasticity
R/tool focusInline browser R labs (base R + ggplot2, no install needed); optional desktop RStudio/Positron for the capstone; pnorm(), qnorm(), cor(), ggplot2 scatterplots/density curves, lm(), summary(lm()), fitted/residual plots; broom::tidy() on desktop only
Interactive assignmentNormal-model area check, correlation/causation fallacy gallery, residual diagnosis cards, regression diagnostic packet

Student learning outcomes

By the end of Module 4, students can…
  1. Interpret correlation within its linear-only limits and recognize the relationships r cannot see.
  2. Fit a simple regression in R and interpret slope, fitted values, and residuals in the context of the question.
  3. Diagnose residual patterns — nonlinearity, non-constant variance, leverage, and influence — and state which conclusions each pattern breaks.
  4. Detect confounding and Simpson-style reversals by comparing pooled and within-group fits.
  5. Write causal-boundary language that matches what the study design can support.

Interactive reading

Module 4 joins continuous probability models with association, regression, and causal-boundary reasoning. Select True or False before opening each explanation, then check the statistical rule.

Normal model · predict before reveal

A value is converted to \( z = \dfrac{x-\mu}{\sigma} \). Computing a z-score establishes that the data follow a normal distribution. True or false?

False. A z-score measures how many standard deviations the value is above or below the model mean. It does not prove the data are normal; it only translates the value onto the scale of a proposed normal model.
Area · predict before reveal

On a continuous density curve, the probability of exactly one value — \(P(X=70)\), say — equals zero. True or false?

True. For a continuous variable, probability is represented by area over an interval. A single exact point has area zero, so \(P(X=70)=0\) under a continuous model. Meaningful probability statements use ranges, such as \(P(X \le 70)\) or \(P(65 \le X \le 75)\).
Association · predict before reveal

A correlation of \(r = 0.82\) between practice time and exam score is strong enough to establish that more practice time causes higher scores. True or false?

False. Correlation is association, not causation. Confounders, selection effects, measurement issues, and study design must be considered before causal language is justified. The strongest defensible claim may be predictive or associational, not causal.
Residuals · predict before reveal

A fitted regression line has a high \(R^2\), but the residual plot bends in a clear curve. The high \(R^2\) makes the straight-line model valid anyway. True or false?

False. A high \(R^2\) does not erase diagnostic failure. Curved residuals indicate that the linear form may be misspecified, so the defense must revise the model, transform variables, use another model, or limit the claim.
Causal boundary · predict before reveal

An AI assistant writes: “Because response time predicts final score, reducing response time will cause scores to improve.” That sentence has crossed from prediction into intervention, and needs a causal-boundary audit before it can be used. True or false?

True. The assistant moved from prediction to intervention. Observational prediction does not prove that changing response time would cause a score change. The audit should name possible confounding variables, state the study design, and rewrite the claim with a causal boundary.

Continuous model lab

Open the continuous normal-model lab · about 8 minutes

Use a normal model as a checkable approximation: translate values to z-scores, interpret area as probability, and state when the model is only a rough fit.

The continuous normal-model lab needs JavaScript.

Confounding lab

Open the confounding and pooled-correlation lab · about 10 minutes

The confounding lab needs JavaScript.

Open the correlation-versus-causation challenge · about 10 minutes

Residual diagnosis cards

Open the residual-diagnosis challenge · about 10 minutes

The residual diagnosis cards need JavaScript.

AI-output audit

Open the AI-output audit · about 10 minutes

Your task: Evaluate all six proposals and use the feedback to review your decisions.

Lab complete when every proposal has a verdict and the audit table below the cards is filled in — that table is the AI-use log your Project 2 packet asks for.

This AI audit needs JavaScript.

Hands-on visual lab

Open the leverage visual lab · about 8 minutes

Move one observation out of twenty-five and watch the fitted slope follow it. Leverage is a property of the predictor alone, so a point can dominate a model before you have looked at its outcome. Then run the browser-R lab below without leaving QuantegyAI.

This visual lab needs JavaScript.

GeoGebra lab

Open the GeoGebra regression lab · about 10 minutes
Interactive Math with AI / GeoGebra strand

Embedded GeoGebra applet loading…

R lab

You predicted the relationship and its causal boundary. Now verify them computationally. Run the provided students data, inspect the fitted and residual plots, test the track sensitivity, and carry the checked evidence into DV07 and DV08.
Open the runnable browser-R lab · about 15 minutes

Your task Write the prediction, run the code sections in order, and read the printed output.

Lab complete when the scatterplot and the residual plot have both drawn and you have written the interpretation note at the foot of the lab — three sentences: the relationship, the residual check, and the causal boundary.

Keep the calculation simple. Run the provided code and interpret the rounded output. Two sensible decimal places are enough; no assessed question requires hand-calculating a correlation, regression coefficient, z-score, or percentile.
R / Quarto / Shiny workflow

Use `pnorm()`/`qnorm()` to verify normal-model area claims, then fit `lm()`, use `summary(lm())` (or `broom::tidy()` in desktop R), create fitted and residual plots, and write a model-appropriateness note.

Browser R note: this inline lab runs base R plus ggplot2 (installed automatically on first run — one short wait, cached after that). Everything this module requires — including lm(), residual plots, and normal-model checks — runs right here in the browser; broom::tidy() is a desktop convenience.

# Suggested R workflow scaffold
library(ggplot2)
set.seed(42)

# Dataset: study hours and exam scores across four exam tracks.
# All tracks share a base slope of 2 points per study hour, with modest
# per-track slope shifts and intercept shifts — real but borderline
# effects, so the sensitivity check is a judgment call, not a
# foregone "no difference".
n_per <- 25
students <- data.frame(
  track = rep(c("TExES", "SAT", "ACT", "Generalist"), each = n_per),
  study_hours = round(runif(4 * n_per, 2, 20), 1)
)
intercept_shift <- c(TExES = 0, SAT = 3, ACT = -2, Generalist = 4)
slope_shift <- c(TExES = 0, SAT = 0.35, ACT = -0.25, Generalist = 0.5)
mean_score <- 50 + intercept_shift[students$track] +
  (2 + slope_shift[students$track]) * students$study_hours
students$exam_score <- round(mean_score + rnorm(nrow(students), 0, 8), 1)
# 1. A simple normal-model check (use the output; do not calculate by hand)
z <- (78 - 70) / 8
round(c(z = z, area_below_78 = pnorm(78, mean = 70, sd = 8)), 2)

# 2. Fit and inspect the linear model
fit <- lm(exam_score ~ study_hours, data = students)
round(c(correlation = cor(students$study_hours, students$exam_score),
        intercept = coef(fit)[1], slope = coef(fit)[2],
        r_squared = summary(fit)$r.squared), 2)

# 3. Build the two required visuals
students$fitted <- fitted(fit)
students$residual <- resid(fit)
p_scatter <- ggplot(students, aes(study_hours, exam_score)) +
  geom_point() + geom_smooth(method = "lm", se = TRUE) +
  labs(title = "Exam score and study hours", x = "Study hours", y = "Exam score")
p_residual <- ggplot(students, aes(fitted, residual)) +
  geom_hline(yintercept = 0, colour = "grey60") + geom_point() +
  labs(title = "Residual check", x = "Fitted score", y = "Residual")
print(p_scatter)
print(p_residual)

# 4. Sensitivity check: does the study-hours slope differ by exam track?
#    The additive model is the "one common slope" benchmark; `*` adds the
#    interaction so each track gets its own slope. The F-test and the
#    study_hours:track terms decide whether the slopes truly differ.
fit_additive <- lm(exam_score ~ study_hours + track, data = students)
fit_interaction <- lm(exam_score ~ study_hours * track, data = students)
round(summary(fit_interaction)$coefficients, 2)
anova(fit_additive, fit_interaction)

# 5. Write three plain-language sentences:
# relationship, residual/model-fit check, and causal boundary.
# Log any AI assistance and corrections.

Canvas Project 2 evidence: DV07 and DV08

You are not doing separate assignments here.

DV07 and DV08 are the evidence you will package for Project 2. Complete each item once; your saved work carries forward to the final Canvas packet.

Open the DV07 and DV08 Project 2 evidence checklists

Complete both evidence items below. The Project 2 page packages DV07 and DV08 into one Canvas zip.

Your task Open each card, read the worked exemplar, then tick each check once your own work satisfies it and paste your note or link.

Project 2 ready when both cards read complete. Only then does Project 2 let you download Project_2_Canvas_Packet.zip.

Evidence-readiness decision check

The Module 4 submission decision check needs JavaScript.

Improve your work — not graded

You have already done most of the work.

Now package the association evidence, regression diagnostics, R verification, and causal-boundary reasoning you produced above. This is formative evidence review, not another assignment and not a Canvas submission.

  1. Get feedback
  2. AI reports one or two evidence gaps
  3. Fix them
  4. Recheck
  5. Ready for Project 2
Open the Module 4 AI-feedback workspace

Where do I get my data?

Use students. Everything you need is already provided in the R lab. It contains the study-hours, exam-score, and track variables needed for the guided DV07 and DV08 path.

Advanced option: Apply this workflow to your own data

You may use a built-in R dataset or add a CSV under Your data in the R lab. The provided students dataset remains the recommended route.

How does grading work?

The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.

Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit the Project 2 packet to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like.

The inline Module 4 capstone submission and AI grader need JavaScript.

Adaptive reflection

What I know / where I go next

Mastery check loading…

Module 4 record

Your saved Module 4 record needs JavaScript. You can also see it on your skills dashboard report card.

Final Project connection Add this module’s evidence to your Final Project readiness tracker. If your explanation depends on AI, include the prompt, output, verification method, and final correction.