Module 3 — Discrete Distributions & Group Comparisons
Same average. Different story. Explore probability and spread with one lab of your choice.
Mastery earns up to 30 points and each completed lab earns 40. The module display stops at 100, but extra points count toward the shared Project 1 grade. Earlier earned credit is preserved.
Start Module 3 →Canvas Project 1 — Modules 1, 2, and 3
Module 3 is part of Project 1. Combine DV05 (Distribution Model Defense) and DV06 (Group Comparison Defense) with DV01–DV04 from Modules 1 and 2.
Download Project_1_Canvas_Packet.zip and upload it to Canvas Project 1. Your running grade adds mastery and lab points across Modules 1–3. AI feedback remains separate from submission.
Open Project 1 instructions and packetStudent learning outcomes
- Identify a discrete random variable and its support.
- Validate a probability table and PMF.
- Construct and interpret a discrete CDF.
- Compute and interpret expected value, variance, and standard deviation.
- Defend or reject a Binomial, Geometric, or Poisson model using its data-generating assumptions and empirical evidence.
- Test whether a histogram feature persists across two defensible bin widths.
- Compare groups using shape, center, spread, overlap, and outliers.
- Compute and interpret a standardized effect size without claiming individual separation or causation.
- Explain what a box plot encodes and which distribution shapes it can hide.
- Use AI to draft, R to verify, and their own reasoning to correct and defend a conclusion.
Interactive reading
Five prompts covering the four things a distribution can tell you — shape, modality, spread, and extremes — plus the group-overlap question this module turns on. Select True or False before opening each explanation; the point is to find out which of these you can already read off the numbers without a chart.
Course completion times have a mean of 14.2 weeks and a median of 11.5. The gap is a clue that high values may be pulling the mean upward, but you should inspect the distribution before naming right skew or choosing the center to emphasize. True or false?
A histogram of exam scores shows a single peak at a bin width of 12 points and two peaks at a bin width of 4. The narrower bins resolve more detail, so the two-peak reading is the one that belongs in the report. True or false?
Two campuses report the same mean score of 68. One has \( \mathrm{SD} = 4 \), the other \( \mathrm{SD} = 16 \). Because the means match, the two campuses can be described to a reader the same way. True or false?
A single value of 240 appears in a set otherwise ranging from 30 to 95. It sits so far outside the rest of the data that deleting it is the defensible first move. True or false?
Two groups can have clearly different means and still overlap heavily in a visualization, because a difference in centres says nothing about the variation within each group. True or false?
Interactive reading: PMF and CDF worked example
Open the five-step PMF/CDF worked example · about 12 minutes
Many applied statistics and biostatistics sequences teach “the probability distribution of a discrete random variable” without ever naming the probability mass function or the cumulative distribution function, so the labels can be new even when the idea is not. Work this example once, by hand, before the reactive table lab — the lab uses the same five probabilities, so every number you compute here should reappear on screen.
Let \( X \) be the number of retakes a randomly chosen student needs after the first sitting, recorded as \( 0, 1, 2, 3, 4 \) (where 0 means the student passed on the first sitting and needs no retake). A probability table assigns one probability to each value:
| \( x \) | 0 | 1 | 2 | 3 | 4 | Total |
|---|---|---|---|---|---|---|
| \( P(X = x) \) | 0.08 | 0.20 | 0.34 | 0.26 | 0.12 | 1.00 |
That row is the PMF. Formally, the probability mass function is \( p(x) = P(X = x) \), and a table is a valid PMF exactly when \( 0 \le p(x) \le 1 \) for every value and \( \sum_x p(x) = 1 \). Here \( 0.08 + 0.20 + 0.34 + 0.26 + 0.12 = 1.00 \), so the table is valid. If the total were \( 0.97 \) or \( 1.05 \), nothing downstream — expected value, spread, or any chart — would be interpretable until the table was repaired.
If you worked Module 2's raw-tally example — the observed attempts \( 1, 1, 2, 3, 3, 3 \) became a table by tallying each value and dividing by \( n = 6 \) — you have already built a probability table the long way: count, then divide. The table above hands you the probabilities directly, and the two validity checks are the same either way. In Module 2 the entries came from data; here they describe a model of the process. The reactive probability table lab edits this same table live.
The cumulative distribution function answers “what is the probability of at most this many retakes?”: \( F(x) = P(X \le x) = \sum_{k \le x} p(k) \). You get it by running a cumulative sum down the PMF:
| \( x \) | \( p(x) = P(X = x) \) | Running sum | \( F(x) = P(X \le x) \) |
|---|---|---|---|
| 0 | 0.08 | \( 0.08 \) | 0.08 |
| 1 | 0.20 | \( 0.08 + 0.20 \) | 0.28 |
| 2 | 0.34 | \( 0.28 + 0.34 \) | 0.62 |
| 3 | 0.26 | \( 0.62 + 0.26 \) | 0.88 |
| 4 | 0.12 | \( 0.88 + 0.12 \) | 1.00 |
Three checks that the CDF is right: it never decreases, it ends at exactly 1, and each PMF value can be recovered as a difference of neighbours, \( p(x) = F(x) - F(x-1) \). For example \( P(X = 2) = F(2) - F(1) = 0.62 - 0.28 = 0.34 \). For a discrete variable the CDF is a step function — flat between the allowed values, jumping by \( p(x) \) at each one — which is what the right-hand panel of the lab draws.
Reading probabilities from the CDF instead of adding PMF cells:
- \( P(X \le 2) = F(2) = 0.62 \) — 62% of students need at most two retakes.
- \( P(X > 2) = 1 - F(2) = 0.38 \).
- \( P(1 \le X \le 3) = F(3) - F(0) = 0.88 - 0.08 = 0.80 \). Subtract \( F \) at the value just below the lower limit, because \( F(1) \) would remove \( P(X = 1) \) itself. Unpacked, the difference keeps exactly the interval's PMF cells — \( P(X=1) + P(X=2) + P(X=3) = 0.20 + 0.34 + 0.26 = 0.80 \) — so the subtraction rule and adding the cells are the same computation seen two ways.
The expected value is the probability-weighted average of the values, \( E(X) = \sum_x x\,p(x) \):
\[ E(X) = 0(0.08) + 1(0.20) + 2(0.34) + 3(0.26) + 4(0.12) = 0.20 + 0.68 + 0.78 + 0.48 = 2.14 \]
Over many students the mean number of retakes is 2.14. Note that \( E(X) = 2.14 \) is not a value \( X \) can take — an expected value is a long-run average, not a prediction for one student.
The variance uses the shortcut \( \operatorname{Var}(X) = E(X^2) - [E(X)]^2 \). First \( E(X^2) = \sum_x x^2 p(x) \):
\[ E(X^2) = 0(0.08) + 1(0.20) + 4(0.34) + 9(0.26) + 16(0.12) = 0.20 + 1.36 + 2.34 + 1.92 = 5.82 \]
\[ \operatorname{Var}(X) = 5.82 - 2.14^2 = 5.82 - 4.5796 = 1.2404, \qquad \mathrm{SD}(X) = \sqrt{1.2404} \approx 1.114 \]
Interpretation: over many students the mean number of retakes is 2.14, with a standard deviation of about 1.1 retakes. Open the reactive table lab without changing anything and confirm the rounded story: total probability is 1, the mean is about 2.14 retakes, the variance is about 1.24, and the standard deviation is about 1.11.
How Spread Out Are the Data? Understanding Variance in R. By the end of this lab you should be able to: explain variance in your own words; calculate variance using R; compare two groups using variance; explain why a larger variance means more spread; and connect variance to standard deviation.
Part 1 — Create two data sets. The starter defines two groups with the same mean: group_a <- c(9, 10, 10, 11, 10) and group_b <- c(2, 6, 10, 14, 18). Run the starter to see each group’s mean printed in real R. Part 2 — Look at the data visually. The same starter loads ggplot2 and draws a dot plot of both groups — one dot per value. Which group looks more spread out? Part 3 — Calculate variance. var(group_a) and var(group_b) print each group’s variance, and the box is yours from there: create your own data set and calculate its variance. The code box below is editable — change any number, add values, or invent your own data set, then press Run and watch the printed values and the dot plot update instantly.
Module 3 asks you to defend or reject a model, which only makes sense once you separate two objects that both get drawn as bar charts:
| Empirical distribution | Theoretical model | |
|---|---|---|
| Where it comes from | Your data. \( \hat{p}(x) = \dfrac{\text{count of } x}{n} \) — relative frequencies from the sample. | A formula chosen from the data-generating situation, e.g. Poisson: \( P(X = k) = \dfrac{e^{-\lambda}\lambda^k}{k!} \), or Binomial \( (n, p) \), or Geometric \( (p) \). |
| What it describes | This sample of \( n \) students. A new sample gives different proportions. | The assumed population process. It does not change when you resample. |
| Free parameters | None — every proportion is read off the data. | One or two (\( \lambda \), or \( n \) and \( p \)), usually estimated from the data, e.g. \( \hat{\lambda} = \bar{x} \). |
| In R | table(attempts) / length(attempts), cumsum(), ecdf() | dpois(), ppois(), dbinom(), pbinom(), dgeom() |
| What can go wrong | Small \( n \): proportions are noisy and rare values may not appear at all. | Wrong assumptions: dependence, a rate that changes across students, or more spread than the model allows (over-dispersion). |
Worked comparison. Suppose 40 students at one campus needed the following retakes: 0 retakes for 6 students, 1 for 11, 2 for 10, 3 for 8, 4 for 3, and exactly 5 for 2. The empirical table divides each count by 40; the sample mean is \( \bar{x} = 77/40 = 1.925 \), so a Poisson model with \( \hat{\lambda} = 1.925 \) (computed with dpois(0:5, 1.925)) gives:
| \( x \) | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| Count | 6 | 11 | 10 | 8 | 3 | 2 |
| Empirical \( \hat{p}(x) \) | 0.150 | 0.275 | 0.250 | 0.200 | 0.075 | 0.050 |
| Poisson(1.925) \( p(x) \) | 0.146 | 0.281 | 0.270 | 0.173 | 0.083 | 0.032 |
Across the displayed exact counts, the two rows disagree by at most about 0.03. The fitted Poisson model also assigns about 0.014 probability to counts above 5; the observed sample has none. A single student changes an empirical cell by 0.025 when \( n = 40 \), so the displayed differences are consistent with ordinary sampling variation. Thus, “the Poisson model is consistent with these data” is defensible — not a proof that retake counts are Poisson. A model is challenged when the empirical row shows a feature the model cannot readily produce: a second peak, far more zeros than \( e^{-\hat{\lambda}} \) allows, or variance far larger than the mean. For a Poisson model, \( \operatorname{Var}(X) = \lambda \), so \( s^2 \gg \bar{x} \) is an overdispersion warning. Say which checks you used in your Project 1 defense.
# PMF -> CDF -> E(X) -> Var(X), from a probability table
x <- 0:4
p <- c(0.08, 0.20, 0.34, 0.26, 0.12)
sum(p) # validity check: must print 1
cdf <- cumsum(p); cdf # F(x): 0.08 0.28 0.62 0.88 1.00
ex <- sum(x * p); ex # E(X) = 2.14
vx <- sum(x^2 * p) - ex^2 # Var(X) = 1.2404
c(variance = vx, sd = sqrt(vx))
# Empirical table versus a fitted Poisson model
attempts <- rep(0:5, times = c(6, 11, 10, 8, 3, 2)) # the 40-student sample above
n <- length(attempts)
empirical <- as.numeric(table(factor(attempts, levels = 0:5))) / n
lambda_hat <- mean(attempts); lambda_hat # 1.925
model <- dpois(0:5, lambda = lambda_hat)
round(rbind(x = 0:5, empirical = empirical, poisson = model), 3)
ppois(5, lambda = lambda_hat, lower.tail = FALSE) # model probability above 5: about 0.014
c(sample_mean = mean(attempts), sample_var = var(attempts)) # Poisson needs these close
Every line prints something you can quote in the “Discrete distribution evidence” field of the graded assignment: the validity total, the CDF, \( E(X) \), \( \operatorname{Var}(X) \), the empirical-versus-model table, and the mean-versus-variance check.
Reactive probability table lab
Open the reactive PMF/CDF table · about 8 minutes
Build a valid discrete distribution before comparing visual shape. Adjust the probabilities and watch the total, expected value, variance, PMF, and CDF update immediately.
If the interactive does not load or is not accessible with your device or assistive technology, use the worked two-toss table. Verify that probabilities are between \(0\) and \(1\), add them, compute the cumulative column, then compute \(E(X)\), \(\operatorname{Var}(X)\), and SD. Equivalent R calculations or written work are accepted in the Canvas packet.
Distribution shape lab
Open the distribution-shape comparison · about 8 minutes
Use the static summaries and descriptions on this page. For each campus, record center, spread, skew or symmetry, peaks, gaps, tails, outliers, and overlap. State which conclusion a box plot supports and which shape details require a histogram, density, violin, or dot plot.
Effect-size lab
Open the effect-size interpretation challenge · about 8 minutes
Use \(d=(\bar{x}_2-\bar{x}_1)/s_{\text{pooled}}\). Show the raw difference, divide by the pooled SD, interpret the standardized magnitude in context, describe overlap, and state that the comparison does not establish causation or individual separation.
AI-output audit
Open the AI-output audit · about 10 minutes
Write the AI claim, the R output or authoritative course evidence used to check it, and the correction you made. Your own explanation is required; an AI response is not evidence by itself.
Hands-on visual lab
Open the histogram bin-width explorer · about 8 minutes
These scores came from two distinct populations. Bin width is a parameter you choose, and it decides whether the reader can see that at all — including whether the second population survives the chart. Then run the browser-R lab below without leaving QuantegyAI.
Create or inspect the same distribution at two defensible bin widths. Record both widths, describe the mode, gap, skew, or subgroup pattern at each width, and state whether the structure persists. If the conclusion changes, label it uncertain. Screenshots or equivalent written descriptions are accepted.
R lab
students data, compare the model with the observed distribution, and carry the checked evidence into DV05 and DV06.Verify the three discrete models in browser R · about 10 minutes
Recompute the binomial, geometric, and Poisson diagrams from the Binomial, geometric, or Poisson? page in real R. Compare the printed table with the chart bars, then defend which model fits a scenario and why.
Browser R note: this inline lab runs base R in the browser — no install needed. Read the printed table against the three diagrams on that page.
Open the runnable browser-R lab · about 15 minutes
The lab is complete as provided. Press Run, read the three printed lines — center, spread, overlap — then answer the True/False statements under the output.
Browser R note: this inline lab runs base R in the browser — no install, no download. For the graded DV05 and DV06 evidence steps, use the Canvas evidence section above the record; its checklists say exactly what each packet needs.
Preserved Project 1 evidence: DV05 and DV06
DV05 and DV06 are the Module 3 evidence included in Canvas Project 1. Complete each item once; combine it with Modules 1 and 2 in the same packet.
Open the DV05 and DV06 Canvas evidence checklists
Complete DV05 and DV06 for Canvas Project 1. Modules 1, 2, and 3 belong in one Project_1_Canvas_Packet.zip.
DV05 minimum evidence and rubric
- Valid PMF and validity check
- CDF with a probability interpreted in words
- Expected value, variance, and SD with context
- Empirical-versus-theoretical model check
- Reproducible R evidence
- A specific limitation
Strong: correct evidence, interpretation, and a defensible model claim. Developing: mostly correct calculations but incomplete interpretation or checking. Insufficient: invalid/missing evidence or an unsupported model claim.
DV06 minimum evidence and rubric
- Named groups, quantitative outcome, target contrast, and each group’s sample size
- Center, spread, and shape comparison
- Overlap and unusual-observation statement
- One appropriate visual with its relevant display choices stated
- Signed raw contrast and effect-size interpretation in context
- Descriptive or inferential status stated explicitly
- For population claims, an uncertainty interval and its assumptions
- Reproducible R evidence
- A limitation and a statement avoiding individual or causal overclaiming
Strong: integrates distribution evidence, magnitude, overlap, and limits. Developing: includes evidence but relies too heavily on means or leaves interpretation incomplete. Insufficient: reports numbers/charts without a defensible comparison.
Reveal generic strong and weak examples after drafting
Stronger: “South’s mean is higher, but the standardized difference is modest and the distributions overlap; this comparison does not establish a campus effect.”
Needs revision: “South is better because its mean is higher.”
Evidence names: use DV05_distribution-model_[your-name] and DV06_group-comparison_[your-name]. Existing saved drafts and uploaded files remain valid. For Project 1, use the DV05 and DV06 checklists on this page. These are evidence identifiers within Module 3, not links to another numbered module or project.
Evidence-readiness decision check
The Module 3 submission decision check needs JavaScript.
Module 3 AI-feedback workspace
Now package the distribution evidence, group-comparison evidence, R verification, and AI-audit reasoning you produced above. This is formative evidence review, not another assignment and not a Canvas submission: requesting AI feedback here does not submit any project to Canvas. Project 1 combines DV01–DV06 from Modules 1–3 in one packet zip.
Open the shared Module 3 AI-feedback workspace
Where do I get my data?
Use students. Everything you need is already provided in the R lab. It contains the discrete counts and campus groups needed for the guided DV05 and DV06 path.
Advanced option: Apply this workflow to your own data
You may use a built-in R dataset or add a CSV under Your data in the R lab. The provided students dataset remains the recommended route.
How does grading work?
The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.
Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit the corresponding project packet to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like.
Required AI verification record
AI drafts the analysis; R checks it; you explain why the result holds. Complete this table for every material AI-assisted claim.
| AI claim | How I checked it | Correction made |
|---|---|---|
Reveal a generic example after completing your row
AI claim: “The Poisson model fits.” Check: compared the empirical PMF with dpois() and checked mean versus variance. Correction: “The Poisson model is consistent with these data; the comparison does not prove independence or a stable rate.”
If this form is unavailable, complete the DV05 and DV06 minimum-evidence checklists, the AI verification record, and the corresponding Canvas packet. Equivalent calculations, screenshots, R output, and written interpretations are accepted. AI feedback is formative; Dr. Roberts’s Canvas review determines the grade.
Adaptive reflection
Mastery check loading…
Module 3 record
Your saved Module 3 record needs JavaScript. You can also see it on your skills dashboard report card.