← All Data Visualization with AI modules

Module 3 — Discrete Distributions & Group Comparisons

QuantegyAI Module 3 · Canvas Project 1 · Due the date on Canvas

Same average. Different story. Explore probability and spread with one lab of your choice.

A short lesson, one lab, then your evidence and mastery check.

Mastery earns up to 30 points and each completed lab earns 40. The module display stops at 100, but extra points count toward the shared Project 1 grade. Earlier earned credit is preserved.

Start Module 3 →
Canvas Project 1 — Modules 1, 2, and 3

Module 3 is part of Project 1. Combine DV05 (Distribution Model Defense) and DV06 (Group Comparison Defense) with DV01–DV04 from Modules 1 and 2.

Download Project_1_Canvas_Packet.zip and upload it to Canvas Project 1. Your running grade adds mastery and lab points across Modules 1–3. AI feedback remains separate from submission.

Open Project 1 instructions and packet

Student learning outcomes

By the end of Module 3, students can…
  1. Identify a discrete random variable and its support.
  2. Validate a probability table and PMF.
  3. Construct and interpret a discrete CDF.
  4. Compute and interpret expected value, variance, and standard deviation.
  5. Defend or reject a Binomial, Geometric, or Poisson model using its data-generating assumptions and empirical evidence.
  6. Test whether a histogram feature persists across two defensible bin widths.
  7. Compare groups using shape, center, spread, overlap, and outliers.
  8. Compute and interpret a standardized effect size without claiming individual separation or causation.
  9. Explain what a box plot encodes and which distribution shapes it can hide.
  10. Use AI to draft, R to verify, and their own reasoning to correct and defend a conclusion.

Interactive reading

Five prompts covering the four things a distribution can tell you — shape, modality, spread, and extremes — plus the group-overlap question this module turns on. Select True or False before opening each explanation; the point is to find out which of these you can already read off the numbers without a chart.

Skew · predict before reveal

Course completion times have a mean of 14.2 weeks and a median of 11.5. The gap is a clue that high values may be pulling the mean upward, but you should inspect the distribution before naming right skew or choosing the center to emphasize. True or false?

True. A mean above the median is a clue that high values may be pulling the mean upward, but the ordering alone does not establish right skew. Inspect the distribution and influential observations first. If a plot confirms strong skew and the goal is a typical individual outcome, lead with the median and IQR; if the goal concerns totals or an additive expectation, the mean may be the relevant estimand.
Modality · predict before reveal

A histogram of exam scores shows a single peak at a bin width of 12 points and two peaks at a bin width of 4. The narrower bins resolve more detail, so the two-peak reading is the one that belongs in the report. True or false?

False. Bin width is a parameter you choose, so a mode that appears at one width and disappears at another is partly a claim about the smoothing — neither view belongs in the report on its own. Show the distribution at more than one width and state whether the second mode survives. If it does, go looking for the subgroup that produces it — bimodality may indicate populations or processes that were pooled, which is exactly what the bin-width explorer below is built from.
Spread · predict before reveal

Two campuses report the same mean score of 68. One has \( \mathrm{SD} = 4 \), the other \( \mathrm{SD} = 16 \). Because the means match, the two campuses can be described to a reader the same way. True or false?

False. Nearly everything changes. If each campus is approximately normal, about two-thirds of students fall within one standard deviation of the mean: 64 to 72 when \( \mathrm{SD} = 4 \), versus 52 to 84 when \( \mathrm{SD} = 16 \). Without an approximately normal shape, that two-thirds rule is not licensed. The identical means license an identical headline and describe very different institutions — one with scores tightly concentrated around the middle, and one with scores much more dispersed. The summaries alone do not determine the detailed shape. A mean reported without a spread is half a sentence.
Outliers · predict before reveal

A single value of 240 appears in a set otherwise ranging from 30 to 95. It sits so far outside the rest of the data that deleting it is the defensible first move. True or false?

False. Diagnose before deciding. Is it a data-entry slip (a 24.0 that lost its decimal point), a different unit, a genuine extreme, or a member of a different population altogether? Only the first two justify correcting the value, and none justifies deleting it silently. If it is genuine, retain it unless a pre-specified rule supports exclusion; assess its influence and report a labeled with-and-without sensitivity analysis when the conclusion changes materially. “An outlier was removed”, with no diagnosis attached, is the sentence a reviewer stops on.
Group overlap · predict before reveal

Two groups can have clearly different means and still overlap heavily in a visualization, because a difference in centres says nothing about the variation within each group. True or false?

True. Mean difference describes center, not the full distribution. Large within-group variation can create substantial overlap even when centers differ — which is why this module asks for an overlap or effect-size statement alongside any difference in means.

Interactive reading: PMF and CDF worked example

Open the five-step PMF/CDF worked example · about 12 minutes
Keep the arithmetic simple. Use R to do the calculation, then interpret the result. Round probabilities, means, variances, standard deviations, and effect sizes to two sensible decimal places unless a prompt says otherwise. Small rounding differences are not mistakes, and no assessed question requires memorizing the decimals in this example.

Many applied statistics and biostatistics sequences teach “the probability distribution of a discrete random variable” without ever naming the probability mass function or the cumulative distribution function, so the labels can be new even when the idea is not. Work this example once, by hand, before the reactive table lab — the lab uses the same five probabilities, so every number you compute here should reappear on screen.

Step 1 · The random variable and its probability table

Let \( X \) be the number of retakes a randomly chosen student needs after the first sitting, recorded as \( 0, 1, 2, 3, 4 \) (where 0 means the student passed on the first sitting and needs no retake). A probability table assigns one probability to each value:

\( x \)01234Total
\( P(X = x) \)0.080.200.340.260.121.00

That row is the PMF. Formally, the probability mass function is \( p(x) = P(X = x) \), and a table is a valid PMF exactly when \( 0 \le p(x) \le 1 \) for every value and \( \sum_x p(x) = 1 \). Here \( 0.08 + 0.20 + 0.34 + 0.26 + 0.12 = 1.00 \), so the table is valid. If the total were \( 0.97 \) or \( 1.05 \), nothing downstream — expected value, spread, or any chart — would be interpretable until the table was repaired.

If you worked Module 2's raw-tally example — the observed attempts \( 1, 1, 2, 3, 3, 3 \) became a table by tallying each value and dividing by \( n = 6 \) — you have already built a probability table the long way: count, then divide. The table above hands you the probabilities directly, and the two validity checks are the same either way. In Module 2 the entries came from data; here they describe a model of the process. The reactive probability table lab edits this same table live.

Step 2 · Build the CDF by accumulating

The cumulative distribution function answers “what is the probability of at most this many retakes?”: \( F(x) = P(X \le x) = \sum_{k \le x} p(k) \). You get it by running a cumulative sum down the PMF:

\( x \)\( p(x) = P(X = x) \)Running sum\( F(x) = P(X \le x) \)
00.08\( 0.08 \)0.08
10.20\( 0.08 + 0.20 \)0.28
20.34\( 0.28 + 0.34 \)0.62
30.26\( 0.62 + 0.26 \)0.88
40.12\( 0.88 + 0.12 \)1.00

Three checks that the CDF is right: it never decreases, it ends at exactly 1, and each PMF value can be recovered as a difference of neighbours, \( p(x) = F(x) - F(x-1) \). For example \( P(X = 2) = F(2) - F(1) = 0.62 - 0.28 = 0.34 \). For a discrete variable the CDF is a step function — flat between the allowed values, jumping by \( p(x) \) at each one — which is what the right-hand panel of the lab draws.

Reading probabilities from the CDF instead of adding PMF cells:

  • \( P(X \le 2) = F(2) = 0.62 \) — 62% of students need at most two retakes.
  • \( P(X > 2) = 1 - F(2) = 0.38 \).
  • \( P(1 \le X \le 3) = F(3) - F(0) = 0.88 - 0.08 = 0.80 \). Subtract \( F \) at the value just below the lower limit, because \( F(1) \) would remove \( P(X = 1) \) itself. Unpacked, the difference keeps exactly the interval's PMF cells — \( P(X=1) + P(X=2) + P(X=3) = 0.20 + 0.34 + 0.26 = 0.80 \) — so the subtraction rule and adding the cells are the same computation seen two ways.
Step 3 · Expected value from the table

The expected value is the probability-weighted average of the values, \( E(X) = \sum_x x\,p(x) \):

\[ E(X) = 0(0.08) + 1(0.20) + 2(0.34) + 3(0.26) + 4(0.12) = 0.20 + 0.68 + 0.78 + 0.48 = 2.14 \]

Over many students the mean number of retakes is 2.14. Note that \( E(X) = 2.14 \) is not a value \( X \) can take — an expected value is a long-run average, not a prediction for one student.

Step 3 (continued) · Variance and standard deviation from the table

The variance uses the shortcut \( \operatorname{Var}(X) = E(X^2) - [E(X)]^2 \). First \( E(X^2) = \sum_x x^2 p(x) \):

\[ E(X^2) = 0(0.08) + 1(0.20) + 4(0.34) + 9(0.26) + 16(0.12) = 0.20 + 1.36 + 2.34 + 1.92 = 5.82 \]

\[ \operatorname{Var}(X) = 5.82 - 2.14^2 = 5.82 - 4.5796 = 1.2404, \qquad \mathrm{SD}(X) = \sqrt{1.2404} \approx 1.114 \]

Interpretation: over many students the mean number of retakes is 2.14, with a standard deviation of about 1.1 retakes. Open the reactive table lab without changing anything and confirm the rounded story: total probability is 1, the mean is about 2.14 retakes, the variance is about 1.24, and the standard deviation is about 1.11.

Online R Lab · Variance

How Spread Out Are the Data? Understanding Variance in R. By the end of this lab you should be able to: explain variance in your own words; calculate variance using R; compare two groups using variance; explain why a larger variance means more spread; and connect variance to standard deviation.

Part 1 — Create two data sets. The starter defines two groups with the same mean: group_a <- c(9, 10, 10, 11, 10) and group_b <- c(2, 6, 10, 14, 18). Run the starter to see each group’s mean printed in real R. Part 2 — Look at the data visually. The same starter loads ggplot2 and draws a dot plot of both groups — one dot per value. Which group looks more spread out? Part 3 — Calculate variance. var(group_a) and var(group_b) print each group’s variance, and the box is yours from there: create your own data set and calculate its variance. The code box below is editable — change any number, add values, or invent your own data set, then press Run and watch the printed values and the dot plot update instantly.

Step 4 · Empirical probability table versus theoretical model

Module 3 asks you to defend or reject a model, which only makes sense once you separate two objects that both get drawn as bar charts:

Empirical distributionTheoretical model
Where it comes fromYour data. \( \hat{p}(x) = \dfrac{\text{count of } x}{n} \) — relative frequencies from the sample.A formula chosen from the data-generating situation, e.g. Poisson: \( P(X = k) = \dfrac{e^{-\lambda}\lambda^k}{k!} \), or Binomial \( (n, p) \), or Geometric \( (p) \).
What it describesThis sample of \( n \) students. A new sample gives different proportions.The assumed population process. It does not change when you resample.
Free parametersNone — every proportion is read off the data.One or two (\( \lambda \), or \( n \) and \( p \)), usually estimated from the data, e.g. \( \hat{\lambda} = \bar{x} \).
In Rtable(attempts) / length(attempts), cumsum(), ecdf()dpois(), ppois(), dbinom(), pbinom(), dgeom()
What can go wrongSmall \( n \): proportions are noisy and rare values may not appear at all.Wrong assumptions: dependence, a rate that changes across students, or more spread than the model allows (over-dispersion).

Worked comparison. Suppose 40 students at one campus needed the following retakes: 0 retakes for 6 students, 1 for 11, 2 for 10, 3 for 8, 4 for 3, and exactly 5 for 2. The empirical table divides each count by 40; the sample mean is \( \bar{x} = 77/40 = 1.925 \), so a Poisson model with \( \hat{\lambda} = 1.925 \) (computed with dpois(0:5, 1.925)) gives:

\( x \)012345
Count61110832
Empirical \( \hat{p}(x) \)0.1500.2750.2500.2000.0750.050
Poisson(1.925) \( p(x) \)0.1460.2810.2700.1730.0830.032

Across the displayed exact counts, the two rows disagree by at most about 0.03. The fitted Poisson model also assigns about 0.014 probability to counts above 5; the observed sample has none. A single student changes an empirical cell by 0.025 when \( n = 40 \), so the displayed differences are consistent with ordinary sampling variation. Thus, “the Poisson model is consistent with these data” is defensible — not a proof that retake counts are Poisson. A model is challenged when the empirical row shows a feature the model cannot readily produce: a second peak, far more zeros than \( e^{-\hat{\lambda}} \) allows, or variance far larger than the mean. For a Poisson model, \( \operatorname{Var}(X) = \lambda \), so \( s^2 \gg \bar{x} \) is an overdispersion warning. Say which checks you used in your Project 1 defense.

Step 5 · The same calculation in R (paste into the browser lab)
# PMF -> CDF -> E(X) -> Var(X), from a probability table
x <- 0:4
p <- c(0.08, 0.20, 0.34, 0.26, 0.12)
sum(p)                        # validity check: must print 1
cdf <- cumsum(p); cdf         # F(x): 0.08 0.28 0.62 0.88 1.00
ex  <- sum(x * p); ex         # E(X) = 2.14
vx  <- sum(x^2 * p) - ex^2    # Var(X) = 1.2404
c(variance = vx, sd = sqrt(vx))

# Empirical table versus a fitted Poisson model
attempts <- rep(0:5, times = c(6, 11, 10, 8, 3, 2))   # the 40-student sample above
n <- length(attempts)
empirical <- as.numeric(table(factor(attempts, levels = 0:5))) / n
lambda_hat <- mean(attempts); lambda_hat            # 1.925
model <- dpois(0:5, lambda = lambda_hat)
round(rbind(x = 0:5, empirical = empirical, poisson = model), 3)
ppois(5, lambda = lambda_hat, lower.tail = FALSE) # model probability above 5: about 0.014
c(sample_mean = mean(attempts), sample_var = var(attempts))  # Poisson needs these close

Every line prints something you can quote in the “Discrete distribution evidence” field of the graded assignment: the validity total, the CDF, \( E(X) \), \( \operatorname{Var}(X) \), the empirical-versus-model table, and the mean-versus-variance check.

Reactive probability table lab

Open the reactive PMF/CDF table · about 8 minutes

Build a valid discrete distribution before comparing visual shape. Adjust the probabilities and watch the total, expected value, variance, PMF, and CDF update immediately.

Accessible fallback: probability table worksheet

If the interactive does not load or is not accessible with your device or assistive technology, use the worked two-toss table. Verify that probabilities are between \(0\) and \(1\), add them, compute the cumulative column, then compute \(E(X)\), \(\operatorname{Var}(X)\), and SD. Equivalent R calculations or written work are accepted in the Canvas packet.

Distribution shape lab

Open the distribution-shape comparison · about 8 minutes
Accessible fallback: distribution comparison

Use the static summaries and descriptions on this page. For each campus, record center, spread, skew or symmetry, peaks, gaps, tails, outliers, and overlap. State which conclusion a box plot supports and which shape details require a histogram, density, violin, or dot plot.

Effect-size lab

Open the effect-size interpretation challenge · about 8 minutes
Accessible fallback: effect-size worksheet

Use \(d=(\bar{x}_2-\bar{x}_1)/s_{\text{pooled}}\). Show the raw difference, divide by the pooled SD, interpret the standardized magnitude in context, describe overlap, and state that the comparison does not establish causation or individual separation.

AI-output audit

Open the AI-output audit · about 10 minutes
Accessible fallback: AI verification record

Write the AI claim, the R output or authoritative course evidence used to check it, and the correction you made. Your own explanation is required; an AI response is not evidence by itself.

Hands-on visual lab

Open the histogram bin-width explorer · about 8 minutes

These scores came from two distinct populations. Bin width is a parameter you choose, and it decides whether the reader can see that at all — including whether the second population survives the chart. Then run the browser-R lab below without leaving QuantegyAI.

Accessible fallback: two-width histogram investigation

Create or inspect the same distribution at two defensible bin widths. Record both widths, describe the mode, gap, skew, or subgroup pattern at each width, and state whether the structure persists. If the conclusion changes, label it uncertain. Screenshots or equivalent written descriptions are accepted.

R lab

You predicted the model and comparison story. Now verify it computationally. Run the provided students data, compare the model with the observed distribution, and carry the checked evidence into DV05 and DV06.
Verify the three discrete models in browser R · about 10 minutes
Model-choice verification lab

Recompute the binomial, geometric, and Poisson diagrams from the Binomial, geometric, or Poisson? page in real R. Compare the printed table with the chart bars, then defend which model fits a scenario and why.

Browser R note: this inline lab runs base R in the browser — no install needed. Read the printed table against the three diagrams on that page.

Open the runnable browser-R lab · about 15 minutes
R / Quarto / Shiny workflow

The lab is complete as provided. Press Run, read the three printed lines — center, spread, overlap — then answer the True/False statements under the output.

Browser R note: this inline lab runs base R in the browser — no install, no download. For the graded DV05 and DV06 evidence steps, use the Canvas evidence section above the record; its checklists say exactly what each packet needs.

Preserved Project 1 evidence: DV05 and DV06

You are not doing separate assignments here.

DV05 and DV06 are the Module 3 evidence included in Canvas Project 1. Complete each item once; combine it with Modules 1 and 2 in the same packet.

Open the DV05 and DV06 Canvas evidence checklists

Complete DV05 and DV06 for Canvas Project 1. Modules 1, 2, and 3 belong in one Project_1_Canvas_Packet.zip.

DV05 minimum evidence and rubric

  • Valid PMF and validity check
  • CDF with a probability interpreted in words
  • Expected value, variance, and SD with context
  • Empirical-versus-theoretical model check
  • Reproducible R evidence
  • A specific limitation

Strong: correct evidence, interpretation, and a defensible model claim. Developing: mostly correct calculations but incomplete interpretation or checking. Insufficient: invalid/missing evidence or an unsupported model claim.

DV06 minimum evidence and rubric

  • Named groups, quantitative outcome, target contrast, and each group’s sample size
  • Center, spread, and shape comparison
  • Overlap and unusual-observation statement
  • One appropriate visual with its relevant display choices stated
  • Signed raw contrast and effect-size interpretation in context
  • Descriptive or inferential status stated explicitly
  • For population claims, an uncertainty interval and its assumptions
  • Reproducible R evidence
  • A limitation and a statement avoiding individual or causal overclaiming

Strong: integrates distribution evidence, magnitude, overlap, and limits. Developing: includes evidence but relies too heavily on means or leaves interpretation incomplete. Insufficient: reports numbers/charts without a defensible comparison.

Reveal generic strong and weak examples after drafting

Stronger: “South’s mean is higher, but the standardized difference is modest and the distributions overlap; this comparison does not establish a campus effect.”

Needs revision: “South is better because its mean is higher.”

Evidence names: use DV05_distribution-model_[your-name] and DV06_group-comparison_[your-name]. Existing saved drafts and uploaded files remain valid. For Project 1, use the DV05 and DV06 checklists on this page. These are evidence identifiers within Module 3, not links to another numbered module or project.

Evidence-readiness decision check

The Module 3 submission decision check needs JavaScript.

Module 3 AI-feedback workspace

You have already done most of the work.

Now package the distribution evidence, group-comparison evidence, R verification, and AI-audit reasoning you produced above. This is formative evidence review, not another assignment and not a Canvas submission: requesting AI feedback here does not submit any project to Canvas. Project 1 combines DV01–DV06 from Modules 1–3 in one packet zip.

Open the shared Module 3 AI-feedback workspace

Where do I get my data?

Use students. Everything you need is already provided in the R lab. It contains the discrete counts and campus groups needed for the guided DV05 and DV06 path.

Advanced option: Apply this workflow to your own data

You may use a built-in R dataset or add a CSV under Your data in the R lab. The provided students dataset remains the recommended route.

How does grading work?

The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.

Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit the corresponding project packet to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like.

Required AI verification record

AI drafts the analysis; R checks it; you explain why the result holds. Complete this table for every material AI-assisted claim.

AI claimHow I checked itCorrection made
Reveal a generic example after completing your row

AI claim: “The Poisson model fits.” Check: compared the empirical PMF with dpois() and checked mean versus variance. Correction: “The Poisson model is consistent with these data; the comparison does not prove independence or a stable rate.”

Accessible submission fallback

If this form is unavailable, complete the DV05 and DV06 minimum-evidence checklists, the AI verification record, and the corresponding Canvas packet. Equivalent calculations, screenshots, R output, and written interpretations are accepted. AI feedback is formative; Dr. Roberts’s Canvas review determines the grade.

Adaptive reflection

What I know / where I go next

Mastery check loading…

Module 3 record

Your saved Module 3 record needs JavaScript. You can also see it on your skills dashboard report card.

Final Project connection Add this module’s evidence to your Final Project readiness tracker. If your explanation depends on AI, include the prompt, output, verification method, and final correction.