Module 2 — Descriptive Statistics & Visual Encoding
Use center, spread, robust summaries, axes, scales, and labels to make visual structure mathematically honest.
Module 1 ended with a clean dataset and a defensible question. Module 2 asks what that dataset is allowed to say and how to draw it honestly. In the first half we work with descriptive statistics: the mean, median, standard deviation, and interquartile range, which of them survive an extreme value and which do not, and how to check a summary table an assistant hands us instead of trusting it. We also meet the empirical probability table, the bridge to Module 3, by counting outcomes and dividing by the sample size.
In the second half, we turn numbers into visual marks. We will learn the main visual channels a chart can use, including position, length, angle, area, and color. We will also learn which channels readers interpret most accurately, when a bar chart must start at zero, when a dot plot does not need to, and how to fix a misleading chart even when none of the labels are technically wrong.
We will test our design choices in browser-based R and record two pieces of evidence for Project 1:
- DV03: a descriptive evidence memo
- DV04: a visual encoding defense
Each lab explains key terms before asking you to answer questions. If you get stuck, a True/False option will help you finish the step and keep moving forward.
By the end of this module, you will be able to
- Compute and compare center, spread, and shape, and say which summaries survive an extreme value.
- Verify the numerical distribution before selecting any chart.
- Defend axes, scales, baselines, and labels, and state when a non-zero axis is justified.
- Choose visual channels — position, length, angle, area, color — by how accurately readers judge them.
- Accept, revise, or reject AI chart advice with reasons, and fix a misleading chart even when no label is technically wrong.
- Build two competing views of the same verified data and defend the more honest one.
- Apply all of this to Project 1: your DV03 Descriptive Evidence Memo and DV04 Visual Encoding Defense, due the date on Canvas.
Your goal
Verify the summaries, then choose and defend the chart that answers the question honestly.
What you’ll learn
Statistics focus: mean, median, IQR, standard deviation, robust summaries, scale and perception.
R/tool focus: Inline browser R labs (base R + ggplot2, no install needed); optional desktop RStudio/Positron with tidyverse for the Project 1 assignment; ggplot2 aesthetics, scales, labels, and transformations.
What students will cover. Module 2 turns cleaned data into defensible visual choices. Students connect the statistical question to the variable types, then decide whether the audience needs a comparison, distribution, relationship, trend, part-to-whole statement, or uncertainty-aware summary.
- Chart choice: bar charts, dot plots, histograms, boxplots, scatterplots, line plots, small multiples, and when not to use pie charts.
- Visual encoding: position, length, color, size, shape, area, angle, facets, labels, scale, baseline, and perceptual accuracy.
- Grammar of graphics:
data,aes(),geom_*, scales, coordinates, facets, labels, and themes inggplot2. - AI critique: audit model-generated chart recommendations for mismatched graph types, misleading scales, hidden denominators, and causal overreach.
- Interactive math: use GeoGebra to test how axes, slope, proportion, area, and truncation change interpretation without changing the data.
Student learning outcomes. By the end of Module 2, students can…
- Match statistical questions to appropriate visualization types based on the variables and comparison being made.
- Use data type and measurement level to justify chart choice before opening a graphing tool.
- Explain how visual encodings communicate information through position, length, color, size, shape, area, angle, and facets.
- Compare encoding accuracy and explain why common-axis position and length are usually more precise than area, angle, or decorative color.
- Build basic visualizations in browser R using
ggplot2mappings, geoms, scales, labels, and themes. - Revise misleading charts by changing the mark, scale, baseline, denominator, label, or grouping structure.
- Critique AI-generated chart recommendations using the original data, statistical question, and visual-encoding principles.
- Defend a final visualization in writing with code, interpretation, limitations, and an AI-audit reflection.
About your submission
Project 1 evidence: complete DV03 and DV04 on this page, request formative AI feedback in the Module 2 evidence form, then complete all remaining evidence from Modules 1–3 in Module 1 if you have not already. When DV01–DV04 are complete, open Project 1, download its Canvas Packet zip, and upload it to Canvas.
Interactive assignment: Robustness explorer, summary-stat verifier, encoding channel lab, chart repair lab, axis-truncation explorer.
Deliverables: verified summary table with visual justification, before/after chart redesign with a written defence of scale, baseline, and encoding.
Important: “Get AI feedback →” saves formative feedback inside QuantegyAI and does not submit to Canvas. Project 1 is one final zip containing DV01–DV06 from Modules 1–3; only that packet is submitted to Canvas.
Keep it simple: use the provided data and R output, focus on what each result means, and write in a few clear sentences. No question requires long arithmetic or exact decimal matching.
Interactive reading
Five short claims about summaries and encodings. Select True or False before opening each explanation, then read the explanation either way. Reading ahead is allowed; the check simply waits for your answer.
Terms used in these checks.
- Mean and median
- The mean is the arithmetic average; the median is the middle value once the data are sorted. They agree for symmetric data and separate when there is a long tail.
- Standard deviation (SD)
- A measure of spread built from every observation’s distance to the mean, so one extreme value inflates it.
- Interquartile range (IQR)
- The distance between the 25th and 75th percentiles: the span of the middle half of the data.
- Encoding
- The rule that turns a number into something visible — a position, a length, an area, an angle, or a color.
- Baseline
- The value a bar’s length is measured from. Bars carry magnitude by length, so their baseline is normally zero.
- Denominator
- The number a rate is divided by. A rate without its denominator hides how much evidence stands behind it.
- Which summary resists an extreme value. Take the values 4, 6, 7, 9 and 40. The mean is 13.2, above four of the five values, because the 40 pulls it upward. The median stays at 7, because moving the largest value further out does not change which value is in the middle. The median is robust to one extreme value; the mean is not.
- Summaries do not determine shape. The mean and SD do not determine a distribution’s shape. Two datasets can share the same mean and the same SD while one is symmetric with a single peak and the other has two separate clusters. Only a plot of the distribution, such as a histogram, shows the shape.
- Some encodings are read more accurately than others. Readers judge position along a common scale more accurately than length, and length more accurately than angle or area. When the task is to compare close values precisely, put them on a common scale.
- Correct labels do not make a bar honest. Readers compare bars by their length. If the axis starts above zero, every label can be correct while the bar lengths overstate the differences between the values.
A distribution has one extreme high value. Because the mean uses every observation, it is still the more defensible report of the typical value. True or false?
Two datasets can have the same mean and the same standard deviation and still have different shapes, so a histogram may tell a different story from the summary table. True or false?
A chart uses bubble area to compare six close pass rates. For a precise comparison of close values, area is usually the clearest encoding. True or false?
A bar chart compares rates of 92%, 94%, and 96%, and its vertical axis starts at 90%. Because every label on the chart is correct, it cannot mislead a reader. True or false?
One program has 18 completions out of 20 attempts, and another has 180 completions out of 240 attempts. Judged by completion rate rather than by raw count, the smaller program did better. True or false?
Why distributions matter
A distribution is the full pattern of values in a variable — not just its average. It shows where observations cluster, how far they spread, whether the shape is symmetric or skewed, whether there are gaps or multiple peaks, and whether unusual values deserve explanation rather than deletion.
This matters because a single summary can hide the very feature the audience needs to understand. Two classes can have the same mean score but very different spreads; two campuses can have the same median pass rate but different tails; a small group can look extreme because its denominator is unstable. Before students choose a chart, they should ask what the distribution is licensing them to claim.
In the Module 2 graded assignment, distribution thinking becomes the bridge between statistics and design: histograms and density views reveal shape, boxplots emphasize median and spread, dot plots preserve individual cases for small samples, and summary tables verify whether the visual story is numerically defensible.
Interactive reading: from a frequency table to a probability table
The summary table you verify in this module already contains counts and proportions. Module 3 will call the proportion column an empirical probability mass function and its running total a cumulative distribution function. Those names are new for many students even when the arithmetic is not, so meet them here, on six numbers, before the Module 3 labs ask you to compute them.
Suppose the observed number of attempts for six students is \( 1, 1, 2, 3, 3, 3 \). Tally each value and divide by \( n = 6 \):
| Value \( x \) | Count | Relative frequency \( \hat{p}(x) = \dfrac{\text{count}}{n} \) |
|---|---|---|
| 1 | 2 | \( 2/6 \approx 0.333 \) |
| 2 | 1 | \( 1/6 \approx 0.167 \) |
| 3 | 3 | \( 3/6 = 0.500 \) |
| Total | 6 | \( 6/6 = 1 \) |
The third column is the empirical PMF: for each value, the share of the data that took that value. Two checks make it a valid probability table — every entry lies between 0 and 1, and the entries sum to exactly 1. The mean you would compute from the raw data is the same number you get by weighting each value by its probability: \( 1(2/6) + 2(1/6) + 3(3/6) = 13/6 \approx 2.17 \), which is exactly \( \bar{x} = 13/6 \). A probability table does not lose the mean; it reorganises it.
The CDF answers “what share of the data is at most this value?”: \( \hat{F}(x) = P(X \le x) \), a running sum down the PMF column.
| Value \( x \) | PMF \( \hat{p}(x) \) | CDF \( \hat{F}(x) = P(X \le x) \) |
|---|---|---|
| 1 | \( 2/6 \) | \( 2/6 \approx 0.333 \) |
| 2 | \( 1/6 \) | \( 3/6 = 0.500 \) |
| 3 | \( 3/6 \) | \( 6/6 = 1 \) |
Read it two ways. Numerically: \( P(X \le 2) = \hat{F}(2) = 0.5 \), so half the students needed at most two attempts, and \( P(X = 3) = \hat{F}(3) - \hat{F}(2) = 1 - 0.5 = 0.5 \) recovers the PMF entry as a difference of neighbours. Visually: a PMF is drawn as a bar chart whose bar length is the probability — so this module's zero-baseline rule applies with full force — while a CDF is a step function that climbs from 0 to 1 and never decreases. If your CDF ever goes down, or ends anywhere but 1, the table underneath it is wrong.
- An empirical probability table is calculated from observed data. Its entries are relative frequencies. Collect six more students and every entry can change.
- A theoretical probability model is based on assumptions about how the data-generating process works. If the six numbers above were rolls of a fair die, the model says \( p(x) = 1/6 \) for each of \( x = 1, \dots, 6 \), before any roll is observed. It has no counts in it at all.
- They can be compared, but they are not the same thing. The empirical table above puts probability \( 0 \) on the values 4, 5, and 6 and \( 1/2 \) on the value 3; the fair-die model puts \( 1/6 \) on every face. With \( n = 6 \) that gap is sampling noise, not evidence against the die — and telling those two apart is exactly the “defend or reject the model” task in Module 3.
In Module 2 you only need the empirical table, which is descriptive statistics. Keep the label straight in your writing: “the proportion of students needing three attempts was 0.50” is an empirical statement about this sample; “the probability a student needs three attempts is 0.50” is a claim about a process, and needs a model or a much larger sample behind it.
# Base R only — press Run in the lab below to run it (same code, preloaded)
x <- c(1, 1, 2, 3, 3, 3)
counts <- table(x); counts # 1:2 2:1 3:3
pmf_hat <- prop.table(counts); round(pmf_hat, 3) # empirical PMF: 0.333 0.167 0.500
cdf_hat <- cumsum(pmf_hat); round(cdf_hat, 3) # empirical CDF: 0.333 0.500 1.000
sum(as.numeric(names(pmf_hat)) * pmf_hat) # mean from the table = 13/6 = 2.167 = mean(x)
ecdf(x)(2) # P(X <= 2) read from the empirical CDF = 0.5
fair_die <- setNames(rep(1/6, 6), 1:6); round(fair_die, 3) # a theoretical model, no data used
Module 3 extends this with expected value and variance from a table, and with the Poisson and binomial models you compare an empirical table against. If those names are new, the Module 3 worked example walks through every step with the same notation.
The die-compare lab needs JavaScript.
See the output yourself: press Run and R executes the code above in your browser — the printed table and the plot appear below the code. The starter is complete as provided; edit it only if you want to test a different idea.
Two-coin lab: run the distinction yourself
Why this lab Run the distinction between independent events yourself.
The reading above named the difference between an empirical probability table (relative frequencies counted from data you actually have) and a theoretical probability model (probabilities written down from assumptions before any data exist). This lab makes you run both on the same random process — tossing two fair coins — so the distinction is something you operated, not just read. The sample space is the list of every equally likely outcome of one toss; counting it correctly is the whole prediction. Then the R scaffold below turns it into project evidence.
Roll two dice and record the sum. The sums 2 to 12 are eleven possible values, but they are not equally likely. The equally likely outcomes are the 36 ordered pairs (first die, second die). A sum of 2 arises from one pair, (1, 1), so \( P(\text{sum} = 2) = 1/36 \). A sum of 7 arises from six pairs, (1, 6), (2, 5), (3, 4), (4, 3), (5, 2) and (6, 1), so \( P(\text{sum} = 7) = 6/36 \). The pairs (1, 6) and (6, 1) count separately because the two dice are different objects.
So list the equally likely paths first, then count how many paths give each value. A value’s probability is the number of its paths divided by the total number of paths.
The two-coin probability lab needs JavaScript.
Probability R scaffold
Simulate the same two-coin process in R: a theoretical probability table written from the fair-coin model, an empirical probability table computed from your simulated tosses, and one graph comparing the two.
- Click Run. The code is complete as provided.
- Compare the empirical bars with the theoretical diamonds. Rounded values are enough.
- Write three short sentences: what the model predicts, what the simulation produced, and why a small gap is ordinary sampling variation.
Summary statistics lab
Why this lab See how mean and spread describe the same data differently.
Numerical summaries are the structure a chart is built on. Before choosing an encoding, establish which summaries this distribution licenses — and check that the ones you were handed are actually right.
The summaries you will move, then verify. The explorer asks which of these follow one extreme value; the verifier asks you to recompute each one from nine numbers.
- Mean
- Add every value and divide by the count \( n \). One extreme value can pull it a long way.
- Median
- The middle value once the data are sorted (with \( n = 9 \), the fifth value). An extreme value cannot move it past its neighbours.
- Mode
- The value that occurs most often. It is not the smallest value.
- Range
- A distance: the largest value minus the smallest, not the largest value on its own.
- IQR
- The third quartile minus the first quartile, the span of the middle half. Quartile conventions differ between software packages, so name yours.
- Standard deviation
- For a sample, \( s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}} \). Dividing by \( n \) instead of \( n-1 \) gives a smaller, biased value.
- Robust
- A summary is robust when there is a ceiling on how far one observation can move it. Median and IQR are robust; mean and SD are not.
The summary statistics lab needs JavaScript.
Encoding channel lab
Why this lab Match each message to the visual channel that carries it best.
Position, length, angle, area, hue, lightness — readers decode these with very different accuracy, and each one carries its own promise about the data. Match the task to the channel it licenses.
The six channels, from most to least accurately read. Each task in the lab is best served by exactly one of them.
- Position on a common scale
- Where a mark sits along a shared axis — a dot plot or scatterplot. The most accurately decoded channel; it carries no promise about zero.
- Length from a zero baseline
- How long a bar is. Readers compare lengths as ratios, so the bar must start at zero for “twice as long” to mean “twice as much”.
- Angle or slice
- The angle of a pie slice. Read least accurately of the quantitative channels, and worse as slices multiply.
- Area
- The size of a circle or square, used when position is already spent (for example on a map). Area grows with the square of the radius.
- Colour hue
- Red versus blue versus green. Separates categories without implying that one is larger than another.
- Colour lightness
- Light to dark within one hue. Reads as “less” to “more”, so it can carry an ordered scale.
The encoding channel lab needs JavaScript.
Chart repair lab
Why this lab Repair a broken chart so it tells the truth accurately.
Nine flawed charts, each with a tempting fix that leaves the defect in place. Diagnose before you repair — the obvious correction is often the wrong one.
Terms the repair options use.
- Truncated axis
- An axis that starts above zero. Harmless for dots and lines when disclosed; it breaks a bar’s length encoding.
- Dual y-axis
- Two series drawn against two different vertical scales on one chart. The relative scaling is the analyst’s free choice, so any apparent relationship can be manufactured.
- Log scale
- An axis where each tick multiplies the previous one (1, 10, 100). Right for multiplicative growth, but it must be labelled with real values and named as logarithmic.
- Sequential palette
- A single hue running from light to dark, used for ordered quantities. A rainbow is not perceptually ordered.
- Data-ink ratio
- The share of the ink on a chart that carries data rather than decoration. Gradients, shadows, and heavy borders lower it.
- Overplotting
- So many overlapping points that density and unusual cases disappear. Transparency, jitter, or binning restores them.
- Visual hierarchy
- Giving the element that answers the question the strongest contrast and muting the context around it.
- Small multiples
- A grid of small panels, one per series or group, drawn on the same shared scales. Each series keeps an honest axis of its own, and the panels can still be compared side by side.
The chart repair lab needs JavaScript.
Data Detective Challenge: What’s Wrong With This Graph?
Why this lab Find every misleading choice hiding in a published graph.
Inspect the deliberately poor visualization, select every defect you can defend, and then repair the same comparison. More than one answer is correct. A defect changes what a reader takes away or blocks them from reading the chart at all; a stylistic preference (such as alphabetical ordering) does neither, and ticking one costs a mark. If the diagnosis has you stuck, the True/False option under it closes the challenge so you can move on.
Before you diagnose: five more ways a graph misleads or cannot be read. The truncated axis from the chart repair lab is one; here are the others.
- Distorted proportions
- A graph is honest when the drawn ratio matches the data ratio: a bar twice as tall should stand for a value twice as large. Truncated axes, 3-D effects and area marks can all break that match.
- 3-D perspective
- Depth, tilt and side faces make nearer bars look larger and hide where a bar’s top meets the axis. Three dimensions add no data to a two-variable comparison.
- Colour that carries nothing
- Colour should carry data, such as group or order. Colours that encode nothing compete for attention without informing, so a reader looks for a meaning that is not there.
- Missing labels
- Every axis and every group needs a label. A reader who cannot tell which bar is which, or what the axis measures, cannot read the chart at all.
- Legend
- A legend must name what each colour means. A legend such as “Series?” leaves the colours unexplained.
The Data Detective challenge needs JavaScript.
AI-output audit
Why this lab Rule on AI claims before trusting them with your data.
You have verified the summaries and matched tasks to channels. Now an assistant proposes six changes to your draft chart. Rule on each one with the encoding rules you just used; the first proposal is exactly the axis truncation the reading checks warned about, and the visual lab that follows lets you test it.
Three ideas you need before you rule. The other proposals use ideas you have already met: robust summaries and dual axes.
- Spread is not shape
- A small standard deviation says the values sit close to their mean. A small SD says nothing about whether the distribution is normal, skewed or has two peaks; only a plot of the distribution shows its shape.
- Part-to-whole
- A pie shows how one whole divides into parts that sum to 100%. Rates for separate groups, such as six campuses’ pass rates, are not parts of one whole, so they do not belong in one pie.
- Formatting versus data
- Rounding tick labels changes how the axis reads, not the data: the stored values keep their full precision. Rounding the stored values themselves throws measurement away.
This AI audit needs JavaScript.
Hands-on visual lab
Why this lab Test the claim you refused with a visual experiment.
In the audit the assistant proposed starting the bar chart’s y-axis at 92 so the difference would show, and you rejected it. Now test that refusal directly: one design decision, six real pass rates, and a ratio that comes apart from the data. Move the baseline and watch how far a reader can be misled without a single incorrect label.
This visual lab needs JavaScript.
GeoGebra lab
Why this lab Explore how changing an encoding changes what you see.
The same truncation, drawn as a geometric object. Use GeoGebra to make scale, proportion, and visual distortion visible. The AI can draft construction commands, but you must verify the visual claim against the data.
Embedded GeoGebra applet loading…
R lab
Why this lab Reproduce the descriptive evidence with real R output.
Compute mean, median, SD, and IQR; build multiple ggplot2 views of the same data; then write a Chart Choice Defense explaining why one visualization answers the question more honestly than another.
Browser R note: this inline lab runs base R plus ggplot2 (downloaded automatically on first run — one short wait, cached after that; if the download is blocked the lab status says so and Run retries it). The full tidyverse collection is a desktop tool for RStudio/Positron. Everything this module requires runs right here in the browser.
The scaffold below already runs end to end — that is deliberate, so the first thing you see is working output, not an error. Every code block is provided and runs as pasted. Output notes tell you what to look for before submitting. Your Project 1 evidence comes from what you add:
- Verified summaries — keep step 1, and use the provided
spreadoutput containing the \( \mathrm{IQR} \) and standard deviation. Paste the printed values into Numerical summaries and say which summary the distribution licenses (median or mean, and why). - The provided third view of the same data, saved as
p_third— so that you can show how one encoding or baseline change alters the reader's impression while the data stay fixed. - The Chart Choice Defense (step 4) — 100–150 words naming which chart answers the question honestly, which claim the other view would invite, and one limitation (six campuses, unequal
testedcounts). This is the Final defense field of the graded assignment. - The AI audit — if you asked an assistant which chart to use, paste the prompt and its recommendation, then state what you verified and corrected. This is the AI audit field.
Which package: base R for every summary (summary(), mean(), median(), IQR(), sd()) and ggplot2 for every plot. Nothing in this module needs dplyr, tidyr, readr, or tibble, and those are not installed in browser R — if an assistant suggests them, that is a finding for your AI audit, not a reason to install anything.
Use the Copy code and Copy output buttons on the inline lab: the code goes in your reproducibility section, the printed output is your evidence.
# Module 2 R scaffold — complete and ready to run
# Use the output to make a chart-choice decision; do not chase exact decimals.
library(ggplot2) # the only package this module needs; base R does the summaries
## PROVIDED — the dataset. One row = one campus; pass_rate = passed / tested.
scores <- data.frame(
campus = c("North", "South", "East", "West", "Central", "Online"),
pass_rate = c(0.92, 0.94, 0.95, 0.96, 0.93, 0.95),
tested = c(84, 122, 76, 98, 111, 64)
)
## PROVIDED — step 1: verify the summaries before you visualize anything
summary(scores$pass_rate)
## EXPECTED Min. 0.9200 1st Qu. 0.9325 Median 0.9450 Mean 0.9417 3rd Qu. 0.9500 Max. 0.9600
cat("Mean pass rate:", round(mean(scores$pass_rate), 3), "\n") # 0.942
cat("Median pass rate:", round(median(scores$pass_rate), 3), "\n") # 0.945
cat("Range:", paste(range(scores$pass_rate), collapse = " to "), "\n") # 0.92 to 0.96
## PROVIDED — step 1b: spread, rounded for interpretation
spread <- c(iqr = IQR(scores$pass_rate), sd = sd(scores$pass_rate))
round(spread, 2) # sensible rounding is enough; no four-decimal matching required
## PROVIDED — step 2: full-baseline bar chart (length from zero — the honest bar)
p_bar <- ggplot(scores, aes(x = reorder(campus, pass_rate), y = pass_rate)) +
geom_col() +
coord_flip() +
scale_y_continuous(limits = c(0, 1)) +
labs(title = "Full-baseline bar chart", x = "Campus", y = "Pass rate")
## PROVIDED — step 3: dot plot on an annotated narrow range (position, not length)
p_dot <- ggplot(scores, aes(x = pass_rate, y = reorder(campus, pass_rate))) +
geom_point(size = 3) +
scale_x_continuous(limits = c(0.90, 0.97)) +
labs(title = "Dot plot for close pass-rate comparisons", x = "Pass rate (axis runs 0.90-0.97)", y = "Campus")
print(p_bar)
print(p_dot)
## EXPECTED two plots: six bars that look almost equal, then six dots spread across the axis
## PROVIDED — step 3b: the misleading comparison view, ready to run
p_third <- ggplot(scores, aes(x = reorder(campus, pass_rate), y = pass_rate)) +
geom_col() +
coord_cartesian(ylim = c(0.90, 0.97)) +
labs(title = "Truncated bar chart (what not to publish)", x = "Campus", y = "Pass rate")
print(p_third)
## INTERPRET the values did not change, but the truncated bars exaggerate the gap
## WRITE — three short sentences in the Final Chart Choice Defense field:
## 1. which chart answers the question honestly; 2. what the other view exaggerates;
## 3. one limitation, such as unequal tested counts.
- The lab status reads “R ran successfully” and there is no red
Errorline in the output. Iflibrary(ggplot2)fails, the status tells you it is a download problem and Run R retries it; the base-R summaries are still valid evidence. - Every
## EXPECTEDvalue matches your console to the printed rounding. Runsum(scores$tested)— it must print555; anything else means the dataset was edited. spreadandp_thirdexist:exists("spread")andexists("p_third")both printTRUE.- Your defense names the encoding channel (position versus length), the baseline decision, the axis range you disclosed, and one limitation. The graded assignment's rubric checks for exactly those four things.
Project 1 evidence: DV03 and DV04
DV03 and DV04 are the evidence you will use in the Module 2 evidence form below. Complete and save both once; your work carries forward.
Module 2 evidence and AI feedback
Now package the verified summaries, competing charts, and AI-audit reasoning you produced above. This is one evidence-review step, not another assignment and not a Canvas submission.
Where do I get my data?
Use scores. Everything you need is already provided in the R lab. It contains the verified campus summaries needed for DV03 and DV04.
Advanced option: Use your own dataset
You may use a built-in R dataset or add a CSV under Your data in the R lab. The provided scores dataset remains the recommended route.
How does grading work?
The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.
Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit Project 1 to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like. If a long textbox is the only thing in your way, a short True/False concept check can close it.
- Comparing a few values across groups: a bar chart from zero, or a dot plot. For close values, use a dot plot on a disclosed narrow scale.
- The distribution of one variable: a histogram shows the shape; a box plot shows the median, the spread and unusual values.
- A relationship between two quantitative variables: a scatterplot, with one point per observational unit and both axes labelled with units.
- Comparing compositions or patterns across several groups: small multiples on shared scales, one panel per group, rather than several pies.
- A claim you cannot verify: if your computed summaries and a chart or caption cannot be reconciled, do not publish the chart. Report the checked summary table and state the limitation.
The inline Module 2 graded assignment and AI grader need JavaScript.
Prepare for the mastery check
Mastery check loading…
What Module 2 saved
Your saved Module 2 evidence needs JavaScript. You can also see it later on your course report card.