← All Data Visualization with AI modules

Module 2 — Descriptive Statistics & Visual Encoding

QuantegyAI Module 2 · Canvas Project 1 · Due the date on Canvas

Use center, spread, robust summaries, axes, scales, and labels to make visual structure mathematically honest.

What we will study in this module

Module 1 ended with a clean dataset and a defensible question. Module 2 asks what that dataset is allowed to say and how to draw it honestly. In the first half we work with descriptive statistics: the mean, median, standard deviation, and interquartile range, which of them survive an extreme value and which do not, and how to check a summary table an assistant hands us instead of trusting it. We also meet the empirical probability table, the bridge to Module 3, by counting outcomes and dividing by the sample size.

In the second half, we turn numbers into visual marks. We will learn the main visual channels a chart can use, including position, length, angle, area, and color. We will also learn which channels readers interpret most accurately, when a bar chart must start at zero, when a dot plot does not need to, and how to fix a misleading chart even when none of the labels are technically wrong.

We will test our design choices in browser-based R and record two pieces of evidence for Project 1:

Each lab explains key terms before asking you to answer questions. If you get stuck, a True/False option will help you finish the step and keep moving forward.

By the end of this module, you will be able to

Start here

Your goal

Verify the summaries, then choose and defend the chart that answers the question honestly.

TimeAbout 90–120 minutes
We will saveDV03 + DV04 evidence
Start Module 2 →
What you’ll learn

Statistics focus: mean, median, IQR, standard deviation, robust summaries, scale and perception.

R/tool focus: Inline browser R labs (base R + ggplot2, no install needed); optional desktop RStudio/Positron with tidyverse for the Project 1 assignment; ggplot2 aesthetics, scales, labels, and transformations.

What students will cover. Module 2 turns cleaned data into defensible visual choices. Students connect the statistical question to the variable types, then decide whether the audience needs a comparison, distribution, relationship, trend, part-to-whole statement, or uncertainty-aware summary.

  • Chart choice: bar charts, dot plots, histograms, boxplots, scatterplots, line plots, small multiples, and when not to use pie charts.
  • Visual encoding: position, length, color, size, shape, area, angle, facets, labels, scale, baseline, and perceptual accuracy.
  • Grammar of graphics: data, aes(), geom_*, scales, coordinates, facets, labels, and themes in ggplot2.
  • AI critique: audit model-generated chart recommendations for mismatched graph types, misleading scales, hidden denominators, and causal overreach.
  • Interactive math: use GeoGebra to test how axes, slope, proportion, area, and truncation change interpretation without changing the data.

Student learning outcomes. By the end of Module 2, students can…

  1. Match statistical questions to appropriate visualization types based on the variables and comparison being made.
  2. Use data type and measurement level to justify chart choice before opening a graphing tool.
  3. Explain how visual encodings communicate information through position, length, color, size, shape, area, angle, and facets.
  4. Compare encoding accuracy and explain why common-axis position and length are usually more precise than area, angle, or decorative color.
  5. Build basic visualizations in browser R using ggplot2 mappings, geoms, scales, labels, and themes.
  6. Revise misleading charts by changing the mark, scale, baseline, denominator, label, or grouping structure.
  7. Critique AI-generated chart recommendations using the original data, statistical question, and visual-encoding principles.
  8. Defend a final visualization in writing with code, interpretation, limitations, and an AI-audit reflection.
About your submission

Project 1 evidence: complete DV03 and DV04 on this page, request formative AI feedback in the Module 2 evidence form, then complete all remaining evidence from Modules 1–3 in Module 1 if you have not already. When DV01–DV04 are complete, open Project 1, download its Canvas Packet zip, and upload it to Canvas.

Interactive assignment: Robustness explorer, summary-stat verifier, encoding channel lab, chart repair lab, axis-truncation explorer.

Deliverables: verified summary table with visual justification, before/after chart redesign with a written defence of scale, baseline, and encoding.

Important: “Get AI feedback →” saves formative feedback inside QuantegyAI and does not submit to Canvas. Project 1 is one final zip containing DV01–DV06 from Modules 1–3; only that packet is submitted to Canvas.

Keep it simple: use the provided data and R output, focus on what each result means, and write in a few clear sentences. No question requires long arithmetic or exact decimal matching.

Interactive reading

Five short claims about summaries and encodings. Select True or False before opening each explanation, then read the explanation either way. Reading ahead is allowed; the check simply waits for your answer.

Terms used in these checks.

Mean and median
The mean is the arithmetic average; the median is the middle value once the data are sorted. They agree for symmetric data and separate when there is a long tail.
Standard deviation (SD)
A measure of spread built from every observation’s distance to the mean, so one extreme value inflates it.
Interquartile range (IQR)
The distance between the 25th and 75th percentiles: the span of the middle half of the data.
Encoding
The rule that turns a number into something visible — a position, a length, an area, an angle, or a color.
Baseline
The value a bar’s length is measured from. Bars carry magnitude by length, so their baseline is normally zero.
Denominator
The number a rate is divided by. A rate without its denominator hides how much evidence stands behind it.
Before the checks · four ideas
  1. Which summary resists an extreme value. Take the values 4, 6, 7, 9 and 40. The mean is 13.2, above four of the five values, because the 40 pulls it upward. The median stays at 7, because moving the largest value further out does not change which value is in the middle. The median is robust to one extreme value; the mean is not.
  2. Summaries do not determine shape. The mean and SD do not determine a distribution’s shape. Two datasets can share the same mean and the same SD while one is symmetric with a single peak and the other has two separate clusters. Only a plot of the distribution, such as a histogram, shows the shape.
  3. Some encodings are read more accurately than others. Readers judge position along a common scale more accurately than length, and length more accurately than angle or area. When the task is to compare close values precisely, put them on a common scale.
  4. Correct labels do not make a bar honest. Readers compare bars by their length. If the axis starts above zero, every label can be correct while the bar lengths overstate the differences between the values.
Robustness · predict before reveal

A distribution has one extreme high value. Because the mean uses every observation, it is still the more defensible report of the typical value. True or false?

False. The median usually reports the typical value more defensibly because it resists one extreme value. The mean can be pulled arbitrarily far by one extreme observation; this is why statisticians say the mean has a breakdown point of zero — the fraction of the data that can be corrupted before the summary becomes arbitrary. The chart should show the outlier rather than letting it silently dominate the summary.
Shape · predict before reveal

Two datasets can have the same mean and the same standard deviation and still have different shapes, so a histogram may tell a different story from the summary table. True or false?

True. The same mean and standard deviation can hide different shapes, clusters, gaps, skew, or outliers. A responsible chart-choice defense checks the full distribution, not only the summary table.
Encoding · predict before reveal

A chart uses bubble area to compare six close pass rates. For a precise comparison of close values, area is usually the clearest encoding. True or false?

False. People compare position on a common scale more accurately than area. If the task is precise comparison of similar values, a dot plot or aligned bar chart is usually easier to defend than area-based marks.
Axis scale · predict before reveal

A bar chart compares rates of 92%, 94%, and 96%, and its vertical axis starts at 90%. Because every label on the chart is correct, it cannot mislead a reader. True or false?

False. The truncated baseline visually exaggerates small differences even though every label is correct: readers compare bar lengths, and the lengths no longer match the values. Truncation is not always forbidden, but the designer must make the scale obvious, and for bars the honest repair is usually to change the mark to a dot plot.
Denominator · predict before reveal

One program has 18 completions out of 20 attempts, and another has 180 completions out of 240 attempts. Judged by completion rate rather than by raw count, the smaller program did better. True or false?

True. The raw count favors the larger program, but the rate tells a different story: \(18/20 = 90\%\) and \(180/240 = 75\%\). A defensible visualization states whether the question is about total volume, rate, or both, and publishes the denominator with the rate.
When the reading checks are complete: Continue → Why distributions matter

Why distributions matter

From one number to the whole pattern

A distribution is the full pattern of values in a variable — not just its average. It shows where observations cluster, how far they spread, whether the shape is symmetric or skewed, whether there are gaps or multiple peaks, and whether unusual values deserve explanation rather than deletion.

This matters because a single summary can hide the very feature the audience needs to understand. Two classes can have the same mean score but very different spreads; two campuses can have the same median pass rate but different tails; a small group can look extreme because its denominator is unstable. Before students choose a chart, they should ask what the distribution is licensing them to claim.

In the Module 2 graded assignment, distribution thinking becomes the bridge between statistics and design: histograms and density views reveal shape, boxplots emphasize median and spread, dot plots preserve individual cases for small samples, and summary tables verify whether the visual story is numerically defensible.

When you have read this: Continue → From a frequency table to a probability table

Interactive reading: from a frequency table to a probability table

Keep the arithmetic simple. Use the table or R output and focus on what the result means. Rounded values are acceptable; assessed choices do not require reproducing long decimals by hand.

The summary table you verify in this module already contains counts and proportions. Module 3 will call the proportion column an empirical probability mass function and its running total a cumulative distribution function. Those names are new for many students even when the arithmetic is not, so meet them here, on six numbers, before the Module 3 labs ask you to compute them.

Step 1 · Count, then divide by \( n \)

Suppose the observed number of attempts for six students is \( 1, 1, 2, 3, 3, 3 \). Tally each value and divide by \( n = 6 \):

Value \( x \)CountRelative frequency \( \hat{p}(x) = \dfrac{\text{count}}{n} \)
12\( 2/6 \approx 0.333 \)
21\( 1/6 \approx 0.167 \)
33\( 3/6 = 0.500 \)
Total6\( 6/6 = 1 \)

The third column is the empirical PMF: for each value, the share of the data that took that value. Two checks make it a valid probability table — every entry lies between 0 and 1, and the entries sum to exactly 1. The mean you would compute from the raw data is the same number you get by weighting each value by its probability: \( 1(2/6) + 2(1/6) + 3(3/6) = 13/6 \approx 2.17 \), which is exactly \( \bar{x} = 13/6 \). A probability table does not lose the mean; it reorganises it.

Step 2 · Accumulate to get the CDF

The CDF answers “what share of the data is at most this value?”: \( \hat{F}(x) = P(X \le x) \), a running sum down the PMF column.

Value \( x \)PMF \( \hat{p}(x) \)CDF \( \hat{F}(x) = P(X \le x) \)
1\( 2/6 \)\( 2/6 \approx 0.333 \)
2\( 1/6 \)\( 3/6 = 0.500 \)
3\( 3/6 \)\( 6/6 = 1 \)

Read it two ways. Numerically: \( P(X \le 2) = \hat{F}(2) = 0.5 \), so half the students needed at most two attempts, and \( P(X = 3) = \hat{F}(3) - \hat{F}(2) = 1 - 0.5 = 0.5 \) recovers the PMF entry as a difference of neighbours. Visually: a PMF is drawn as a bar chart whose bar length is the probability — so this module's zero-baseline rule applies with full force — while a CDF is a step function that climbs from 0 to 1 and never decreases. If your CDF ever goes down, or ends anywhere but 1, the table underneath it is wrong.

Step 3 · Empirical table versus theoretical model — not the same object

In Module 2 you only need the empirical table, which is descriptive statistics. Keep the label straight in your writing: “the proportion of students needing three attempts was 0.50” is an empirical statement about this sample; “the probability a student needs three attempts is 0.50” is a claim about a process, and needs a model or a much larger sample behind it.

# Base R only — press Run in the lab below to run it (same code, preloaded)
x <- c(1, 1, 2, 3, 3, 3)
counts  <- table(x); counts                  # 1:2  2:1  3:3
pmf_hat <- prop.table(counts); round(pmf_hat, 3)   # empirical PMF: 0.333 0.167 0.500
cdf_hat <- cumsum(pmf_hat); round(cdf_hat, 3)      # empirical CDF: 0.333 0.500 1.000
sum(as.numeric(names(pmf_hat)) * pmf_hat)    # mean from the table = 13/6 = 2.167 = mean(x)
ecdf(x)(2)                                   # P(X <= 2) read from the empirical CDF = 0.5
fair_die <- setNames(rep(1/6, 6), 1:6); round(fair_die, 3)   # a theoretical model, no data used

Module 3 extends this with expected value and variance from a table, and with the Poisson and binomial models you compare an empirical table against. If those names are new, the Module 3 worked example walks through every step with the same notation.

The die-compare lab needs JavaScript.

R / Quarto / Shiny workflow

See the output yourself: press Run and R executes the code above in your browser — the printed table and the plot appear below the code. The starter is complete as provided; edit it only if you want to test a different idea.

When you have read this: Continue → Two-coin lab

Two-coin lab: run the distinction yourself

Why this lab Run the distinction between independent events yourself.

The reading above named the difference between an empirical probability table (relative frequencies counted from data you actually have) and a theoretical probability model (probabilities written down from assumptions before any data exist). This lab makes you run both on the same random process — tossing two fair coins — so the distinction is something you operated, not just read. The sample space is the list of every equally likely outcome of one toss; counting it correctly is the whole prediction. Then the R scaffold below turns it into project evidence.

Before the prediction · count paths, not values

Roll two dice and record the sum. The sums 2 to 12 are eleven possible values, but they are not equally likely. The equally likely outcomes are the 36 ordered pairs (first die, second die). A sum of 2 arises from one pair, (1, 1), so \( P(\text{sum} = 2) = 1/36 \). A sum of 7 arises from six pairs, (1, 6), (2, 5), (3, 4), (4, 3), (5, 2) and (6, 1), so \( P(\text{sum} = 7) = 6/36 \). The pairs (1, 6) and (6, 1) count separately because the two dice are different objects.

So list the equally likely paths first, then count how many paths give each value. A value’s probability is the number of its paths divided by the total number of paths.

The two-coin probability lab needs JavaScript.

When this activity is complete: Continue → Probability R scaffold

Probability R scaffold

R / Quarto / Shiny workflow

Simulate the same two-coin process in R: a theoretical probability table written from the fair-coin model, an empirical probability table computed from your simulated tosses, and one graph comparing the two.

Three simple steps
  1. Click Run. The code is complete as provided.
  2. Compare the empirical bars with the theoretical diamonds. Rounded values are enough.
  3. Write three short sentences: what the model predicts, what the simulation produced, and why a small gap is ordinary sampling variation.
Quick check: both probability rows should total 1. The empirical bars may differ slightly from the model diamonds; that difference is expected.
When the R scaffold has run: Continue → Summary statistics lab

Summary statistics lab

Why this lab See how mean and spread describe the same data differently.

Numerical summaries are the structure a chart is built on. Before choosing an encoding, establish which summaries this distribution licenses — and check that the ones you were handed are actually right.

The summaries you will move, then verify. The explorer asks which of these follow one extreme value; the verifier asks you to recompute each one from nine numbers.

Mean
Add every value and divide by the count \( n \). One extreme value can pull it a long way.
Median
The middle value once the data are sorted (with \( n = 9 \), the fifth value). An extreme value cannot move it past its neighbours.
Mode
The value that occurs most often. It is not the smallest value.
Range
A distance: the largest value minus the smallest, not the largest value on its own.
IQR
The third quartile minus the first quartile, the span of the middle half. Quartile conventions differ between software packages, so name yours.
Standard deviation
For a sample, \( s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}} \). Dividing by \( n \) instead of \( n-1 \) gives a smaller, biased value.
Robust
A summary is robust when there is a ceiling on how far one observation can move it. Median and IQR are robust; mean and SD are not.

The summary statistics lab needs JavaScript.

When this lab is complete: Continue → Encoding channel lab

Encoding channel lab

Why this lab Match each message to the visual channel that carries it best.

Position, length, angle, area, hue, lightness — readers decode these with very different accuracy, and each one carries its own promise about the data. Match the task to the channel it licenses.

The six channels, from most to least accurately read. Each task in the lab is best served by exactly one of them.

Position on a common scale
Where a mark sits along a shared axis — a dot plot or scatterplot. The most accurately decoded channel; it carries no promise about zero.
Length from a zero baseline
How long a bar is. Readers compare lengths as ratios, so the bar must start at zero for “twice as long” to mean “twice as much”.
Angle or slice
The angle of a pie slice. Read least accurately of the quantitative channels, and worse as slices multiply.
Area
The size of a circle or square, used when position is already spent (for example on a map). Area grows with the square of the radius.
Colour hue
Red versus blue versus green. Separates categories without implying that one is larger than another.
Colour lightness
Light to dark within one hue. Reads as “less” to “more”, so it can carry an ordered scale.

The encoding channel lab needs JavaScript.

When this lab is complete: Continue → Chart repair lab

Chart repair lab

Why this lab Repair a broken chart so it tells the truth accurately.

Nine flawed charts, each with a tempting fix that leaves the defect in place. Diagnose before you repair — the obvious correction is often the wrong one.

Terms the repair options use.

Truncated axis
An axis that starts above zero. Harmless for dots and lines when disclosed; it breaks a bar’s length encoding.
Dual y-axis
Two series drawn against two different vertical scales on one chart. The relative scaling is the analyst’s free choice, so any apparent relationship can be manufactured.
Log scale
An axis where each tick multiplies the previous one (1, 10, 100). Right for multiplicative growth, but it must be labelled with real values and named as logarithmic.
Sequential palette
A single hue running from light to dark, used for ordered quantities. A rainbow is not perceptually ordered.
Data-ink ratio
The share of the ink on a chart that carries data rather than decoration. Gradients, shadows, and heavy borders lower it.
Overplotting
So many overlapping points that density and unusual cases disappear. Transparency, jitter, or binning restores them.
Visual hierarchy
Giving the element that answers the question the strongest contrast and muting the context around it.
Small multiples
A grid of small panels, one per series or group, drawn on the same shared scales. Each series keeps an honest axis of its own, and the panels can still be compared side by side.

The chart repair lab needs JavaScript.

When this lab is complete: Continue → Data Detective challenge

Data Detective Challenge: What’s Wrong With This Graph?

Why this lab Find every misleading choice hiding in a published graph.

Inspect the deliberately poor visualization, select every defect you can defend, and then repair the same comparison. More than one answer is correct. A defect changes what a reader takes away or blocks them from reading the chart at all; a stylistic preference (such as alphabetical ordering) does neither, and ticking one costs a mark. If the diagnosis has you stuck, the True/False option under it closes the challenge so you can move on.

Before you diagnose: five more ways a graph misleads or cannot be read. The truncated axis from the chart repair lab is one; here are the others.

Distorted proportions
A graph is honest when the drawn ratio matches the data ratio: a bar twice as tall should stand for a value twice as large. Truncated axes, 3-D effects and area marks can all break that match.
3-D perspective
Depth, tilt and side faces make nearer bars look larger and hide where a bar’s top meets the axis. Three dimensions add no data to a two-variable comparison.
Colour that carries nothing
Colour should carry data, such as group or order. Colours that encode nothing compete for attention without informing, so a reader looks for a meaning that is not there.
Missing labels
Every axis and every group needs a label. A reader who cannot tell which bar is which, or what the axis measures, cannot read the chart at all.
Legend
A legend must name what each colour means. A legend such as “Series?” leaves the colours unexplained.

The Data Detective challenge needs JavaScript.

When the diagnosis is complete: Continue → AI-output audit

AI-output audit

Why this lab Rule on AI claims before trusting them with your data.

You have verified the summaries and matched tasks to channels. Now an assistant proposes six changes to your draft chart. Rule on each one with the encoding rules you just used; the first proposal is exactly the axis truncation the reading checks warned about, and the visual lab that follows lets you test it.

Three ideas you need before you rule. The other proposals use ideas you have already met: robust summaries and dual axes.

Spread is not shape
A small standard deviation says the values sit close to their mean. A small SD says nothing about whether the distribution is normal, skewed or has two peaks; only a plot of the distribution shows its shape.
Part-to-whole
A pie shows how one whole divides into parts that sum to 100%. Rates for separate groups, such as six campuses’ pass rates, are not parts of one whole, so they do not belong in one pie.
Formatting versus data
Rounding tick labels changes how the axis reads, not the data: the stored values keep their full precision. Rounding the stored values themselves throws measurement away.

This AI audit needs JavaScript.

When every proposal has a verdict: Continue → Test the truncation claim

Hands-on visual lab

Why this lab Test the claim you refused with a visual experiment.

In the audit the assistant proposed starting the bar chart’s y-axis at 92 so the difference would show, and you rejected it. Now test that refusal directly: one design decision, six real pass rates, and a ratio that comes apart from the data. Move the baseline and watch how far a reader can be misled without a single incorrect label.

This visual lab needs JavaScript.

When you have answered the baseline question: Continue → GeoGebra lab

GeoGebra lab

Why this lab Explore how changing an encoding changes what you see.

The same truncation, drawn as a geometric object. Use GeoGebra to make scale, proportion, and visual distortion visible. The AI can draft construction commands, but you must verify the visual claim against the data.

Interactive Math with AI / GeoGebra strand

Embedded GeoGebra applet loading…

When the construction check is done: Continue → R lab

R lab

Why this lab Reproduce the descriptive evidence with real R output.

You predicted which summaries and encodings would be defensible. Now verify them computationally. Run the provided data, compare the competing views, and carry the checked result into DV03 and DV04.
R / Quarto / Shiny workflow

Compute mean, median, SD, and IQR; build multiple ggplot2 views of the same data; then write a Chart Choice Defense explaining why one visualization answers the question more honestly than another.

Browser R note: this inline lab runs base R plus ggplot2 (downloaded automatically on first run — one short wait, cached after that; if the download is blocked the lab status says so and Run retries it). The full tidyverse collection is a desktop tool for RStudio/Positron. Everything this module requires runs right here in the browser.

What you are expected to add to this scaffold

The scaffold below already runs end to end — that is deliberate, so the first thing you see is working output, not an error. Every code block is provided and runs as pasted. Output notes tell you what to look for before submitting. Your Project 1 evidence comes from what you add:

  1. Verified summaries — keep step 1, and use the provided spread output containing the \( \mathrm{IQR} \) and standard deviation. Paste the printed values into Numerical summaries and say which summary the distribution licenses (median or mean, and why).
  2. The provided third view of the same data, saved as p_third — so that you can show how one encoding or baseline change alters the reader's impression while the data stay fixed.
  3. The Chart Choice Defense (step 4) — 100–150 words naming which chart answers the question honestly, which claim the other view would invite, and one limitation (six campuses, unequal tested counts). This is the Final defense field of the graded assignment.
  4. The AI audit — if you asked an assistant which chart to use, paste the prompt and its recommendation, then state what you verified and corrected. This is the AI audit field.

Which package: base R for every summary (summary(), mean(), median(), IQR(), sd()) and ggplot2 for every plot. Nothing in this module needs dplyr, tidyr, readr, or tibble, and those are not installed in browser R — if an assistant suggests them, that is a finding for your AI audit, not a reason to install anything.

Use the Copy code and Copy output buttons on the inline lab: the code goes in your reproducibility section, the printed output is your evidence.

# Module 2 R scaffold — complete and ready to run
# Use the output to make a chart-choice decision; do not chase exact decimals.
library(ggplot2)      # the only package this module needs; base R does the summaries

## PROVIDED — the dataset. One row = one campus; pass_rate = passed / tested.
scores <- data.frame(
  campus = c("North", "South", "East", "West", "Central", "Online"),
  pass_rate = c(0.92, 0.94, 0.95, 0.96, 0.93, 0.95),
  tested = c(84, 122, 76, 98, 111, 64)
)

## PROVIDED — step 1: verify the summaries before you visualize anything
summary(scores$pass_rate)
## EXPECTED   Min. 0.9200   1st Qu. 0.9325   Median 0.9450   Mean 0.9417   3rd Qu. 0.9500   Max. 0.9600
cat("Mean pass rate:", round(mean(scores$pass_rate), 3), "\n")      # 0.942
cat("Median pass rate:", round(median(scores$pass_rate), 3), "\n")  # 0.945
cat("Range:", paste(range(scores$pass_rate), collapse = " to "), "\n")  # 0.92 to 0.96

## PROVIDED — step 1b: spread, rounded for interpretation
spread <- c(iqr = IQR(scores$pass_rate), sd = sd(scores$pass_rate))
round(spread, 2)  # sensible rounding is enough; no four-decimal matching required

## PROVIDED — step 2: full-baseline bar chart (length from zero — the honest bar)
p_bar <- ggplot(scores, aes(x = reorder(campus, pass_rate), y = pass_rate)) +
  geom_col() +
  coord_flip() +
  scale_y_continuous(limits = c(0, 1)) +
  labs(title = "Full-baseline bar chart", x = "Campus", y = "Pass rate")

## PROVIDED — step 3: dot plot on an annotated narrow range (position, not length)
p_dot <- ggplot(scores, aes(x = pass_rate, y = reorder(campus, pass_rate))) +
  geom_point(size = 3) +
  scale_x_continuous(limits = c(0.90, 0.97)) +
  labs(title = "Dot plot for close pass-rate comparisons", x = "Pass rate (axis runs 0.90-0.97)", y = "Campus")

print(p_bar)
print(p_dot)
## EXPECTED   two plots: six bars that look almost equal, then six dots spread across the axis

## PROVIDED — step 3b: the misleading comparison view, ready to run
p_third <- ggplot(scores, aes(x = reorder(campus, pass_rate), y = pass_rate)) +
  geom_col() +
  coord_cartesian(ylim = c(0.90, 0.97)) +
  labs(title = "Truncated bar chart (what not to publish)", x = "Campus", y = "Pass rate")
print(p_third)
## INTERPRET   the values did not change, but the truncated bars exaggerate the gap

## WRITE — three short sentences in the Final Chart Choice Defense field:
## 1. which chart answers the question honestly; 2. what the other view exaggerates;
## 3. one limitation, such as unequal tested counts.
How to check your work before you submit
  • The lab status reads “R ran successfully” and there is no red Error line in the output. If library(ggplot2) fails, the status tells you it is a download problem and Run R retries it; the base-R summaries are still valid evidence.
  • Every ## EXPECTED value matches your console to the printed rounding. Run sum(scores$tested) — it must print 555; anything else means the dataset was edited.
  • spread and p_third exist: exists("spread") and exists("p_third") both print TRUE.
  • Your defense names the encoding channel (position versus length), the baseline decision, the axis range you disclosed, and one limitation. The graded assignment's rubric checks for exactly those four things.
Prediction checked against the data. Use that verified evidence next. Continue → Build DV03 & DV04

Project 1 evidence: DV03 and DV04

You are not doing separate assignments here.

DV03 and DV04 are the evidence you will use in the Module 2 evidence form below. Complete and save both once; your work carries forward.

When DV03 and DV04 are saved: Continue → Get AI feedback

Module 2 evidence and AI feedback

You have already done most of the work.

Now package the verified summaries, competing charts, and AI-audit reasoning you produced above. This is one evidence-review step, not another assignment and not a Canvas submission.

Where do I get my data?

Use scores. Everything you need is already provided in the R lab. It contains the verified campus summaries needed for DV03 and DV04.

Advanced option: Use your own dataset

You may use a built-in R dataset or add a CSV under Your data in the R lab. The provided scores dataset remains the recommended route.

How does grading work?

The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.

Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit Project 1 to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like. If a long textbox is the only thing in your way, a short True/False concept check can close it.

Before the decision check · match the display to the question

The inline Module 2 graded assignment and AI grader need JavaScript.

After you review the AI feedback: Continue → Mastery check

Prepare for the mastery check

Adaptive reflection · What I know / where I go next

Mastery check loading…

What Module 2 saved

Your saved Module 2 evidence needs JavaScript. You can also see it later on your course report card.

After Module 2 Add this module’s evidence to the Project 1 packet. If your explanation depends on AI, include the prompt, output, verification method, and final correction.
When the mastery check is saved: Continue → Project 1 Canvas Packet