Module 1 — Questions, Data & Cleaning
From statistical question framing to data dictionaries, cleaning logs, and missingness audits.
Every data visualization begins with a question, and every honest chart depends on data that has been inspected before it is drawn. Module 1 covers those two foundations. In the first week we learn to frame a statistical question precisely: name the population, the observational unit, and the variables, decide whether the question is descriptive, comparative, relational, or causal, and match the claim we intend to make to the study design that produced the data. In the second week we turn to the data itself: build a data dictionary, diagnose why values are missing and whom the missingness affects, separate unit and entry errors from genuine outliers, and record every cleaning decision with its limitation so a reviewer can retrace it.
We practise these decisions in short interactive labs, verify them in the browser R lab, and then record two Project 1 evidence items: DV01, a statistical question and data contract, and DV02, a cleaning and missingness audit. The module mastery check confirms both skills. By the end we should be able to defend not only what a dataset shows but whether it was ever able to answer the question we asked of it. Every lab explains its terms before it asks a question, and if a question has us stuck a True/False option closes it so we can move on.
By the end of this module, you will be able to
- Frame a statistical question by naming its population, observational unit, and variables.
- Match the claim type — descriptive, comparative, relational, or causal — to the study design that produced the data.
- Build a data dictionary and data contract that fix what every variable means and where inference must stop.
- Classify data defects instead of silently deleting rows, and separate unit and entry errors from genuine outliers.
- Diagnose missingness: describe its pattern and whom it affects.
- Log every cleaning decision, including AI advice you accepted, corrected, or rejected, so a reviewer can retrace your steps.
- Apply all of this to Project 1: your DV01 Statistical Question & Data Contract and DV02 Cleaning & Missingness Audit, due the date on Canvas.
Your goal
Turn messy data into a defensible statistical question and cleaning plan.
What you’ll learn
Statistics focus: variables, observational units, sampling frames, missingness, outliers, and defensible inference boundaries.
Tool focus: Inline browser R labs with base R and ggplot2; no installation is required. An optional desktop RStudio/Positron setup with tidyverse is available for the Project 1 assignment, but nothing here requires it.
- Turn a vague topic into an answerable statistical question naming the population, observational unit, variables, time frame, and inference boundary.
- Classify each variable’s type and measurement level and state which summaries and visual encodings it makes legal.
- Distinguish unit errors, implausible values, duplicates, and genuine outliers — and justify a different repair for each.
- Diagnose a missingness pattern and explain when complete-case deletion is safe and when it biases the estimate.
- Verify an AI-proposed cleaning protocol step by step, recording Accept/Modify/Reject verdicts with evidence.
- Assemble a data dictionary and cleaning log another analyst could audit and reproduce.
About your submission
Project 1 evidence: complete DV01 and DV02, then request formative AI feedback before you submit the final Project 1 packet to Canvas.
Your packet includes: a data dictionary, cleaning log, missingness explanation, and an AI accepted/rejected table.
What is submitted where: Project 1 is one final zip containing DV01–DV06 from Modules 1–3, built on the Project 1 page and uploaded to Canvas. Requesting AI feedback on this page saves formative coaching inside QuantegyAI and does not submit to Canvas.
Keep it simple: use the provided data and R output, choose the defensible category or repair, and explain the decision in a few clear sentences. No question requires long arithmetic or exact decimal matching.
Statistical question formation
Academic anchor: before any chart or cleaning decision, students must turn a vague topic into a researchable statistical question by naming the population, observational unit, variables, time frame, and inference boundary.
- Observational unit. The unit is what one row of the data represents. If the question is about students but one student can appear in several rows, for example one row per tutoring session, then counting rows as if each were a separate student inflates the evidence. State the unit before any count, average or rate.
- Inference boundary. The boundary says who the data represent and which claims the design can support. A convenience sample, such as the students who happened to join one online group, describes itself; it cannot stand in for a wider population.
- Association versus cause. Only a design with a randomized or well-matched comparison group supports a causal claim. Two variables that move together in observational data support an association. A before-and-after change in the same students, with no comparison group, supports a description of the change, not a claim about what caused it.
The statistical question formation builder needs JavaScript.
Interactive reading
Five short claims about question framing, observational units, variable types, cleaning, and inference boundaries. Select True or False before opening each explanation, then read the explanation either way. Reading ahead is allowed; the check simply waits for your answer. Module 1 is about turning messy ideas into defensible statistical questions before AI or visualization enters the workflow.
Two more terms these checks use. The observational unit and the inference boundary are defined in the step before this one.
- Categorical variable
- A variable whose values are labels, such as a campus name or a practice path. Labels are counted, and compared as proportions; they are never averaged, because a label has no numeric value to average.
- Complete-case deletion
- Dropping every row that has a missing value. It is safe only when values are missing for reasons unrelated to the data. When missingness is related to the outcome, for example when weaker students skip the post-test more often, deleting incomplete rows can bias the results.
A student asks: “Does screen time affect grades?” Naming the population it is about is enough to make this a defensible statistical question. True or false?
A spreadsheet has one row per tutoring session, but the research question asks about students. If one student can appear in several session rows, summarizing the rows as if each were one student inflates the evidence. True or false?
A column records practice path as “Algebra,” “Geometry,” or “Statistics.” An AI assistant may report the mean of this variable. True or false?
An AI assistant says: “Delete every row with a missing value before visualizing.” That is a safe default. True or false?
A small convenience sample from one online practice group shows a higher completion rate after a reminder. The strongest defensible claim describes this sample and flags an association worth studying, but stops short of a population-wide or causal claim. True or false?
Question triage lab
Why this lab Sort questions by what they ask of the data.
Visualization begins with a statistical question, not a chart type. Sort each statement into what it actually is — and notice how often a causal verb sneaks into a design that cannot support it.
Before you sort, here are the five categories. Every statement in the lab belongs to exactly one of them.
- Topic (not yet a question)
- A subject area with no population, unit, or measured variable. No dataset can answer it yet.
- Descriptive question
- Summarises one variable in one population: a median, a proportion, a typical value.
- Comparative question
- Sets two or more defined groups side by side on one outcome and asks whether they differ.
- Relational question
- Asks how two variables measured on the same units move together — “is associated with”.
- Causal claim
- Asserts that one thing produces a change in another (“raises”, “improves”). Only a design built for it can support it.
If a statement has you stuck, use the True/False option that appears under it and move on.
The question triage lab needs JavaScript.
Types of data interactive lab
Why this lab Match each variable to its data type so the right checks apply.
Before choosing a visualization, classify the variable correctly. The data type determines what summaries are legal, which chart encodings are honest, and what an AI-generated claim is allowed to say. Then build the data dictionary yourself — it is the contract every later chart in your project has to satisfy.
The data-types lab needs JavaScript.
Dirty-data audit
Why this lab Find every defect that would poison a real analysis.
A real extract arrives with inconsistent labels, missing values, duplicates, implausible entries, and unit errors mixed together. Diagnose each row before repairing it: the wrong diagnosis produces the wrong repair, and a deleted row is evidence you can never get back.
The six labels you will choose from. Read the row against the rest of its column, then name what is wrong with it.
- Inconsistent coding or format
- The same thing written two ways — “retaker” beside “Retaker”, or a date in a different format. Software treats them as different values.
- Missing value
- A cell that was never recorded (shown as
NA). Whether it can be dropped depends on why it is missing. - Duplicate record
- The same row entered twice. It inflates the sample size without adding information.
- Implausible value
- A number that cannot be right for this variable and has no unit that explains it — most likely a typing slip.
- Unit mismatch
- A number recorded in the wrong unit, such as minutes in an hours column. The value is real; the unit is wrong.
- No defect
- The row is clean. Certifying clean rows matters: an audit that only ever finds faults over-cleans.
The dirty-data audit needs JavaScript.
AI-output audit
Why this lab Rule on AI proposals before trusting them with your data.
An assistant proposes a six-step cleaning protocol. Rule on every step and build the “AI suggestions accepted / rejected” table that ships with your data dictionary and cleaning log.
Two ideas you need before you rule. The other proposals use the defect types from the dirty-data audit.
- Mean imputation
- Replacing every missing value with the column mean. It invents values that all sit exactly at the centre, so it understates the spread: the standard deviation shrinks and every interval built from it is narrower than the data earned. It also hides which values were never recorded.
- Rounding for display
- Rounding the stored or recorded values throws away measurement for good. To make a chart easier to read, round only the axis and label formatting and keep the data at full precision.
The AI cleaning-log audit needs JavaScript.
Hands-on visual lab
Why this lab Test the claim you refused with a visual experiment.
In Lab 4 the assistant claimed that dropping every incomplete row always improves accuracy, and you refused it. Now test that refusal directly: hold the data fixed, change only the reason values go missing, and watch the complete-case estimate drift away from the truth.
Three terms the readout uses.
- Complete-case deletion
- Keeping only the rows with no missing values and dropping the rest. It is what “delete every row with a missing value” means.
- Missing completely at random (MCAR)
- Values go missing for reasons unrelated to anything in the data. Deletion then costs precision (a smaller sample) but does not shift the estimate.
- Systematic missingness
- Values go missing more often for some students than others — here, more often for low pre-test scores. Deletion then removes a particular kind of student, and the estimate shifts. A bigger sample repeats the shift instead of fixing it.
What to do: write your prediction, then drag the slider past 5 and compare the complete-case mean with the true mean.
The missingness explorer needs JavaScript.
Verify with R
Why this lab Reproduce the cleaning result with real R output.
scores, identify which variables contain missing values, and compare the output with your earlier reasoning.Inspect the provided scores dataset (student pre/post with missing values), count missingness by variable, and defend a cleaning decision.
Browser R note: this inline lab runs base R plus ggplot2 (installed automatically on first run). The full tidyverse collection is a desktop tool — use RStudio or Positron for it. Everything this module requires works right here in the browser.
The scores dataset and the full audit sequence — inspection, missingness counts, complete pairs, and the plot — are already loaded in this workspace. Press Run, then write three short sentences: what is missing, the rule you used, and one limitation. Log any AI assistance and corrections.
Project 1 evidence: DV01 and DV02
DV01 and DV02 are the evidence you will use in the Module 1 evidence form below. Complete and save both once; your work carries forward.
Module 1 evidence and AI feedback
Now package the question, data decisions, R evidence, and AI-audit reasoning you produced above. This is one evidence-review step, not another assignment and not a Canvas submission.
How does grading work?
The AI score is formative feedback, not your grade. Use it to find gaps, revise, and request updated feedback — you have unlimited attempts, and only your best work matters. A low AI score simply means the rubric wants more detail or evidence somewhere; it is never locked in.
Your official grade comes from Dr. Roberts's review of your Canvas packet. The AI grader exists to coach you toward a complete, defensible packet before you submit Project 1 to Canvas — read its feedback, strengthen the weak spots, and request updated feedback as many times as you like.
Where do I get my data?
Use scores. Everything you need is already provided in the Verify with R workspace. It is the small pre/post score dataset with built-in NA values, quantitative variables, a categorical identifier, and missingness to audit.
Step by step:
- Continue to Verify with R →
- Press Run — the
scoresdataset and the audit code are already loaded in that workspace. - Run
str(scores)andsummary(scores)to inspect the structure. - Run
colSums(is.na(scores))to get missingness counts per variable. - Build your data dictionary and dirty-data audit from the output.
- Record your results in the Module 1 evidence fields below and select Get AI feedback →.
Advanced option: Use your own dataset
You may use a built-in R dataset such as mpg, storms, or diamonds, or add a CSV under Your data in the R lab. The default scores path is already complete and remains the recommended route.
This assignment turns the Module 1 labs into a Project 1 evidence packet. You'll refine one statistical question, classify the variables, audit the data quality, and document which AI cleaning suggestions were accepted or rejected.
Required evidence
The six required items are listed as the Module 1 evidence checklist below — tick each one as your evidence is ready, and the status line tells you when you can request AI feedback.
Mastery evidence: the final packet should show that the student can move from a vague topic to a defensible question, from raw variables to a usable data dictionary, and from AI-generated suggestions to verified cleaning decisions.
The inline Module 1 graded assignment and AI grader need JavaScript.
Prepare for the mastery check
Mastery check loading…
What Module 1 saved
Your saved Module 1 evidence needs JavaScript. You can also see it later on your course report card.