Module 8 — Reproducibility, AI Audit & Statistical Defense
Make your statistical story independently reproducible, fully auditable, and defensible under challenge—then assemble the strongest evidence from Modules 1–8 into the later course synthesis and final project.
Student learning outcomes
- Package data, code, seeds, and environment so an independent reviewer can regenerate every figure.
- Maintain an AI prompt-and-correction log that preserves the complete decision chain.
- Review a peer’s analysis against a reproducibility rubric, separating different failure types that need different repairs.
- Read a specification curve and report how a conclusion depends on defensible analytic choices.
- Assemble the five-phase evidence spine the final project composes.
Interactive reading
Select True or False before opening each explanation. Module 8 distinguishes work that merely looks finished from evidence another person can reproduce and defend.
A prompt log on its own is not enough disclosure for graduate-level AI-assisted statistical work. True or false?
A fixed random seed does not make a report reproducible when the raw data are absent. True or false?
An AI-generated confidence interval is plausible and agrees with the chart, so it has been verified. True or false?
When four defensible choices give four different estimates, reporting the preferred specification alone is sufficient. True or false?
When a reviewer challenges a causal claim, the strongest reply is to show how precisely the effect was estimated. True or false?
Peer-review rubric lab
Open the DV15 peer-review rubric lab · about 15 minutes
The peer-review rubric lab needs JavaScript.
Defense rehearsal
Open the DV16 statistical-defense rehearsal · about 15 minutes
The defense rehearsal needs JavaScript.
AI-output audit
Open the reproducibility AI-output audit · about 10 minutes
This AI audit needs JavaScript.
Hands-on visual lab
Open the specification-curve visual lab · about 12 minutes
One dataset and four binary analytic choices produce sixteen illustrated specifications. Each choice needs justification; the normal approximation and mean imputation have limitations. Find your specification on the curve, then ask what reporting only that one would tell a reader.
This visual lab needs JavaScript.
R lab
Open the runnable reproducibility R labs · about 15 minutes
Finalize a reproducible Quarto report or Shiny artifact with source code, prompt log, verification notes, and statistical defense.
R Lab A — DV15: Source lineage
Your task The editor opens with the raw scores frame already loaded. Press Run: the lab writes scores-v1.csv, reads it back, and applies one documented exclusion rule, to demonstrate a traceable transformation of a synthetic teaching fixture. Its fixed recording date is fictional, not a retrieval date. Excluding scores below 50 is illustrative and needs an independent justification.
Lab complete when you can identify the synthetic source, fixture version, fictional recording date, and exclusion rule. Writing a CSV does not establish external provenance.
Press Run: the lab writes and reads back the versioned CSV, then prints the synthetic source file, fictional recording date and the rows the illustrative exclusion drops. Nothing needs copying.
R Lab B — DV16: Specification defense
Your task The editor opens with the same frame and the four analytic choices already loaded. Press Run: the lab computes four Before-versus-After descriptive contrasts. Means, medians and trimmed means summarize different quantities; exclusions change the population summarized.
Lab complete when you can state which specification you report, and what reporting only that one would hide from a reader.
Press Run: the lab prints the four estimates and the spread across them. Their range is not a confidence interval. Justify the target summary and any exclusion independently of the result.
Project 5 evidence: DV15 and DV16
Open the DV15 and DV16 Project 5 evidence workspace
Package the reproducibility and AI-audit evidence that will govern the final project.
Where do I get my data?
Best route: reuse data and code from an earlier module assignment. This module is about reproducibility, so taking an analysis you already did and packaging it until someone else can regenerate every result from source is the point of the exercise.
Starting from the module's own data? The R lab editor holds claim_data and its seeded bootstrap, which is enough for the uncertainty and clean-run evidence. It is a synthetic teaching fixture defined by the script; no external retrieval occurred. Preserve the script version as its source. Two dedicated labs illustrate file lineage and changes to the summary: R Lab A — DV15 opens with the versioned-CSV lineage preloaded — it writes and reads back scores-v1.csv with a fictional recording date and one illustrative exclusion — and R Lab B — DV16 opens with the same frame plus the four-specification comparison. Press Run; nothing needs copying.
Want to see the wider picture first? The specification-curve lab shows the sensitivity of sixteen illustrated specifications, each requiring justification.
Prefer your own dataset? You can use any built-in R dataset or add your own CSV under Your data in the R lab (files up to 500 MB stay in your browser). But reusing a prior module's data is ideal — it shows you can make existing work reproducible.
Project 5 evidence items: DV15 and DV16
Module 8 AI-feedback workspace
Open the Module 8 AI-feedback workspace
The Module 8 workspace needs JavaScript.
Final project integration planner
Open the later final-project planning extension · optional
Select the completed evidence that will carry the final argument. The planner checks that the project spans foundations, modeling, advanced visual evidence, ethics, and reproducibility.
The final project integration planner needs JavaScript.
Adaptive reflection
Mastery check loading…
Module 8 record
Your saved Module 8 record needs JavaScript. You can also see it on your skills dashboard report card.