← All Data Visualization with AI modules

Uncertainty & Simulation Workspace

QuantegyAI Module 5 · Canvas Project 3 · Due the date on Canvas

Make uncertainty visible through confidence intervals, bootstrap intervals, simulation error, Monte Carlo convergence, and probability visuals that separate defensible inference from false certainty.

This Module 5 workspace is a guided path: 24 short pages, one idea each. Five short lessons teach the five reading checks; every other page ends with a guided check. Start the uncertainty and simulation path →
What to do for Canvas Project 3
  1. Work through Module 5 and save its required evidence and files.
  2. Complete the associated workspaces: Module 5. Previously earned completion remains valid; additional work is optional review.
  3. Open Project 3, check the packet status, and optionally download Project_3_Canvas_Packet.zip. For Canvas, upload at least one artifact from your work or a screenshot of your grade. The full ZIP is optional. A download is not a Canvas submission. Submit by the date on Canvas, only if you have not already submitted.

QuantegyAI feedback does not submit to Canvas. Module numbers, evidence identifiers, and Canvas weekly resource numbers are distinct.

Statistics focusstandard error, confidence intervals, bootstrap intervals, simulation, sampling distributions, law of large numbers
R/tool focusInline browser R labs; optional desktop RStudio/Positron for capstone; bootstrap workflows; infer; simulation loops or purrr; interval and convergence plots
Interactive assignmentConfidence interval cards, coverage explorer, Monte Carlo simulation lab, bootstrap interval defense, AI simulation audit

What students will cover

Review the Module 5 topic map
  • Confidence intervals: method-level coverage, interval width, standard error, and the difference between data spread and estimate precision.
  • Bootstrap reasoning: resampling logic, bootstrap distributions, percentile intervals, and what bootstrap intervals can and cannot repair.
  • Monte Carlo simulation: trial count, seed, running estimate, simulation standard error, convergence, and the cost of reducing error.
  • Central Limit Theorem: sampling distributions, why sample means become approximately normal as n grows under independent, identically distributed sampling with finite positive variance, and the relationship between the CLT and standard error.
  • Probability visuals: coverage plots, interval bands, uncertainty ribbons, and convergence charts with honest captions.
  • AI audit: checking AI-generated uncertainty claims for overconfidence, missing denominators, unsupported probability language, and unreported simulation error.

Student learning outcomes

Review the learning outcomes and evidence standard
  • Interpret the 95% in a confidence interval as long-run coverage of a method, not a probability attached to one fixed parameter.
  • Distinguish the spread of individual observations from the uncertainty in an estimated mean or proportion.
  • State the Central Limit Theorem and explain why the sampling distribution of the mean becomes approximately normal as sample size grows under independent, identically distributed sampling with finite positive variance, even when the population is not normal.
  • Build and interpret uncertainty visuals that show intervals, bands, or simulated sampling variation.
  • Choose uncertainty visuals deliberately — interval plots, distribution plots, histograms, fan charts, error bars, or scenario bands — matched to what the claim needs.
  • Identify visualizations that hide uncertainty or create false confidence, and correct them.
  • Distinguish showing what happened from what is likely and what could happen.
  • Communicate risk and uncertainty clearly to a non-technical audience.
  • Use bootstrap or Monte Carlo simulation to estimate uncertainty and report the simulation design.
  • Explain how simulation standard error changes with the number of trials.
  • Audit an AI-generated uncertainty claim and revise it into defensible statistical language.
  • Write a final uncertainty defense that states what the evidence supports and one limitation it does not remove.

Interactive reading

This reading uses quick prediction checks before the explanation is revealed. Work each prompt before opening the answer; the goal is to build mathematically precise language for uncertainty, not just to recognize vocabulary.

Coverage · predict before reveal

A 95% confidence interval computed as \( \bar{x} \pm t^\star \dfrac{s}{\sqrt{n}} \) means there is a 95% probability that this already-computed interval contains the fixed population mean. True or false?

False in the classical interpretation. Here is where the formula comes from: \( \bar{x} \) is the sample mean; \( s \) estimates the spread of individual observations; and \( s/\sqrt{n} \) estimates how much the sample mean varies from sample to sample (its standard error). Under independent sampling with an approximately normal sampling distribution of the mean, \( (\bar{X}-\mu)/(S/\sqrt{n}) \) follows a \( t \) distribution with \( n-1 \) degrees of freedom when the population is normal; for suitable large samples this is an approximation. Choose \( t^\star \) so that the central 95% of that distribution lies between \( -t^\star \) and \( t^\star \). Rearranging \( -t^\star \leq (\bar{X}-\mu)/(S/\sqrt{n}) \leq t^\star \) gives \( \bar{X} \pm t^\star S/\sqrt{n} \).

The population mean \( \mu \) is fixed; an interval already computed either contains it or it does not. The 95% describes the method: over repeated samples under those assumptions, about 95% of the resulting intervals contain \( \mu \). A defensible sentence is: “This procedure has about 95% long-run coverage under the stated assumptions.”

Precision · predict before reveal

A large sample can produce a narrow interval for the mean even when the individual observations are widely spread out. True or false?

True. The sample standard deviation \( s \) describes the spread of individual observations. The standard error \( \mathrm{SE}_{\bar{x}} = \dfrac{s}{\sqrt{n}} \) describes the sampling variability of the estimated mean. A large sample can produce a narrow interval even when individual observations are widely spread out. Use SD, IQR, percentile ranges, or prediction intervals for data spread; use SE and confidence intervals for estimate precision.
Bootstrap · predict before reveal

A bootstrap interval with \( B = 5000 \) resamples automatically fixes biased sampling because the computer resampled the data many times. True or false?

False. Bootstrap resampling estimates variability around the empirical sample. It can approximate the sampling distribution of a statistic, but it cannot repair biased collection, missing cases, measurement problems, or a statistic that does not answer the question. A graduate-level bootstrap caption names the statistic, the resampling scheme, \( B \), the interval rule, and one limitation of the original data.
Simulation error · predict before reveal

A Monte Carlo estimate of \( \hat{p} = 0.64 \) from \( N = 100 \) trials is precise enough to report as “the probability is 64%.” True or false?

Usually false. A simulated proportion has simulation uncertainty. A useful approximation is \( \mathrm{SE}_{\mathrm{sim}} = \sqrt{\dfrac{\hat{p}(1-\hat{p})}{N}} \). With only \( N = 100 \), the Monte Carlo error is large enough that a headline percentage can imply false precision. Report the model, trial count, estimate, simulation standard error, and a convergence visual when the estimate drives the conclusion.
Error reduction · predict before reveal

To cut Monte Carlo standard error in half you need about four times as many trials. True or false?

True. Monte Carlo standard error shrinks at the rate \( 1/\sqrt{N} \), so halving it costs four times the trials — doubling them divides the error by \( \sqrt{2} \approx 1.41 \) only. More trials reduce simulation error, but they do not fix the wrong probability model, biased assumptions, or an unclear estimand.
Central Limit Theorem · the theorem in full

The theorem explains why averages can have an approximately normal sampling distribution. Open the formal statement before the sampling-distribution lab; the bootstrap and simulation checks above use different reasoning.

Central Limit Theorem (Lindeberg–Lévy). Let \( X_1, X_2, \ldots, X_n \) be independent and identically distributed random variables with mean \( \mu \) and finite variance \( \sigma^2 \). Then as \( n \to \infty \),

\[ \sqrt{n}\,\frac{\bar{X}_n - \mu}{\sigma} \;\xrightarrow{d}\; N(0,\,1) \]

Equivalently, for large \( n \), the sample mean is approximately normal:

\[ \bar{X}_n \;\stackrel{\text{approx}}{\sim}\; N\!\left(\mu,\; \frac{\sigma^2}{n}\right) \]

This holds regardless of the shape of the parent distribution, provided \( \sigma^2 < \infty \). The parent can be skewed, discrete, or bimodal — the distribution of the sample mean still converges to a normal bell. This is why many normal-theory procedures for sample means can remain approximately valid even when the underlying observations are not normally distributed, provided the sampling assumptions and sample size are appropriate.

Caveat: "Large \( n \)" depends on the parent. For a near-symmetric parent, \( n \approx 10 \) may suffice. For heavily skewed parents, \( n \geq 30 \) is a common rule of thumb, not a guarantee. Thirty is not a universal CLT threshold. For extreme skew or heavy tails, convergence is slower and you may need \( n \gg 30 \).

Central Limit Theorem lab

Open the Central Limit Theorem lab · about 10 minutes

The parent distribution is Exponential(1) — right-skewed, with \( \mu = 1 \) and \( \sigma = 1 \). Draw samples of size \( n \), compute each sample mean, and watch the sampling distribution of \( \bar{X} \) take shape. The amber curve is the theoretical \( N(\mu,\, \sigma^2/n) \) the CLT predicts. Compare the simulated SD of the means to the theoretical SE \( \sigma/\sqrt{n} \).

The Central Limit Theorem lab needs JavaScript.

R lab: CLT verification

Two different results. For independent observations with a common mean and finite variance \(\sigma^2\), the variance identity gives \(\operatorname{Var}(\bar X)=\sigma^2/n\), hence \(\operatorname{SE}(\bar X)=\sigma/\sqrt n\), exactly at each sample size. For independent, identically distributed observations with finite positive variance, the CLT instead states \(\sqrt n(\bar X-\mu)/\sigma\xrightarrow{d}N(0,1)\) as \(n\) grows. Inspect the histogram shape for the latter and compare the SE estimates separately. Increasing the number of simulated replications reduces Monte Carlo uncertainty; it does not change the exact identity.

Open the runnable CLT verification lab · about 12 minutes

Your task Write your prediction, run the script unchanged, and compare the printed simulated_SE with theoretical_SE for each sample size.

Lab complete when you explain the size and direction of each simulated minus theoretical SE difference, allowing Monte Carlo variability, and distinguish the exact SE identity from the CLT approximation. Matching rounded values is not required.

R / Quarto / Shiny workflow

Verify the Central Limit Theorem yourself in R. The script draws 5,000 sample means at each of three sample sizes from an Exponential(1) parent, overlays the normal curve the theorem predicts, and prints the simulated standard error beside the theoretical \( \sigma/\sqrt{n} \). It runs here in the browser; desktop RStudio/Positron works too if you prefer it.

Interval interpretation cards

Open the confidence-interval interpretation challenge · about 10 minutes

The interval interpretation cards need JavaScript.

Monte Carlo lab

Open the Monte Carlo convergence lab · about 10 minutes

The Monte Carlo lab needs JavaScript.

Bootstrap lab

Open the bootstrap percentile-interval lab · about 10 minutes

Bootstrap resampling is one option for your DV10 simulation evidence; Monte Carlo simulation is another. Use this lab to explore the bootstrap option. Completing both methods is not an additional evidence requirement. Resample the same 25-score sample with replacement and watch the distribution of resample means take shape; the middle 95% of that distribution is the percentile interval you report.

The bootstrap lab needs JavaScript.

AI-output audit

Open the AI uncertainty audit · about 10 minutes

Your task Rule Accept, Modify, or Reject on all six AI uncertainty claims. Modify is for a claim whose statistics are right but whose wording promises more than they support.

Lab complete when every proposal has a verdict and the audit table below the cards is filled in — that table is the AI-use log your Project 3 packet asks for.

This AI audit needs JavaScript.

Hands-on visual lab

Open the uncertainty visual lab · about 8 minutes

Your task Write the prediction, then step through the samples and count how many of the forty intervals miss the true mean.

Lab complete when you explain why individual misses can occur with a correctly calibrated procedure, and why this one set of forty cannot establish long-run coverage. An expected miss count is not a guaranteed result.

Forty samples from a population whose true mean we know, each with its own 95% interval. Correctly computed intervals may miss. For a properly calibrated 95% procedure, coverage refers to repeated experiments over the long run; any one set of forty has a variable miss count and cannot by itself establish calibration. Then run the browser-R lab below without leaving QuantegyAI.

This visual lab needs JavaScript.

R labs

Two separate labs, one per Project 3 evidence item. Each editor is already loaded with its own code and data — open the lab your evidence item names and press Run. Nothing needs installing and nothing is calculated by hand.

R Lab A — DV09: Uncertainty interval

Open R Lab A — DV09: Uncertainty interval · about 10 minutes

Your task Six steps, all already loaded in the editor: 1 inspect the distribution, 2 calculate the estimate, 3 bootstrap, 4 build the confidence interval, 5 visualise the uncertainty, 6 write the interpretation.

Lab complete when the 95% interval has printed and you have written what it promises about the procedure across repeated samples — and what it does not promise about this one sample.

DV09 · Uncertainty interval

The editor already holds the response_times data frame — 96 practice questions, with the simulated response times in milliseconds stored in its response_ms column. Press Run and work top to bottom; the 96 values are response_times$response_ms. The values are drawn with rlnorm(), so describe them as simulated rather than collected.

Browser R note: this inline lab runs base R plus ggplot2 (installed automatically on first run — one short wait, cached after that). The full tidyverse collection is a desktop tool for RStudio/Positron. Everything this module requires runs right here in the browser.

R Lab B — DV10: Monte Carlo convergence

Open R Lab B — DV10: Monte Carlo convergence · about 15 minutes

Your task Six steps, all already loaded in the editor: 1 predict, 2 simulate, 3 plot the cumulative estimate, 4 calculate the simulation SE, 5 increase N, 6 defend the convergence.

Lab complete when the convergence output has printed and you have written the interpretation note at the foot of the lab — three sentences: the estimate, the simulation error, and what more trials would and would not fix.

A recorded successful Lab B run with output earns 40 practice points under the existing policy. Interpretation and DV10 evidence requirements remain separate.

DV10 · Monte Carlo convergence

The editor already holds the die-roll simulation: rolls, 2,000 trials under set.seed(44), the pooled estimate, and chunks — ten independent 200-roll estimates that show how far a short run can sit from the long-run value. Press Run and work top to bottom.

Browser R note: this inline lab runs base R plus ggplot2 (installed automatically on first run — one short wait, cached after that). The full tidyverse collection is a desktop tool for RStudio/Positron. Everything this module requires runs right here in the browser.

Canvas Project 3 evidence: DV09 and DV10

Open the DV09 and DV10 Project 3 evidence checklists

Complete the Project 3 evidence items below. For Canvas, upload at least one artifact from your work or a screenshot of your grade. The full ZIP is optional. A download is not a Canvas submission.

Your task Open each card, read the worked exemplar, then tick each check once your own work satisfies it. Add an optional note or link if helpful.

Project 3 ready when both required evidence cards are complete, or your previously earned full Module 5 credit preserves packet readiness. An earned grade never fills in missing checklist items, and an incomplete checklist never revokes that earned credit. Project_3_Canvas_Packet.zip is an optional way to bundle your work; an artifact or a screenshot of your grade also satisfies the upload instruction.

Submission decision check

Open the submission decision check

The uncertainty and simulation decision check needs JavaScript.

Shared Module 5 AI-feedback workspace

Open the shared Module 5 AI-feedback workspace

Where do I get my data?

For an interval or a bootstrap (DV09): R Lab A already holds the response_times data frame in its editor — 96 practice questions, with the simulated response times in milliseconds stored in its response_ms column. Open it and press Run; the 96 values you need are response_times$response_ms. The values are drawn with rlnorm(), so describe them as simulated rather than collected.

For Monte Carlo evidence (DV10): R Lab B already holds the die-roll simulation in its editor. Running it gives rolls, 2,000 trials under set.seed(44), the pooled estimate, and chunks, ten independent 200-roll estimates that show how far a short run can sit from the long-run value.

Rather not write R? The bootstrap lab reports a percentile interval from a fixed sample of 25 exam scores, and the Monte Carlo lab draws the convergence path.

Prefer your own dataset? You can use any built-in R dataset (run data() to list them) or add your own CSV under Your data in the R lab (files up to 500 MB stay in your browser).

Uncertainty and simulation capstone loading…

Adaptive reflection

What I know / where I go next

Mastery check loading…

Uncertainty & Simulation workspace record

Your saved uncertainty and simulation record needs JavaScript. You can also see it on your skills dashboard report card.

Project 3 connection Complete DV09 and DV10 here, then use the Project 3 page to verify readiness and choose your Canvas submission format: at least one artifact from your work or a screenshot of your grade; the full ZIP is optional. If your explanation depends on AI, include the prompt, output, verification method, and final correction.