Module 15 · Project 7 · Due the date on Canvas · Final Project Synthesis & Evidence Workshop · adaptive competency DV23
Concept review — Probability, distributions and sampling
Mission progress0%
Begin by selecting your strongest evidence.
Below is a compact reference for every major mathematical concept in the course. Use it to refresh your memory before defending your portfolio. Each entry explains what the concept is and when you would use it. The notation uses plain Unicode symbols so it copies cleanly into Canvas or any text editor.
Probability & Random Variables
- Random Variable
- A function that assigns a number to each outcome of a random process. Discrete random variables take countable values (like 0, 1, 2); continuous ones take any value in an interval. You used random variables throughout Modules 3–5 to model uncertainty in data.
- Sample Space
- The set of all possible outcomes of a random experiment. Every probability calculation starts here — if you cannot enumerate or describe the sample space, you cannot compute probabilities correctly.
- Probability Mass Function (PMF)
- For a discrete random variable, the PMF gives P(X = x) for each possible value x. It must be non-negative and sum to 1. You built interactive probability tables in Module 3 to explore how PMFs change as parameters shift.
- Cumulative Distribution Function (CDF)
- The CDF gives F(x) = P(X ≤ x). It accumulates probability up to a value. For discrete variables it is a step function; for continuous variables it is a smooth curve. You used CDFs in Modules 3–4 to move between individual outcomes and cumulative probabilities.
- Probability Density Function (PDF)
- For a continuous probability distribution, probabilities are represented by areas under the probability density function, rather than by the height of the density curve at a particular point. The total area under the PDF equals 1.
- Expected Value
- The long-run average of a random variable. For discrete: E[X] = Σ x · P(X=x). For continuous: E[X] = ∫ x · f(x) dx. It is the center of mass of the distribution — where the distribution balances.
- Variance and Standard Deviation
- Variance \( \operatorname{Var}(X) = E[(X - \mu)^2] \) measures spread around the mean. Standard deviation \( \sigma = \sqrt{\operatorname{Var}(X)} \) is in the same units as the data, making it more interpretable. Together with the mean, these two numbers summarize the center and spread of a distribution.
Discrete Distribution Models
- Bernoulli Distribution
- The simplest random variable: a single trial with two outcomes, success (1) with probability p and failure (0) with probability 1−p. Mean is p, variance is p(1−p). Every more complex binomial model builds on this.
- Binomial Distribution
- Counts the number of successes in n independent Bernoulli trials: P(X=k) = C(n,k) · pk · (1−p)n−k. Mean = np, variance = np(1−p). Used in Module 3 to model repeated yes/no experiments and in Module 9 as the likelihood in Bayesian updating.
- Geometric Distribution
- Counts the number of trials until the first success: P(X=k) = (1−p)k−1 · p. Mean = 1/p. Models waiting times — how many attempts before something happens. You encountered this in Module 3 alongside the binomial.
- Poisson Distribution
- Models the number of rare events in a fixed interval at rate λ: P(X=k) = (λk · e−λ) / k!. Mean and variance are both λ. Used for counts of infrequent events — arrivals, defects, clicks per minute.
Continuous Distributions
- Normal Distribution
- The bell curve N(μ, σ²), symmetric about its mean, fully described by μ and σ. Approximately 68% of data falls within one standard deviation, 95% within two. It is the backbone of the CLT and nearly all classical inference.
- Z-Score
- Standardizes a value: \( z=(x-\mu)/\sigma \). A z-score tells you how many standard deviations an observation is above or below the mean, and lets you compare values from different normal distributions on a common scale.
- Percentile
- The value below which a given percentage of the distribution falls. The 50th percentile is the median. Percentiles are read from the CDF and are essential for interpreting positions within a distribution, including standardized test scores.
- Exponential Distribution
- For a Poisson process with rate \( \lambda>0 \), models the waiting time between events with density \( f(x)=\lambda e^{-\lambda x} \) for \( x\ge 0 \). It is memoryless — the expected remaining wait does not depend on how long you have already waited. You encountered it in Module 5 as a contrast to the Poisson count model.
Sampling and Inference
- Central Limit Theorem (CLT)
- For independent and identically distributed observations with finite mean \( \mu \) and variance \( \sigma^2 \), the sample mean is approximately \( \bar X_n \sim N(\mu,\sigma^2/n) \) for large \(n\). This is what makes normal-based inference valid even when the data are not normal — provided the sample is large, independent, and not heavy-tailed. You built an interactive CLT simulation in Module 5.
- Sampling Distribution
- The distribution of a statistic (like the sample mean) over all possible samples of the same size. The CLT describes its shape. Understanding sampling distributions is the difference between knowing your estimate and knowing how much to trust it.
- Standard Error (SE)
- The standard deviation of a sampling distribution. For the sample mean of independent observations, \( \mathrm{SE}(\bar X)=\sigma/\sqrt{n} \). SE shrinks as sample size grows — larger samples mean more precise estimates.
- Confidence Interval
- For an unknown population SD, a classical mean interval uses the Student-t critical value: \( \bar{x} \pm t^*_{n-1}s/\sqrt{n} \). A z critical value applies when \( \sigma \) is known or as a justified large-sample approximation. A 95% confidence interval means that if you repeated the sampling many times, 95% of the intervals would contain the true parameter. It quantifies uncertainty; it does not give the probability that the parameter is inside this particular interval.
- Bootstrap
- A resampling method that draws samples with replacement from the observed data to approximate the sampling distribution. It works when no formula for the SE exists or when CLT assumptions are questionable. But it cannot fix a non-representative sample.
- Monte Carlo Simulation
- Estimates quantities by repeated random sampling. The estimate converges as the number of replications increases. You used Monte Carlo simulation in Module 5 to demonstrate convergence and to approximate probabilities that have no closed-form solution.