Mathematical Theory Review — The Complete Toolkit
Begin by selecting your strongest evidence.
This review summarizes every major mathematical construct you have encountered across Modules 1–14. You will need to identify, explain, and defend these concepts as part of your synthesis narrative. None of this is new — but you should be able to state what each one is, when it applies, and what it cannot do.
1. Probability Foundations and Discrete Distributions (Modules 3–5)
A random variable \(X\) maps outcomes from a sample space to real numbers. A probability mass function (PMF) \( f(x) = P(X = x) \) assigns probability to each discrete outcome and must satisfy two axioms: \( f(x) \geq 0 \) for every \( x \), and \( \sum_x f(x) = 1 \). The cumulative distribution function (CDF) is \( F(x) = P(X \leq x) \). The expected value is \( E[X] = \sum_x x \, f(x) \), and the variance is \( \mathrm{Var}(X) = E[X^2] - (E[X])^2 \), with standard deviation \( \sigma = \sqrt{\mathrm{Var}(X)} \).
You studied four discrete models. A Bernoulli trial has outcomes 0 and 1 with \( P(X=1) = p \), giving \( E[X] = p \) and \( \mathrm{Var}(X) = p(1-p) \). The binomial distribution counts successes in \( n \) independent trials: \( P(X=k) = \binom{n}{k} p^k (1-p)^{n-k} \), with \( E[X] = np \) and \( \mathrm{Var}(X) = np(1-p) \). The geometric distribution counts trials until the first success: \( P(X=k) = (1-p)^{k-1} p \), with \( E[X] = 1/p \). The Poisson distribution models rare events at rate \( \lambda \): \( P(X=k) = \frac{\lambda^k e^{-\lambda}}{k!} \), with \( E[X] = \mathrm{Var}(X) = \lambda \).
2. Continuous Distributions and Normal Models (Modules 4–5)
For a continuous probability distribution, probabilities are represented by areas under the probability density function, rather than by the height of the density curve at a particular point. For \( f(x) \), \( P(a \leq X \leq b) = \int_a^b f(x)\,dx \) and the total area equals 1. The normal distribution \( N(\mu, \sigma^2) \) has density \( f(x) = \frac{1}{\sigma\sqrt{2\pi}} \, e^{-(x-\mu)^2 / (2\sigma^2)} \). A z-score standardizes any normal variable: \( z = \frac{x - \mu}{\sigma} \), mapping it to the standard normal \( N(0, 1) \). Percentiles and tail probabilities are read from the standard normal CDF \( \Phi(z) \).
3. Sampling Distributions and the Central Limit Theorem (Module 5)
The Central Limit Theorem (CLT) states that for independent and identically distributed observations with finite mean \( \mu \) and variance \( \sigma^2 \), the sample mean converges in distribution:
\[ \sqrt{n}\,\frac{\bar{X}_n - \mu}{\sigma} \;\xrightarrow{d}\; N(0,\,1) \]
Equivalently, for large \( n \), \( \bar{X}_n \stackrel{\text{approx}}{\sim} N\!\left(\mu,\, \frac{\sigma^2}{n}\right) \). The standard error of the mean is \( \mathrm{SE} = \frac{\sigma}{\sqrt{n}} \). The CLT justifies normal-based inference for means even when the underlying distribution is non-normal, provided the sampling assumptions hold and the sample is large enough for the population tail behavior. It does not repair biased sampling, dependent observations, or heavy-tailed distributions with infinite variance.
When \( \sigma \) is unknown, a classical confidence interval for a normal-population mean is \( \bar{x} \pm t^*_{n-1} \frac{s}{\sqrt{n}} \). When \( \sigma \) is known, use \( \bar{x} \pm z^* \frac{\sigma}{\sqrt{n}} \); \( z^* \approx 1.96 \) for 95%. Large-sample procedures may use the normal approximation when its conditions are defensible. The bootstrap resamples the observed data with replacement to approximate the sampling distribution empirically — useful when the CLT is questionable or no closed-form SE exists. Monte Carlo simulation estimates quantities of interest by repeated random sampling and converges as the number of replications grows.
4. Correlation, Regression, and Model Claims (Module 4)
Covariance measures linear co-movement: \( \mathrm{Cov}(X, Y) = E[(X - \mu_X)(Y - \mu_Y)] \). The population Pearson correlation is \( \rho = \frac{\mathrm{Cov}(X, Y)}{\sigma_X \sigma_Y} \), bounded to \([-1, 1]\); its sample analogue is \( r=s_{XY}/(s_Xs_Y) \). Ordinary least squares (OLS) regression fits a line \( \hat{y} = \hat{\beta}_0 + \hat{\beta}_1 x \) by minimizing the residual sum of squares \( \sum (y_i - \hat{y}_i)^2 \). The slope is \( \hat{\beta}_1 = r \cdot \frac{s_y}{s_x} \). Residuals \( e_i = y_i - \hat{y}_i \) should show no pattern; systematic residual structure signals model misspecification. Heteroscedasticity (non-constant conditional variance) makes conventional homoscedastic standard errors unreliable; they may be too small or too large. Leverage measures how unusual an observation is in predictor space; high leverage creates the potential for influence, while actual influence also depends on the residual — a single high-leverage observation can reverse the slope. You learned to diagnose these issues visually and to defend when a regression claim is justified versus an overreach.
5. Bayesian Updating and Model Criticism (Module 9)
Bayes’ theorem updates a prior belief with observed evidence:
\[ P(\theta \mid \text{data}) = \frac{P(\text{data} \mid \theta) \, P(\theta)}{P(\text{data})} \]
The prior \( P(\theta) \) encodes belief before seeing data. The likelihood \( P(\text{data} \mid \theta) \) is the probability of the observed data given a parameter value. The posterior \( P(\theta \mid \text{data}) \) is the updated belief. With a conjugate prior, the posterior falls in the same family as the prior — for example, a Beta prior paired with a binomial likelihood yields a Beta posterior: \( \mathrm{Beta}(\alpha, \beta) \to \mathrm{Beta}(\alpha + k, \, \beta + n - k) \). Posterior predictive checks simulate new datasets from the posterior and compare them to observed data to assess model fit. A model that cannot generate data resembling what was actually observed should be revised, regardless of how well the posterior concentrates.
6. Linear Algebra and Principal Component Analysis (Modules 6, 10)
PCA reduces dimensionality by finding orthogonal directions of maximum variance. Starting from a centered data matrix—usually standardized as well when variables use incomparable units—PCA computes the eigenvalues and eigenvectors of the covariance matrix \( \mathbf{\Sigma} \). Each eigenvalue \( \lambda_i \) gives the variance explained by its eigenvector (principal component). The proportion of variance explained by component \( i \) is \( \lambda_i / \sum_j \lambda_j \). A scree plot displays these proportions in descending order to help decide how many components to retain. Loadings measure how much each original variable contributes to a component; a biplot overlays scores and loadings to visualize both observations and variables in the reduced space. Rotation can improve interpretability. Rotating a fixed retained subspace preserves the total variance represented by that subspace, although variance assigned to individual rotated axes changes; oblique rotation also permits correlated axes and requires more careful variance accounting.
7. Time-Series Structure (Modules 6, 11)
A time series decomposes into trend (long-term direction), seasonality (repeating cycles), and residual noise. A moving average smooths short-term fluctuation: \( \bar{x}_t = \frac{1}{k} \sum_{i=0}^{k-1} x_{t-i} \). Autocorrelation measures correlation between a series and a lagged version of itself; statistically detectable autocorrelation is evidence against independence. It can invalidate independence-based standard errors and make apparent trend evidence stronger than it really is unless the dependence is modeled.
8. Hypothesis Testing and Inference (Modules 3–5, 9)
A hypothesis test compares a null hypothesis \( H_0 \) to an alternative \( H_1 \). The p-value is the probability, assuming \( H_0 \) and the test model, of a test statistic at least as extreme as the observed one in the direction(s) specified by the alternative. A Type I error rejects a true null (false positive); a Type II error fails to reject a false null (false negative). Statistical power is \( 1 - P(\text{Type II}) \). An effect size quantifies magnitude independent of sample size. You learned that a small p-value alone does not imply a practically meaningful effect, and that failure to reject \( H_0 \) is not evidence for the null — it may simply reflect insufficient power.
How to use this review
As you select evidence for your synthesis, identify which of these mathematical constructs each artifact relies on. Your defense should be able to answer: What distribution or model underlies this analysis? What assumptions does it require? Were those assumptions checked? What would violate them? What does the math allow you to claim — and what does it stop you from claiming? A portfolio that can answer these questions for every artifact is defensible. One that cannot is not.