Module 15 · Project 7 · Due the date on Canvas · Final Project Synthesis & Evidence Workshop · adaptive competency DV23
Concept review — Regression, Bayesian updating and dimension reduction
Mission progress0%
Begin by selecting your strongest evidence.
Correlation and Regression
- Covariance
- Measures how two variables move together: \( \mathrm{Cov}(X,Y)=E[(X-\mu_X)(Y-\mu_Y)] \). Positive covariance means they tend to increase together; negative means one rises as the other falls. It is unstandardized, so its magnitude depends on units.
- Pearson Correlation Coefficient
- Population correlation is \( \rho=\mathrm{Cov}(X,Y)/(\sigma_X\sigma_Y) \); sample correlation is \( r=s_{XY}/(s_Xs_Y) \). Both are bounded to \([-1,1]\). Values near ±1 indicate strong linear association; near 0 indicates weak linear association. Crucially, correlation does not imply causation — a lesson reinforced throughout the course.
- Ordinary Least Squares (OLS) Regression
- Fits \( \hat y=\hat\beta_0+\hat\beta_1x \) by minimizing the sum of squared residuals. In simple regression, \( \hat\beta_1=r(s_y/s_x) \) and \( \hat\beta_0=\bar y-\hat\beta_1\bar x \). OLS is the workhorse of Module 4.
- Residual
- The difference between observed and fitted value: \( e_i=y_i-\hat y_i \). Residuals should look like random noise. Patterns in residuals — curvature, fans, outliers — signal that the model is missing structure.
- Leverage
- How far an observation’s x-value is from the mean of x. High leverage creates the potential to pull the regression line, but influence also depends on the observation’s residual; a single influential observation can reverse the slope. You learned to detect and defend against this in Module 4.
- Heteroscedasticity
- Non-constant variance of residuals across the range of x. It makes conventional homoscedastic standard errors unreliable—they may be too small or too large—so confidence intervals and p-values require a defensible variance model or robust standard errors. Detected visually from a residual-vs-fitted plot — the spread should be roughly even.
Bayesian Statistics
- Bayes’ Theorem
- \( P(\theta\mid\text{data})=P(\text{data}\mid\theta)P(\theta)/P(\text{data}) \). It updates a prior belief with observed evidence to produce a posterior. The central idea of Module 9 and the foundation of all Bayesian inference.
- Prior Distribution
- Your belief about a parameter before seeing data. It can be informative, weakly informative, or a stated reference prior; a flat prior is not universally “uninformative” because flatness depends on parameterization. Under regular conditions, as compatible data accumulate, the likelihood often dominates provided the prior gives support to the relevant parameter region.
- Likelihood
- The probability mass or density assigned to the observed data as a function of a specific parameter value: \( L(\theta;\text{data}) \propto P(\text{data}\mid\theta) \). It is not a probability distribution over θ — it does not integrate to 1 over the parameter. It measures how well each parameter value explains the data.
- Posterior Distribution
- The updated belief after combining prior and likelihood. It is the object you report, visualize, and draw conclusions from. The posterior mean is a common point estimate; credible intervals make posterior-probability statements and therefore have a different interpretation from frequentist confidence intervals.
- Conjugate Prior
- A prior chosen so that the posterior falls in the same distribution family. A Beta prior with a binomial likelihood gives \( \mathrm{Beta}(\alpha,\beta) \to \mathrm{Beta}(\alpha+k,\beta+n-k) \). Conjugacy makes Bayesian updating tractable without simulation.
- Posterior Predictive Check
- Simulates new datasets from the posterior distribution and compares them to the observed data. If the model is adequate, simulated data should resemble observed data. A mismatch means the model cannot generate the patterns you see — it should be revised, no matter how concentrated the posterior is.
Linear Algebra and Dimension Reduction
- Covariance Matrix
- A square matrix where entry \( (i,j) \) is the covariance between variables \(i\) and \(j\). Diagonal entries are variances. PCA operates on this matrix (or its correlation-based equivalent) to find directions of maximum spread.
- Eigenvalue and Eigenvector
- For a matrix \( A \), an eigenvector \( v \) satisfies \( Av=\lambda v \). For PCA specifically, when \( A \) is a covariance or correlation matrix, the eigenvalue \( \lambda \) is the variance captured along that principal direction. In PCA, eigenvectors of the covariance matrix are the principal components; eigenvalues rank them by importance.
- Principal Component Analysis (PCA)
- Projects high-dimensional data onto orthogonal axes that capture the most variance, in decreasing order. It reduces dimensionality while preserving as much information as possible. It is a linear technique — it cannot capture nonlinear structure.
- Scree Plot
- A bar or line chart showing the variance explained by each principal component. The “elbow” — where explained variance drops sharply — helps you decide how many components to keep. There is no single correct cutoff; the decision depends on interpretability and downstream use.
- Loadings
- The weights that connect original variables to a principal component. High loadings tell you which variables drive each component. A biplot overlays observations (as points) and variables (as arrows) in the component space so you can see both at once.
- Projection
- The shadow of a data point onto a lower-dimensional subspace. PCA projections are orthogonal — each component is uncorrelated with the others. The projected coordinates are called scores.