Module 15 · Project 7 · Due the date on Canvas · Final Project Synthesis & Evidence Workshop · adaptive competency DV23
Concept review — Time series, hypothesis testing and description
Mission progress0%
Begin by selecting your strongest evidence.
Time Series
- Trend
- The long-term upward or downward movement in a time series, separate from short-term fluctuations. Detecting a real trend requires enough data and awareness of seasonality — a few data points going up is not a trend.
- Seasonality
- Repeating cycles at a fixed period (daily, weekly, yearly). Seasonal patterns can mask or mimic trends; you must account for both when modeling time series, as you did in Modules 6 and 11.
- Moving Average
- A trailing smoother that averages the last \(k\) observations: \( m_t=k^{-1}\sum_{i=0}^{k-1}x_{t-i} \). It reduces noise but lags behind true changes and can flatten peaks. A moving average is a descriptive smoother; by itself it is not a fully specified forecasting model.
- Autocorrelation
- Correlation between a series and a lagged copy of itself. Statistically detectable autocorrelation is evidence against independence. Positive autocorrelation often makes independence-based intervals too narrow, while other dependence structures can affect uncertainty differently. Testing for autocorrelation is essential before applying inference to time-series data.
Hypothesis Testing
- Null and Alternative Hypotheses
- The null \(H_0\) represents the reference claim; the alternative \(H_1\) specifies the direction(s) or values against which evidence is evaluated. You never “prove” the alternative — you reject or fail to reject the null at a chosen significance level.
- P-Value
- The probability, assuming the null hypothesis and test model, of a test statistic at least as extreme as the observed statistic in the direction(s) specified by the alternative. A small p-value means the data would be surprising under the null. It is not the probability that the null is true, and it says nothing about effect size.
- Type I and Type II Errors
- A Type I error rejects a true null (false positive). A Type II error fails to reject a false null (false negative). Lowering one raises the other for a fixed sample size. You control the Type I rate through the significance level α.
- Statistical Power
- For a specified alternative, power is the probability of rejecting \(H_0\), namely \(1-\beta\). Power increases with sample size and effect size. A non-significant result from an underpowered test is not evidence for the null — it may simply mean you did not collect enough data.
- Effect Size
- Quantifies the magnitude of a difference or relationship, independent of sample size. A statistically significant result with a tiny effect size may be practically meaningless. You encountered effect sizes in Module 3 when defending group comparisons.
Descriptive Statistics and Visualization
- Median and IQR
- The median is the 50th percentile — resistant to outliers. The interquartile range (IQR) is the spread of the middle 50%. Together they summarize center and spread for skewed distributions where the mean and standard deviation would be misleading.
- Five-Number Summary
- Minimum, first quartile, median, third quartile, maximum. It gives a quick picture of distribution shape and is the basis for the boxplot. You used it in Module 2 to compare groups without assuming normality.
- Histogram
- Divides data into bins and counts observations per bin. Bin width changes the story — too wide hides detail, too narrow adds noise. Histograms show shape (symmetric, skewed, bimodal) but are sensitive to binning choices.
- Kernel Density Estimate
- A smooth alternative to the histogram that estimates the PDF by summing small bumps (kernels) centered at each data point. The bandwidth controls smoothness, just as bin width controls a histogram. You used KDEs in Modules 3 and 5.
- Boxplot and Violin Plot
- A Tukey boxplot shows the quartiles and median; its whiskers extend to the most extreme observations within the \( 1.5\,\mathrm{IQR} \) fences, with farther observations plotted separately. A violin plot shows a mirrored density estimate and may overlay summary statistics, depending on its design. Both compare distributions across groups without assuming a specific model.