Empirical Distributions, Skewness & The Mean Lies
In-depth statistical exploration: comparing mean vs median under skewness, the 'Mean Lies' case study, calculating dispersion (variance, std dev, IQR), and identifying outliers via Tukey's 1.5xIQR rule.
Learning Objectives
- •Calculate and interpret mean, median, mode, and trimmed mean across skewed and symmetric data.
- •Demonstrate the 'Mean Lies' effect on skewed delivery time or salary distributions.
- •Diagnose right-skewness versus left-skewness from summary statistics and histograms.
- •Apply Tukey's 1.5 x IQR rule to detect and tag anomalies in tabular datasets.
Essential Prerequisites
- •Basic arithmetic and percentage/percentile concepts
The Core Mental Model
Why This Exists
Machine learning models (linear regression, neural networks) are heavily degraded by unhandled outliers and skewed inputs. Exploratory Data Analysis reveals these anomalies before training broken models.
Beginner Foundation
When exploring a new dataset, you never trust averages blindly. You inspect the distribution shape: is it a symmetrical bell curve, or does it have a long tail of extreme numbers pulling the average off course?
Micro Concepts Decomposition
Central Tendency & 'The Mean Lies'
The arithmetic mean minimizes squared error but is heavily distorted by extreme values. Median is the 50th percentile order-statistic. In right-skewed data: Mean > Median. In left-skewed: Mean < Median.
Skewness: Right (Positive) vs. Left (Negative)
Skewness is the third standardized statistical moment: Gamma_1 = E[((X - mu) / sigma)^3]. Positive skewness indicates a longer tail to the right; negative skewness indicates a longer tail to the left.
Dispersion: Variance, Standard Deviation & IQR
Standard deviation measures average root-squared spread from the mean. Interquartile Range (IQR = Q3 - Q1) measures the span of the middle 50% of sorted values and is robust against outliers.
Outlier Detection: Tukey's 1.5 x IQR Rule
John Tukey established the standard boxplot fence: any data point falling below (Q1 - 1.5 * IQR) or above (Q3 + 1.5 * IQR) is flagged as an empirical outlier. Outlier flags are not automatic deletion rules.
Hardware State Machine Architecture
Interactive Simulator
"The Mean Lies" & Tukey's IQR Outlier Laboratory
The Mean is sensitive to extreme values because it minimizes squared errors ($\sum (x_i - \mu)^2$). A single massive outlier drags the mean significantly toward itself. In contrast, the Median and IQR are rank-based order statistics that remain completely unaffected by the extreme magnitude of outliers.
End-to-End Execution Trace
Step-by-Step Code Execution (PYTHON)
Sandbox Terminal Ready
Click Run Code or press Ctrl+Enter to compile and execute.
Where Students Lose Marks
Active Assessment Quiz
Empirical Distributions, Skewness & The Mean Lies — Practice Questions
If Q1 = 30 and Q3 = 50 in an exam dataset, what is the upper boundary threshold above which a score is considered an outlier according to Tukey's 1.5xIQR rule?