Missing Data Mechanisms (MCAR, MAR, MNAR) & Cleaning Contracts
Understand the statistical taxonomy of missing data, establish defensible cleaning contracts, implement group-wise and KNN imputation, and prevent data leakage during preprocessing.
Learning Objectives
- •Categorize missing values into MCAR, MAR, and MNAR using statistical tests and domain intuition.
- •Construct missingness indicator features (df['col_isna'] = df['col'].isna().astype(int)).
- •Choose appropriate imputation methods (group-wise median, KNN, IterativeImputer).
- •Architect leakage-free preprocessing pipelines using scikit-learn Pipeline and ColumnTransformer.
Essential Prerequisites
- •Descriptive statistics and Pandas DataFrame filtering
The Core Mental Model
Why This Exists
Filling missing values blindly with the mean can severely distort variance, destroy correlations, and lead to catastrophic model failure in production. In credit scoring, MNAR missingness is itself the strongest predictive signal.
Beginner Foundation
Before calling df.fillna(), first calculate the percentage of missing values per column. If a column is missing 80% of data, imputation creates artificial noise. If it is missing 2%, simple median imputation or row deletion may be safe.
Micro Concepts Decomposition
The Three Missingness Mechanisms: MCAR, MAR, MNAR
MCAR (Missing Completely At Random): missingness is pure noise independent of all variables. MAR (Missing At Random): missingness depends on observed features (e.g., younger users report income less often). MNAR (Missing Not At Random): missingness depends on the unobserved value itself (e.g., high debt individuals refuse to state debt).
Cleaning as a Decision System
Cleaning is not making the dataframe pretty. It is deciding which representation is faithful to the analytical question. Every step must document: what problem does it solve, what information does it destroy, and how will it be verified.
Deduplication & Business Keys
There are three types of duplicates: exact byte-level row duplicates, repeated business events, and inconsistent entity spellings. Always deduplicate on explicit business keys (e.g. [customer_id, invoice_id]) rather than blindly dropping all matches.
Data Leakage Prevention in Imputation
Never fit an imputer or scaler on the complete dataset before train/test splitting. Doing so leaks information from the evaluation set into the training set. Always fit statistics on training data and transform the test data.
Hardware State Machine Architecture
Interactive Simulator
"The Mean Lies" & Tukey's IQR Outlier Laboratory
The Mean is sensitive to extreme values because it minimizes squared errors ($\sum (x_i - \mu)^2$). A single massive outlier drags the mean significantly toward itself. In contrast, the Median and IQR are rank-based order statistics that remain completely unaffected by the extreme magnitude of outliers.
End-to-End Execution Trace
Step-by-Step Code Execution (PYTHON)
Sandbox Terminal Ready
Click Run Code or press Ctrl+Enter to compile and execute.
Where Students Lose Marks
Active Assessment Quiz
Missing Data Mechanisms (MCAR, MAR, MNAR) & Cleaning Contracts — Practice Questions
If high-income individuals systematically refuse to disclose their salary on a loan application, what missing data mechanism is this?