IDRASAcademic OS
Unit 3: Data Cleaning Contracts & Missingness Mechanics 35 mins study timeINTERMEDIATE

Missing Data Mechanisms (MCAR, MAR, MNAR) & Cleaning Contracts

Understand the statistical taxonomy of missing data, establish defensible cleaning contracts, implement group-wise and KNN imputation, and prevent data leakage during preprocessing.

Verified: Faculty Peer Review Board

Learning Objectives

  • •Categorize missing values into MCAR, MAR, and MNAR using statistical tests and domain intuition.
  • •Construct missingness indicator features (df['col_isna'] = df['col'].isna().astype(int)).
  • •Choose appropriate imputation methods (group-wise median, KNN, IterativeImputer).
  • •Architect leakage-free preprocessing pipelines using scikit-learn Pipeline and ColumnTransformer.

Essential Prerequisites

  • •Descriptive statistics and Pandas DataFrame filtering
Layer 1: Intuition & Why It Matters

The Core Mental Model

“If exam papers blow out of a professor's car window, the missing grades are MCAR (random accident). If failing students skip the final exam, the missing grades are MNAR (missing because of the low grade). Treating them the same way by filling with class average would give zeros a free passing score!”

Why This Exists

Filling missing values blindly with the mean can severely distort variance, destroy correlations, and lead to catastrophic model failure in production. In credit scoring, MNAR missingness is itself the strongest predictive signal.

Beginner Foundation

Before calling df.fillna(), first calculate the percentage of missing values per column. If a column is missing 80% of data, imputation creates artificial noise. If it is missing 2%, simple median imputation or row deletion may be safe.

Micro Concepts Decomposition

MICRO CONCEPT 1Canonical Object

The Three Missingness Mechanisms: MCAR, MAR, MNAR

MCAR (Missing Completely At Random): missingness is pure noise independent of all variables. MAR (Missing At Random): missingness depends on observed features (e.g., younger users report income less often). MNAR (Missing Not At Random): missingness depends on the unobserved value itself (e.g., high debt individuals refuse to state debt).

Key Takeaway: MCAR: purely random; MAR: explained by other observed columns; MNAR: missingness caused by the missing value itself.
MICRO CONCEPT 2Canonical Object

Cleaning as a Decision System

Cleaning is not making the dataframe pretty. It is deciding which representation is faithful to the analytical question. Every step must document: what problem does it solve, what information does it destroy, and how will it be verified.

Key Takeaway: Cleaning destroys information. Every drop, clip, or impute must be verified with a before-and-after audit.
MICRO CONCEPT 3Canonical Object

Deduplication & Business Keys

There are three types of duplicates: exact byte-level row duplicates, repeated business events, and inconsistent entity spellings. Always deduplicate on explicit business keys (e.g. [customer_id, invoice_id]) rather than blindly dropping all matches.

Key Takeaway: Use df.duplicated(subset=['key_cols'], keep='first') with explicit business keys.
MICRO CONCEPT 4Canonical Object

Data Leakage Prevention in Imputation

Never fit an imputer or scaler on the complete dataset before train/test splitting. Doing so leaks information from the evaluation set into the training set. Always fit statistics on training data and transform the test data.

Key Takeaway: Fit imputers and scalers ONLY on training data; transform validation/test sets without re-fitting.
Layer 3 & 4: Formal Specification & Mechanism

Hardware State Machine Architecture

Rubin's Missing Data Theory proves that under MCAR, listwise deletion yields unbiased parameter estimates (though reduced sample power). Under MAR, maximum likelihood estimation and multiple imputation (MICE) yield valid inference. Under MNAR, joint selection modeling or pattern-mixture models are mandatory.
Aurxon Cleaning Protocol: 1. Audit missingness % by column and row. 2. Create binary missingness flags for MAR/MNAR columns. 3. Verify duplicate rows on domain business keys. 4. Standardize text cases and strip whitespace. 5. Split into train/validation sets before fitting any imputer. 6. Fit imputer on train, transform train and test. 7. Check variance shrinkage post-imputation.
Layer 7: Interactive Laboratory

Interactive Simulator

EDA • LABMissing Data Mechanisms & Imputation Strategy Lab
Launch Fullscreen Lab
EDA • DESCRIPTIVE STATISTICSDistribution & Outlier Sensitivity

"The Mean Lies" & Tukey's IQR Outlier Laboratory

Outliers Found:1 points
IQR Fence Multiplier:1.5 × IQR
Mean (μ)
36.67
Distorted by extreme values
Median (Q2)
28.00
Robust to outliers!
IQR (Q3 - Q1)
4.00
Middle 50% spread
Tukey Fences
[20.0, 36.0]
Values outside are outliers
1D Dot Plot with Fences:
Crucial EDA Principle:

The Mean is sensitive to extreme values because it minimizes squared errors ($\sum (x_i - \mu)^2$). A single massive outlier drags the mean significantly toward itself. In contrast, the Median and IQR are rank-based order statistics that remain completely unaffected by the extreme magnitude of outliers.

Layer 5: Step-by-Step Worked Numerical Example

End-to-End Execution Trace

MCAR vs MNAR Investigation: Data: customer_id employment_status income 0 C01 Employed 75000 1 C02 Employed 82000 2 C03 Unemployed NaN 3 C04 Unemployed NaN 4 C05 Employed 68000 Notice: Income is missing 100% for Unemployed customers! This is MAR (Missing At Random conditioned on Employment Status). Imputing overall median ($75,000) would assign unemployed customers high income!
Layer 6: Active Runtime CodeLab

Step-by-Step Code Execution (PYTHON)

Font
main.pyGlacier Light
Ln 1 • Python 3.12
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
786 chars • 26 lines • Ln 1UTF-8 • 4 Spaces
Interactive Terminal Shell

Sandbox Terminal Ready

Click Run Code or press Ctrl+Enter to compile and execute.

⚡ AURXON Bitstream Runtime v4.8IDRAS Academic Virtual Node
Common Student Pitfalls & Mistakes

Where Students Lose Marks

❌ Mistake: Calling df.fillna(df.mean()) directly without checking if the data is heavily skewed or contains MNAR mechanism.
✓ Correct Understanding: Report missingness pattern first, consider creating an indicator flag, and prefer median or model-based group imputation.
Layer 8: Practice & Knowledge Verification

Active Assessment Quiz

Interactive Assessment EngineQuestion 1 of 1

Missing Data Mechanisms (MCAR, MAR, MNAR) & Cleaning Contracts — Practice Questions

INTERMEDIATE LevelScore: 0/0

If high-income individuals systematically refuse to disclose their salary on a loan application, what missing data mechanism is this?

Academic Evaluation Preparation

Viva Examination & University Scoring Strategy

Standard Viva Examination Questions

Q1: What is data leakage during data cleaning?
Answer: Data leakage occurs when information from outside the training dataset (such as test set mean or scale) is used to create or impute features, causing overoptimistic validation metrics that fail in production.

How to Write High-Scoring University Exam Answers

Classify missing data mechanisms into MCAR, MAR, and MNAR with practical examples. Detail why simple mean imputation distorts sample variance. Provide a Python code blueprint demonstrating group-wise imputation and missingness indicator generation.