Exploratory Data Analysis
A First-Semester Reasoning-First Textbook + Practical Workbook
Authored by Karann • AURXON. Built for engineering students moving from zero Python knowledge to statistical intuition, hands-on Pandas mastery, defensible cleaning contracts, and production ML pipelines.
The 6-Step Analytical Reasoning Cycle
What does one single row represent? What is the column unit and time window? Predict what you expect to see before touching any code.
Worked Example: “The Mean Lies” Simulator
Suppose food delivery times (in minutes) are measured. A normal delivery takes 25–29 minutes, but one bike breakdown causes a 120-minute delivery. Inspect below how one extreme outlier pulls the Mean up by 15 minutes, while the Median tells the truth of the typical customer experience.
Data Cleaning Decision Matrix: MCAR vs. MAR vs. MNAR
Real-World Example: People with exceptionally high credit card debt or failing grades systematically refuse to answer debt/grade questions.
df['debt_missing'] = df['debt'].isna().astype(int). The missingness is itself the most valuable predictive signal!Textbook Units & Laboratory Implementations
Foundations of EDA, Measurement Scales & Environment Setup
EDA is not a collection of plots. It is a disciplined way of asking: What does this dataset actually say, what might be misleading us, and what should we do next?
| Scale | Definition | True Zero? | Permitted Operations | Valid Central Measure |
|---|---|---|---|---|
| Nominal | Categories without natural order (e.g., City, Gender) | No | Equality (==, !=) | Mode |
| Ordinal | Ordered ranks without equal intervals (e.g., 5-Star Rating) | No | Comparison (>, <) | Median, Percentiles |
| Interval | Equal distances, arbitrary zero (e.g., Temperature in Celsius) | No (0°C ≠ absence of heat) | Addition, Subtraction | Mean, Standard Deviation |
| Ratio | Meaningful zero and equal ratios (e.g., Revenue, Kelvin, Mass) | Yes | Multiplication, Division, Ratios | Geometric Mean, CV |
import pandas as pd
import numpy as np
def aurxon_first_pass_profile(df: pd.DataFrame) -> pd.DataFrame:
"""Profiles structure, missingness, cardinality, and dtypes."""
return pd.DataFrame({
"dtype": df.dtypes,
"null_count": df.isna().sum(),
"null_pct": (df.isna().mean() * 100).round(2),
"unique_vals": df.nunique(),
"sample_val": df.iloc[0] if len(df) > 0 else np.nan
})