Exploratory data analysis (EDA) is where you learn what a dataset actually contains. Skipping it leads to wrong conclusions later.
1. Structure
- How many rows and columns?
- What does one row represent?
- Are column types as expected (numbers, dates, text)?
2. Missing and Invalid Values
- Missing values per column, including hidden markers such as
?or-999. - Values outside plausible ranges.
- Duplicated rows or duplicated IDs.
3. Distributions
- Summary statistics for numeric columns: mean, median, spread, minimum, maximum.
- Histograms to see shape, skew and outliers.
- Value counts for categories, including rare categories.
4. Relationships
- Correlations between numeric variables.
- Scatter plots and grouped summaries against the target.
- Cross-tabulations for categories.
5. Time
- Coverage: which dates are included, and are there gaps?
- Trends and seasonality.
- Changes in definitions or collection over time.
6. Segments
- Do patterns hold across regions, products or customer groups?
- Are some groups very small?
7. Sanity Checks
- Do totals match known figures from other sources?
- Do domain experts recognise the patterns?
Record What You Find
Keep notes on quirks, assumptions and cleaning decisions. They become documentation for everyone who uses the data after you.