Descriptive statistics summarise a dataset with a few numbers. They're the first step in any analysis.
Measures of Centre
- Mean: the average. Sensitive to extreme values — one billionaire raises the average income of a room dramatically.
- Median: the middle value when sorted. Robust to outliers; better for skewed data such as incomes and house prices.
- Mode: the most common value. Useful for categories.
Measures of Spread
- Range: maximum minus minimum. Heavily affected by outliers.
- Interquartile range (IQR): the spread of the middle 50% of values.
- Standard deviation: typical distance from the mean. Most meaningful for roughly symmetric data.
Shape
- Skewness: whether the distribution has a long tail to one side.
- Multiple peaks can indicate distinct groups mixed together.
Always look at a histogram as well as the numbers.
Percentiles
The 90th percentile is the value below which 90% of observations fall. Percentiles describe distributions well and are standard for metrics such as response times.
Summarising Categories
Counts and proportions, shown with bar charts.
Common Pitfalls
- Reporting a mean for heavily skewed data.
- Ignoring missing values when calculating statistics.
- Comparing groups of very different sizes without noting it.
df["income"].describe(percentiles=[0.1, 0.5, 0.9])