Fairness can be defined in several mathematically precise ways. Choosing among them is a values decision, not just a technical one.
Common Definitions
- Demographic parity: each group receives positive outcomes at the same rate (for example, the same approval rate).
- Equal opportunity: among people who truly qualify, each group is equally likely to be approved — equal true positive rates.
- Equalised odds: equal true positive rates and equal false positive rates across groups.
- Predictive parity: when the model predicts positive, it's equally likely to be right for each group — equal precision.
- Calibration within groups: a predicted score of 0.7 means the same probability for every group.
They Conflict
When groups have different underlying rates of the outcome, it's generally impossible to satisfy all these definitions at once. Improving one can worsen another.
Choosing a Definition
Ask what harm matters most in context:
- Denying qualified people an opportunity → focus on equal opportunity.
- Wrongly flagging innocent people → focus on false positive rates.
- Scores used directly as probabilities → focus on calibration.
Document the choice and the reasoning, and involve affected stakeholders where possible.
Measuring in Practice
Compute metrics per group with confidence intervals, since small groups give noisy estimates. Libraries such as Fairlearn and AIF360 help.
Beyond Metrics
Fairness metrics can't capture everything: whether the task itself is appropriate, whether people can appeal, and whether the system changes behaviour in harmful ways.