A prediction you can't explain is hard to trust, debug or defend. Interpretability tools help you understand what a model has learned.
Why It Matters
- Debugging: spotting leakage or nonsensical patterns.
- Trust: helping users decide when to rely on predictions.
- Compliance: some decisions, such as credit, require reasons.
- Fairness: checking whether sensitive attributes or proxies drive outcomes.
Inherently Interpretable Models
Linear and logistic regression, shallow decision trees and rule lists can be read directly. When the accuracy cost is small, prefer them.
Global Explanations
- Permutation importance: shuffle one feature at a time on held-out data and measure how much performance drops.
- Partial dependence plots: show how predictions change, on average, as one feature varies.
Local Explanations
- SHAP values attribute a single prediction to each feature, based on cooperative game theory. They add up to the difference between the prediction and the average prediction.
- LIME fits a simple model around one prediction to approximate it locally.
Cautions
- Explanations describe the model, not necessarily cause and effect in the real world.
- Correlated features can split or swap importance in misleading ways.
- Different methods can disagree; use more than one.
For Deep Learning
Techniques such as saliency maps and attention visualisation exist, but interpreting large neural networks remains an open research area.