Alignment means ensuring AI systems act according to the intentions and values of the people they serve, and of society more broadly.
Why It's Hard
- Specifying goals: it's hard to fully describe what we want. Systems optimising imperfect objectives can find unintended shortcuts, known as specification gaming or reward hacking.
- Learning the wrong lesson: training may produce behaviour that looks right in training but generalises differently in new situations.
- Opacity: we can't yet fully inspect what a model has learned or why it acts as it does.
- Evaluation limits: as systems become more capable, it gets harder for people to check their work.
Why It Grows in Importance
With narrow, low-stakes systems, misalignment causes contained mistakes. With general systems acting autonomously in high-stakes domains, consequences could be much larger.
Current Approaches
- Training from human feedback and written principles.
- Red-teaming and evaluations for dangerous behaviour.
- Interpretability research to understand model internals.
- Scalable oversight: using AI to help humans supervise AI.
- Monitoring and control measures in deployment.
An Open Problem
Researchers broadly agree alignment isn't solved for highly capable systems; they disagree about how hard it is.