Skip to content

The AI Alignment Problem

Why making advanced AI systems reliably pursue intended goals is difficult, and why it grows more important with capability.

Editorial team 1 min read

Alignment means ensuring AI systems act according to the intentions and values of the people they serve, and of society more broadly.

Why It's Hard

  • Specifying goals: it's hard to fully describe what we want. Systems optimising imperfect objectives can find unintended shortcuts, known as specification gaming or reward hacking.
  • Learning the wrong lesson: training may produce behaviour that looks right in training but generalises differently in new situations.
  • Opacity: we can't yet fully inspect what a model has learned or why it acts as it does.
  • Evaluation limits: as systems become more capable, it gets harder for people to check their work.

Why It Grows in Importance

With narrow, low-stakes systems, misalignment causes contained mistakes. With general systems acting autonomously in high-stakes domains, consequences could be much larger.

Current Approaches

  • Training from human feedback and written principles.
  • Red-teaming and evaluations for dangerous behaviour.
  • Interpretability research to understand model internals.
  • Scalable oversight: using AI to help humans supervise AI.
  • Monitoring and control measures in deployment.

An Open Problem

Researchers broadly agree alignment isn't solved for highly capable systems; they disagree about how hard it is.

More in AGI

All AGI guides →
AGI Guide · 1 min

How Would We Know AGI Has Arrived?

Tests, benchmarks and practical measures proposed for recognising general intelligence in machines, and their limits.

AGI 1 min read 17 Jul 2025

AGI Guide · 1 min

A Brief History of the Quest for AGI

From the 1956 Dartmouth workshop through AI winters to large language models: how ambitions for general AI evolved.

AGI 1 min read 16 Jul 2025