AI safety research seeks to ensure increasingly capable systems remain beneficial and under meaningful human control.
Interpretability
Understanding what happens inside neural networks — which features and circuits produce behaviour. Progress could let researchers detect deception or dangerous knowledge directly.
Evaluations
Testing models for dangerous capabilities (cyber offence, weapons knowledge, autonomous replication) and concerning tendencies (deception, power-seeking, sabotage) before and after deployment.
Alignment Training
Improving how models learn values and follow intended goals — from human feedback, written principles and better reward design.
Scalable Oversight
Methods for supervising systems that may exceed human ability in some areas, such as AI-assisted evaluation and debate.
Control
Designing deployments so that even a misaligned model can't cause serious harm: monitoring, restricted permissions and the ability to shut down.
Robustness and Security
Resistance to jailbreaks, adversarial inputs and theft of model weights.
Getting Involved
Safety research spans machine learning, security, policy and social science. Many organisations publish research and run fellowships for newcomers.