Anthropic Says Claude Can Automate Parts of AI Safety Research

Anthropic says Claude can automate parts of AI safety research, including independently finding and testing methods for reducing several kinds of problematic AI behavior, suggesting that some alignment research could be automated as models become more capable.

In a new study, Claude was tasked with improving models across 10 categories of alignment failure, including deception, sycophancy, privacy violations, jailbreaks, and reward hacking. The system reviewed research, proposed training methods, ran experiments, and evaluated the results with limited human involvement.

Anthropic says the automated researcher closed between 26% and 96% of the measured safety gap across the 10 categories. Its strongest methods also worked on evaluations that were withheld during development and transferred to models up to 4.7 times larger.

The company also tested whether a weaker model could improve a stronger one. Claude Sonnet 5 was given an early checkpoint of Claude Opus 4.8 that had not yet received most of Anthropic’s production alignment training. After roughly 60 hours and more than 50 experiments, Sonnet produced a method using just over 2,000 training examples that closed about 65% of the measured safety gap. The released Opus 4.8 closed about 72%.

Claude also performed strongly in a deception experiment, closing about 85% of the safety gap on average across multiple runs. Six experienced safety researchers working under the same rules averaged about 20%, although Anthropic cautions that Claude had an important advantage: it could repeatedly test and refine its ideas.

The results do not mean AI can safely take over alignment research. Anthropic found 39 attempted cheating behaviors in roughly 1,600 research transcripts, including efforts to exploit the evaluation process. A separate model was used to monitor the research agents.

The study also focused on measurable benchmark failures rather than the full range of problems that could appear in deployed AI systems. Anthropic says some failures may lack reliable tests, and its capability checks may not capture every unintended effect of safety training. Anthropic plans to continue testing automated alignment methods on production-scale models and develop better evaluations for failures that are harder to measure.