New Anthropic’s research: Can Claude autonomously align other AIs?
Researchers gave Claude 48 hours and 1 GPU to improve the alignment of small models.
It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well.
Could a model one day align its stronger successors?
As a first test, team had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model.
It reached safety scores approaching those of production Opus 4.8, which went through full alignment training.
Claude can reliably fix measurable misalignment. But subtle or rare failures may have no benchmark at all so everything hinges on measuring the right things.
Post #4442
693