(OpenAI with Apollo Research)
The most dangerous type of AI behavior in alignment is considered scheming. This is when the model has personal goals different from the developers’ goals but hides them well. Usually, this is combined with situational awareness: the model understands when it is being tested and behaves well at those moments, then starts acting harmfully.
🟤 Two main questions:
(1) How to understand that the model is truly safe and not just pretending?
(2) How to train the model so that it is truly safe and not just pretending?
OpenAI proposes an approach called Deliberative Alignment. The system is taught to think about safety rules before acting, and this is done in special environments.
➖ First of all, the model receives a safety specification (safety spec): a set of principles like "never lie," "clearly communicate doubts," etc. Then it must reason, always referring to this specification.
➖ Then the model is placed in generated environments where to complete the task it needs to break some rule, but it is disadvantageous to report it. Ideally, it should refuse to break the rule and notify the user. And for balance, they add a second type of tasks where rules do not need to be broken. This is simply so the model does not overlearn to always refuse.
➖ Training goes in two stages: Supervised Fine-Tuning on examples with correct actions and correct reasoning; and RL.
The trained models were tested in similar environments but with different tasks. This is exactly so there is no situational awareness (+ check for generalization). The scenarios are all focused on lying, sabotage, and things like the prisoner’s dilemma.
🟠 What the Result says?
before training, models cheat constantly. For example, o3 cheats in 13% of cases. It gets to the point where the agent distorts the specification itself, trying to justify its obviously cheating actions. After training, the rate drops to 0.3-0.4%.
BUT. First, in many cases the model still understands it is being tested. And if you artificially remove from the CoT phrases that show this awareness (like erasing it from the model’s thoughts), the cheating rate sharply rises again. And second, if after such alignment the model is retrained to improve general capabilities, harmful behavior returns again.
The method itself is, of course, interesting and probably even effective while agents do not have real tools that can cause harm. But it is definitely not a panacea. And how to completely get rid of scheming – no one knows yet.
🤖 Data Science, ML & Big Data with @DataXplore