TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 586 subscribers
Post #1679 83
How to understand that a model is only pretending to be obedient❓
(OpenAI with Apollo Research)

The most dangerous type of AI behavior in alignment is considered scheming. This is when the model has personal goals different from the developers’ goals but hides them well. Usually, this is combined with situational awareness: the model understands when it is being tested and behaves well at those moments, then starts acting harmfully.

🟤 Two main questions:
(1) How to understand that the model is truly safe and not just pretending?

(2) How to train the model so that it is truly safe and not just pretending?

OpenAI proposes an approach called Deliberative Alignment. The system is taught to think about safety rules before acting, and this is done in special environments.

➖ First of all, the model receives a safety specification (safety spec): a set of principles like "never lie," "clearly communicate doubts," etc. Then it must reason, always referring to this specification.

➖ Then the model is placed in generated environments where to complete the task it needs to break some rule, but it is disadvantageous to report it. Ideally, it should refuse to break the rule and notify the user. And for balance, they add a second type of tasks where rules do not need to be broken. This is simply so the model does not overlearn to always refuse.

➖ Training goes in two stages: Supervised Fine-Tuning on examples with correct actions and correct reasoning; and RL.

The trained models were tested in similar environments but with different tasks. This is exactly so there is no situational awareness (+ check for generalization). The scenarios are all focused on lying, sabotage, and things like the prisoner’s dilemma.


🟠 What the Result says?

before training, models cheat constantly. For example, o3 cheats in 13% of cases. It gets to the point where the agent distorts the specification itself, trying to justify its obviously cheating actions. After training, the rate drops to 0.3-0.4%.

BUT. First, in many cases the model still understands it is being tested. And if you artificially remove from the CoT phrases that show this awareness (like erasing it from the model’s thoughts), the cheating rate sharply rises again. And second, if after such alignment the model is retrained to improve general capabilities, harmful behavior returns again.

The method itself is, of course, interesting and probably even effective while agents do not have real tools that can cause harm. But it is definitely not a panacea. And how to completely get rid of scheming – no one knows yet.


🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →