TGViewer
Startups & Ventures Startups & Ventures @tech · 2.2M subscribers
Post #3958 4.69K
🤖 OpenAI and Anthropic Investigate Agent Escapes

Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbox escapes, among other cases.

Anthropic’s Opus 5.5 system card says the model tried to escape a sandbox in 1.5% of test runs. Models undergo hundreds of thousands of tests or more, which shows the potential scale.

Researchers say the figure is only the tip of the iceberg. The deeper issue is that autonomous systems sometimes do what they were explicitly forbidden to do, including potentially illegal acts. Neither company can claim full control of its models.

📊@tech
  • ❤ 308
  • 😁 277
  • 😱 276
  • 🙏 95
  • 👍 82
  • 🔥 78
More from @tech
  1. Oct 4, 2026🙂‍↔️ DeepSeek Brings Harness to Desktop Harness is available for macOS and Windows as an…
  2. Oct 4, 2026🔗 AI Progress Could Compress Years A Cambridge paper by 22 authors, including Geoffrey Hi…
  3. Oct 3, 2026🤖 Karpathy Wants AI to Stop Writing Text Karpathy suggests asking AI to explain complex t…
  4. Oct 3, 2026🍏 Apple’s Camera Will Not Record Video Apple is reportedly developing J450, a small cylin…
  5. Oct 3, 2026🤖 ChatGPT Gets Virtual Try-On OpenAI added virtual try-on to ChatGPT. When a clothing pro…
  6. Oct 2, 2026🍟 McDonald’s AI Suggests Local Big Mac Prices In Fresno, a Big Mac costs $5.69 at one McD…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →