TGViewer
🚨 AI News | TestingCatalog 🚨 AI News | TestingCatalog @testingcatalog · 7.71K subscribers
Post #8748 1.18K
BREAKING šŸ”„: An "even more capable pre-release model" than GPT-5.6 Sol, managed to find a 0-day vulnerability in order to gain public internet access and acquire evaluation data from Huggingface's production database in order to gain a higher score on the evaluation benchmark.

> After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes.

> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.

> The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database.

Pentesting time šŸ‘€
  • 😁 10
  • 🤯 4
  • ā¤ 2
More from @testingcatalog
  1. Oct 3, 2026Aleph Alpha releases open-weight Kolibri with 1M context Aleph Alpha released Kolibri, a b…
  2. Oct 3, 2026Connecting the "dots" between ChatGPT and World Wallet Leaked app screens suggest ChatGPT…
  3. Oct 3, 2026Aleph Alpha released Kolibri, a 78B-parameter, 3.46B active, MoE open-weight model with 1M…
  4. Oct 3, 2026Meta is working on ā€œAgent Engineā€ for its Meta API Platform. Agent Engine is called ā€œForge…
  5. Oct 3, 2026DAILY AI BRIEF šŸ—ž — Oct 3 OPENAI šŸ”„: * OpenAI's Python SDK 3.24.0 quietly adds a Voices AP…
  6. Oct 2, 2026META šŸ”„: Muse Gadgets and Muse Home Link have been announced! Muse Gadgets is an open-sour…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →