Google launched Gemini 3.8 Flash for coding, agents, and multi-step reasoning, plus 3.8 Flash Cyber for vulnerability discovery and patching. Flash keeps 3.7 pricing, while Cyber access is limited to trusted defense users.
OpenAI says Astra has reached its Critical cyber threshold, showing top exploit and zero-day capability on hardened systems. Release is planned soon, with strongest access limited to alpha testers and Daybreak Blue for defense use.
๐จ AI News | TestingCatalogOPENAI ๐ฅ: GPT-6-Astra model slug has been spotted on the APIs. If we will actually get it tomorrow, it would be a huge week. Routing first ๐
Astralogy ๐ฎ
OpenAI employees are teasing an upcoming release of the Astra model.
OPENAI ๐ฅ: > Official โPath to Astraโ post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is โcoming soonโ โ advanced cyber tools stay limited to testers / Daybreak Blue at first. > M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.
GOOGLE ๐ฅ: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.
ANTHROPIC ๐ฅ: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads โ about 25% cheaper typically, up to 45% on heavy agent runs.
META ๐ฅ: > Muse Voice Transcribe is live โ MSLโs first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.
XAI ๐ฅ: > Elon: โGrok 4.7 comes out in 10 days.โ Thatโs ~Sept 12. Reply to Tobi on Grok 4.6.
ALIBABA ๐ฅ: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 โ 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.
WORLD LABS ๐ฅ: > Fei-Fei Liโs lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.
* Too much is happening, and I also have some scoops planned for today. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.
OPENAI ๐ฅ: Astra will be "available soon," but its cybersecurity capabilities will be limited.
> Astra scored 100% on ExploitBench. > OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently. > Astra achieved "much higher arbitrary code-execution rates than GPTโ5.6 Sol". > During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.