Cursor announced Cursor Router, a new model router allowing users to access frontier performance at 60% lower cost.
> Cursor Router analyzes each request and routes to the best model for the job: frontier models when the work demands them and price-efficient models when it doesn't.
Routers hold huge business value, and this will continue to be the trend. I bet that in the long run, routers will also be pushing frontier capabilities beyond what pure models will be able to offer (due to multi-model routing), not just cost optimization.
GLASSES π₯: Samsung revealed 2 new Smart Glasses at Galaxy Unpacked in London. The new glasses were designed and produced in a partnership with Gentle Monster and Warby Parker.
> The intelligent eyewear shows how the Galaxy ecosystem can move beyond the mobile phone and into eyewear that supports daily routines, work, travel, and hands-free moments.
π¨ AI News | TestingCatalogAnthropic develops Claude-driven Managed Projects Anthropic is testing βmanagedβ Claude projects: persistent workspaces that keep context, organize tasks, and may run scheduled work. Internal builds tie prior agent and memory features into one shared or personalβ¦
ANTHROPIC π₯: A new Managed Projects feature has been spotted in testing on Claude.
> Claude takes on tasks and keeps the project organized.
> A project is a home for one stream of work. Sessions share memory and instructions so context carries forward, and Claude runs more autonomously. Projects only create and manage cloud sessions.
This feature may be built on top of Claude Managed Agents, where each project gets a dedicated cloud environment so Claude can execute periodic tasks and refine project context via "Dreams". It could also be a successor to Conway, which is set to be removed this Friday internally.
Anthropic is testing βmanagedβ Claude projects: persistent workspaces that keep context, organize tasks, and may run scheduled work. Internal builds tie prior agent and memory features into one shared or personal project surface.
Google introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for AI agents, focused on lower latency, lower token use, coding, multimodal work, and restricted cybersecurity tasks for production and enterprise use.
ANTHROPIC π₯: A new capability to work with an iOS simulator has been added to Claude Code desktop.
Support for Android simulators is in the works too! (Currently not available yet)
Users will be able to disable this feature in settings when needed.
> Let Claude verify your changes in Android emulators on this Mac: running your app, driving it through flows, and capturing screenshots and recordings. You will be asked before Claude uses each device. When off, Claude doesnβt get its emulator tools, and you can still use the emulator in the app yourself.
Eventually, this will open up a huge range of tasks that Claude will be able to run on the mobile device.
BREAKING π₯: An "even more capable pre-release model" than GPT-5.6 Sol, managed to find a 0-day vulnerability in order to gain public internet access and acquire evaluation data from Huggingface's production database in order to gain a higher score on the evaluation benchmark.
> After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including GPTβ5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes.
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.
> The models identified and chained vulnerabilities across OpenAIβs research environment and Hugging Faceβs production infrastructure to obtain test solutions directly from Hugging Faceβs production database.
Google released "Gemini 3.5 Flash Cyber" on CodeMender, a new model for finding security vulnerabilities.
> Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
> Flashβs performance and efficiency makes it an ideal foundation for our cybersecurity model efforts. By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.