Новости Линукс Linux
По всем вопросам @evgenycarter
Post #20294
77

GPT-6 Astra hits 53% on Rails benchmark while Gemini regresses
GPT-6 Astra climbed from 35% to 53% on the Agents on Rails coding benchmark when reasoning was pushed to maximum, but Gemini went backward despite higher cost and effort. The results show that maximum reasoning is a model-specific tradeoff—not a universal accuracy upgrade—and even rescued GPT-5.6 Luna from a scoreless previous run.
Source
👉@sysadminoff
https://4sysops.com/archives/gpt-6-astra-hits-53-on-rails-benchmark-while-gemini-regresses/
GPT-6 Astra climbed from 35% to 53% on the Agents on Rails coding benchmark when reasoning was pushed to maximum, but Gemini went backward despite higher cost and effort. The results show that maximum reasoning is a model-specific tradeoff—not a universal accuracy upgrade—and even rescued GPT-5.6 Luna from a scoreless previous run.
Source
👉@sysadminoff
https://4sysops.com/archives/gpt-6-astra-hits-53-on-rails-benchmark-while-gemini-regresses/








