a very interesting MoE model with 196 billion total and 11 active parameters.
Authors claim an insane speed of up to 300 tokens per second, and on tasks with code, it supposedly accelerates to 350. For a model of this level, this is very impressive.
🟢 What's going on inside?
Instead of the standard attention mechanism, they used a hybrid scheme: one layer of full attention on 3 sliding window layers, which allowed them to cram a context of 256 thousand tokens into the model without clogging up the memory to the point of failure.
In training, they used the MIS-PO algorithm, which helped solve the problem of losing the thread in long CoTs and simply cuts off options that deviate too much from logic.
The model, as is now fashionable, was tailored for autonomous agents. It can use ten tools at the same time. In Deep Research mode, the model itself googles, plans stages, and writes reports up to 10 thousand words long.
If you need to run a heavy code repository through the model, it handles it without the usual slowdowns that occur when working with voluminous texts.
➜ BENCHMARKS
Step 3.5 Flash scored 97.3 on the AIME 2025 test (and this is bare risoning, without third-party calculators). If you give it access to Python, the result soars to 99.8.
On code benchmarks, the numbers also look impressive: on SWE-bench, it gives 74.4%, and on Terminal-Bench 2.0 - 51.0%.
Of course, in terms of knowledge density, Step 3.5 Flash still lags behind Gemini 3.0 Pro, but the fact that it is available for local use and tests via API, is pleasing.
Article, Model, Demo, Discord Community, GitHub • #AI #ML #LLM #StepFunAI
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
