TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1958 174
How MiniMax M2.1 made?

When they say that one model writes code better than another, they usually mean the SWE-Bench benchmark. The model gets a real bug from a real project on Github, which it has to read, find the error and fix it. This partially resembles a programmer's daily work.

🟢 How MiniMax-AI became a truly universal AI programmer?

SWE-Bench benchmark has its drawbacks. They found the answer and implemented it in their latest model M2.1.

1️⃣ LANGUAGE BARRIER

Problem: SWE-Bench only works with Python. In the real world, developers deal with Java, Go, TypeScript, Rust, C++ and a bunch of other languages.

⁠☞ Solution: SCALING THE ENVIRONMENT

Behind this vague term lies a huge system that operates with popular languages: JS, TS, Python, Java, Go, C++ and Rust.

For this, more than 100 thousand real tasks with a description of the problem, code and tests were collected from GitHub. This was not easy, as complex languages (Java or C++) require setup and each language has its own frameworks and dependency management systems.

To train the model on such a dataset, MiniMax built an infrastructure capable of running more than 5 thousand isolated execution environments in the shortest possible time - 10 seconds.

2️⃣ "BUG-FIX ONLY" TRAP
The Problem: Most benchmarks are is only about fixing bugs, while programmers also write new functions, refactor and optimize.

☞ The Solution: GOING BEYOND BUG FIXES:

MiniMax-M2.1 was also trained to generate tests, and it turned out that this is a critically important skill.

The previous version, M1, wrote too simple tests and often chose the wrong solutions. M2.1 excelled in this and equaled the results of the powerful competitor Claude Sonnet 4.5.

It also learned to optimize code performance - on SWE-Perf it showed an average increase in efficiency of 3.1%.

And finally, M2.1 was taught to do Code Review, for which an internal benchmark SWE-Review was created.

3️⃣ ENVIRONMENT DEPENDENCY
The Problem: A model's results strongly depend on the environment in which the model operates.

⁠☞ The Solution: GENERALIZATION ON OOD SCAFFOLDS.

The model should equally well follow long instructions and adapt to different ways of managing the context of the dialogue.

The team conducted tests in mini-swe-agent, Droid and Claude Code and if you look at the figures from their comparative table, you can see that the model has become much more flexible and versatile.

On the same SWE-Bench, when using Claude Code, MiniMax-M2.1 scored 74 points, which is higher than the model M2 with its 69.2 points, and almost on a par with Claude Sonnet 4.5 and DeepSeek V3.2.

On another test, OctoCodingBench, the gap is even greater: 26.1 for the new model against 13.3 for the old one.


Seems that concept of an "AI coder" is becoming more and more real. Success of MiniMax-M2.1 showed that it's no longer about writing individual lines of code, but about a comprehensive understanding of the entire development process.

#AI #ML #LLM #MiniMaх

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →