Mistral launches Large 4 in API preview, leading a test of finding and fixing security bugs at 81.7%
Artificial Analysis used 131 tasks from C/C++ projects such as FFmpeg and CPython. A pass meant showing a crash, supplying a patch that stopped it, and keeping existing tests passing.
The score can include fixing a different real bug from the one the task targeted.
Mistral plans to release the model for download by the end of October. That could let teams test this repair workflow on their own servers.
Post #824
79
