Wait a second, does Mistral 3 Large use the DeepSeek V3 architecture, including MLA?
I just went through the configurations: the only difference I saw is that there are 2 times fewer experts in Mistral 3 Large, but each expert is 2 times larger.
Maybe it's easier to notice this when comparing the architectures side by side
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1911
382
