I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."
If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?
SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).
Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
Very nice progress. Also I respect putting Kimi on those charts. Regardless of if they are beating the Pareto frontier (not now), model diversity is a good thing for humanity — I’m hopeful for the team to keep increasing their gains.
The main thing I always get away from the comparison tables of these "big" models, is how well Deepseek v4.1 Flash performs. While still being the cheapest model by a long shot.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
I tried it a bit and I like it! It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming. I think I will keep urimg this as my main assistant for few weeks.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Excited to try this. The low costs v. benchmarks alone here are worth a serious test. K3 has been my daily driver for a month or two now and it's dramatically reduced token spend (while not having much of a negative impact on productivity).
This was the era of the AI race I was waiting for.
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
Maybe I'm missing something. Doesn't seem super impressive to me. A proprietary model with performance comparable to GPT-6 Luna and Deepseek 4.1 Flash, but at a higher price than either. The main selling point is that it's made in Europe... not very compelling, globally. I suppose maybe there is some niche where European-hosted open-weight models aren't enough to satisfy some EU regulation, where only the use of European-trained models is in compliance, but as a non-European I have no idea what that niche would be.
Side note: Wish this thread was more focused on talking about the model instead of debating about China and America. Whatever happened to staying on-topic?
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
Still second most expensive open-weight model. I don't care about cybersecurity index.
And still can't beat Chinese models but good to see European in the game.
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.
if it's not available yet why have a 'try it today' header at all?
> "Try it today"
>
> There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Surprisingly it only supports reasoning "none" or reasoning "high".
That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.
The high bicycle frame is better then the none one though.
Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )
I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."
If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?
SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).
Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.
Just ran this through our data analytics benchmark (I work at Plotly).
It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift.
It's not on the Pareto curve yet, but it's good enough for data analytics, and at this rate I suspect it'll be excellent in another few months.
Full write up: https://plotly.com/blog/mistral-large-4-plotly-data-analytic...
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.
Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.
Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.
This is a pretty grim prognosis for European AI.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
The blog post https://mistral.ai/news/mistral-large-4/
Very nice progress. Also I respect putting Kimi on those charts. Regardless of if they are beating the Pareto frontier (not now), model diversity is a good thing for humanity — I’m hopeful for the team to keep increasing their gains.
Open weight, European, competes with GLM-5.3 on cybersecurity. What's not to like?
Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.
This is awesome, one of the coolest Pokemon ever too for those that don't follow that universe :)
https://bulbapedia.bulbagarden.net/wiki/Lechonk_(Pok%C3%A9mo...
The main thing I always get away from the comparison tables of these "big" models, is how well Deepseek v4.1 Flash performs. While still being the cheapest model by a long shot.
-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience
That's not particularly great.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
I tried it a bit and I like it! It is very fast via openrouter (significantly better than Kimi K3) on webui. Very verbose and starts to forget instructions after awhile it seems, but it gave me quite a lot of good info during a half an hour chat on C and embedded programming. I think I will keep urimg this as my main assistant for few weeks.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
Excited to hear this!
I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.
This model looks reasonably cheap. Though not deepseek levels.
Going to test it with Hermes, wondering where it will land in term of capability.
Bon chance, Mistral!
I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).
Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).
After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.
Not sure where to jump.
Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.
One step closer to Le Chaton Fat.
Refreshing to see this.
The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!
But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).
So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.
Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.
I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.
GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.
If the benchmarks are true, I'd be glad to switch entirely to Mistral.
Has anyone tried Mistral for the coding task? How do you find it compare to Claude, Codex or the Chinese models?
Excited to try this. The low costs v. benchmarks alone here are worth a serious test. K3 has been my daily driver for a month or two now and it's dramatically reduced token spend (while not having much of a negative impact on productivity).
This was the era of the AI race I was waiting for.
How could I resist switching to a model named after my cat!?
Le Chaton Fat is here!
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
Maybe I'm missing something. Doesn't seem super impressive to me. A proprietary model with performance comparable to GPT-6 Luna and Deepseek 4.1 Flash, but at a higher price than either. The main selling point is that it's made in Europe... not very compelling, globally. I suppose maybe there is some niche where European-hosted open-weight models aren't enough to satisfy some EU regulation, where only the use of European-trained models is in compliance, but as a non-European I have no idea what that niche would be.
Side note: Wish this thread was more focused on talking about the model instead of debating about China and America. Whatever happened to staying on-topic?
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Tais-toi et prends mon argent!
I don't know why, but I personally find Mistral's marketing strategy much more appealing than that of other companies.
For example, there's something about Anthropic's picked design and their little Claude avatars that's unsettling to me.
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...
Mistral Large 4: 1050B, 49 Active
GLM-5.3: 753B, 40 Active
I was hoping for something that hinted at smaller models too, but I guess not.
Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.
From Guillaume
> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
https://x.com/GuillaumeLample/status/2107461898127954001
Still second most expensive open-weight model. I don't care about cybersecurity index. And still can't beat Chinese models but good to see European in the game.
Off Topic - The Mistral website - Really nice design. My guess, built by a human.
Something looks off in artificial analysis. Benchmarks aren’t everything, but not even close to the Pareto https://artificialanalysis.ai/models/mistral-large-4?cost=in...
I guess lots of token usage.
Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)
We're probably fast approaching the scenario where the cheapest models will win out.
Is that a reference to LeChuck in Monkey Island? Love that game!
Benchmarks are better than expected! And probably got there without distillation ;)
Title is missing the official model name (Le chonk)
Looks like they are doing 50% off to stay price competitive with DS Flash V4.1
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
1 - https://bench.killswitch-lang.org/
Good to see Europe is at least a little bit still in the game.
https://docs.mistral.ai/inference/model-selection-guide?mode...
Cost is stated at half the price of GLM-5.3, which is quite interesting.
The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Not suitable for my purposes I don't think.
At the end of the blog post we get this nugget.
> The pace of progress from here will be fast. Stay tuned.
Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?
https://venturebeat.com/technology/mistral-debuts-large-4-le...
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
no hugging face link :( ... but hey its on openrouter yay
Here are my results
https://dach.peerbench.ai/compare?models=mistralai%2Fmistral...
Looks like a bit better than the recent Kolibri-1 but still below Qwen3.8 27B
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
Excited to see this! Nice that they are saying this is just a first step.
Give them more compute!
le chaton fat is real, my life is complete. Benches look crazy good for 1T.
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
Why is distillation weird?
What's up with the name? It reminds me of my teenage self trying to speak in funny memes.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
I love the name! Teasing the ones making fun of them.
dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.
if it's not available yet why have a 'try it today' header at all?
> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Mistral doesn't publish the science.
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
> Trained from scratch
How are they training without pirating the Z library corpus and all that?
Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats
So about 2 or 3 generations behind, just like they were a year ago?
at 200M tokens for the full AI suite run its not token efficient at all
sorting the charts like that gives off weird vibes
https://mistral.ai/news/mistral-large-4/
Half price on open router right now
Number one in Sovereign AI. Join our Discord.
We actually got Le Chaton Fat before GTA 6
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
sorting the charts like that gives off weird vibes
I'm glad they're keeping at it!
I thought lechonk motto was just a meme!
Previous one is barely in top 50 on arena.ai