What Mistral Large 4 Gets Right Outside The US And China—and What It Doesn’t
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Mistral Large 4 Gets Right Outside The US And China—and What It Doesn’t on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index, a sharp improvement on its predecessor and a strong result for a model developed outside the United States and China. The same benchmark places it below leading US and Chinese models, while its reported task cost and output volume raise questions for agentic workloads.

Mistral has released Large 4, a research-preview model that scored 38.4 on Artificial Analysis’s Intelligence Index—a substantial rise from the company’s previous flagship, but below the leading US and Chinese models in the same benchmark. The result makes Mistral a prominent European contender, while leaving open whether Large 4 is competitive enough in capability, cost and reliability for demanding business uses.

Artificial Analysis’s Intelligence Index v4.3.2 gives Large 4 a score of 38.4. The same source lists leading US systems between 51.8 and 57.6, while several Chinese models also score above Mistral, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The benchmark data therefore supports a narrower claim: Large 4 is a strong model from outside the US and China, but it does not lead the overall field.

The release marks a sizeable improvement over Mistral’s earlier results on the same index: Large 3 scored 9 and Medium 3.5 scored 14. Mistral says Large 4 has one trillion parameters, with 49 billion active, accepts text and images, and has a 512,000-token context window. It is available as a research preview through Mistral’s API. The company says it plans to release the model weights at the end of October; until then, the weights are unavailable and the licence has not been published, according to the source.

Artificial Analysis reports standard API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. The source calculates a cost of $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models score higher on the index, at 41.8 and 39.5 respectively. Mistral is offering a 50% discount for the first two weeks, according to the supplied material.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research preview, with benchmark data showing a major improvement over its earlier models but continued gaps against leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Agentic Work

The benchmark matters because Artificial Analysis’s index includes tasks involving agentic knowledge work, software workflows and coding, not just short question-and-answer exchanges. Scores are not a direct measure of performance in every customer’s workload, but they provide one comparison point for models intended to complete multistep tasks.

The source also reports that Large 4 generated 200 million output tokens while completing the index, against a median of 81 million for comparable models. That observation may matter to buyers: greater output can increase both cost and latency in workflows with repeated model calls. The task-cost figures provide a more direct comparison, but actual spending will depend on how a customer uses the model, including prompt length, caching and the number of calls.

A lower benchmark score does not by itself establish that a model will fail in production. Still, when a task depends on many sequential steps, mistakes or unnecessary output can accumulate. The source author also reports seeing confident false statements in hands-on testing. That is a personal observation, not a result established by the Artificial Analysis index, and it underscores why buyers need to test reliability on their own tasks before relying on a model in an automated workflow.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Large Jump, Not a Lead

The supplied benchmark comparison places Large 4 far ahead of Mistral’s earlier models on the same index, while keeping it below several current US and Chinese systems. For example, the source lists Anthropic’s Claude Opus 5.5 at 57.6 and Google’s Gemini 4 Argon at 52.6. It lists Chinese models GLM-5.3 and Kimi K3 at 44.8 and 43.6. These figures are specific to Artificial Analysis Index v4.3.2; they should not be treated as a universal ranking across every task or evaluation.

The source frames Large 4 as the strongest model from outside the US and China. That comparison is narrow: the material describes the field of comparable competitors elsewhere as limited, rather than showing that the claim settles which model is best for every use. It also says that Large 4’s weights have not yet been released, so comparisons with open-weight models are prospective until those weights become available.

Mistral says reinforcement learning is still in progress and that scores may change. The current benchmark is consequently a snapshot of a preview model, not necessarily its final performance. The supplied material does not include a full account of the evaluation methods beyond naming the index version and its included task suites.

Amazon

large language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Scores and Reliability

Large 4 remains a research preview, and Mistral says its reinforcement-learning work is continuing. It is not yet clear how much the scores will shift, when customers will receive the promised weights, or what licence will govern those weights. The source gives an end-of-October target for the release, but does not establish that date as a completed or guaranteed release.

The source author’s report of confident hallucinations comes from hands-on testing and is not independently quantified in the material. No test details, sample size or reproducible results are provided for that observation. The benchmark figures also cannot establish how Large 4 will perform across individual companies’ tasks, or whether the reported task costs will match their usage patterns.

Amazon

AI text and image processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Updated Evaluations

The next clear milestone is Mistral’s planned release of Large 4’s weights at the end of October, as stated in the supplied source. The licence and the timing of broad availability remain to be confirmed. Updated Artificial Analysis results could also clarify whether the model’s score changes as Mistral continues reinforcement learning.

For prospective users, the relevant next step is to compare the model against alternatives on their own workloads, with attention to accuracy, false claims, output volume, latency and total API costs. Large 4’s preview results establish a notable improvement for Mistral; they do not settle whether it is the right choice for a particular deployment.

Amazon

AI model cost management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What score did Mistral Large 4 receive?

It scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, according to the comparison table in the supplied material.

Is Mistral Large 4 the leading model overall?

No. The cited comparison places several US and Chinese models above it. The source characterizes Large 4 as a leading model from outside those two countries, a narrower distinction.

Are Large 4’s weights available?

Not yet, according to the supplied material. Mistral plans to release them at the end of October, but the licence has not been published there.

How does its reported task cost compare with some alternatives?

The source calculates $1.13 per Intelligence Index task for Large 4, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. These are benchmark task-cost estimates, not a guarantee of costs for every user.

What is still uncertain about the model?

Mistral says training is ongoing, so scores may change. The weights’ release timing and licence remain pending in the source material, and the author’s reported hallucinations are anecdotal rather than a quantified benchmark result.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Is Tokenization in LLM? Unlock AI’s Potential on the Blockchain

Tokenization transforms language learning models, paving the way for enhanced AI capabilities on the blockchain—discover what this means for the future of technology.

Mistral Forge: The Practical Choice For Owning Your AI Models

Mistral announces Forge, a comprehensive platform for building and managing domain-specific AI models, emphasizing ownership and control for enterprise users.

What Is Air-Gapped

What is an air gap and how does it enhance security for sensitive systems while posing unique challenges? Discover the intricacies of this robust protective measure.

Software-Defined Warfare: How Ukraine’s Delta Turned the Battlefield Into a Shared, Real-Time Map

Ukraine’s Delta is a cloud-based, browser-run battlefield system that fuses real-time intelligence, revolutionizing modern warfare and strategic agility.