🔍 Read the full analysis: Mistral Large 4 Is Still Working To Catch The AI Frontier on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
Mistral released Mistral Large 4 as an API preview on October 6, 2026, with open weights scheduled for later in the month. Artificial Analysis gives the preview an Intelligence Index score of 38, below leading US models and several Chinese alternatives. The available benchmark and one reviewer’s experience do not establish how it will perform on every workload.
Mistral released Mistral Large 4 Preview through an API on October 6, but an independent benchmark snapshot puts it behind several leading US and Chinese models, raising questions about its fit for demanding work. The model’s open weights have not yet been released; Mistral says they are scheduled for later in October.
Artificial Analysis assigns the preview an Intelligence Index score of 38 in data available October 7. That matches OpenAI’s GPT-6 Luna at maximum reasoning effort and sits just below DeepSeek V4.1 Flash at maximum effort, scored at 39. The same snapshot gives Anthropic’s Claude Opus 5.5 a score of 58, Google’s Gemini 4 Argon 53, OpenAI’s GPT-6.1 Sol 52, Z.ai’s GLM-5.3 45 and Moonshot AI’s Kimi K3 44. Cohere’s Command A+ scores 13.
The scores are index points, not percentages, and do not predict success on a particular task. The evaluated reasoning settings differ, so the models were not tested under identical compute budgets. The developer locations identify the companies, not where individual API requests are processed. Artificial Analysis describes its index as an aggregate measure; it is not a direct test of reliability on every coding, research or agent workflow.
Mistral describes Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. Those are company statements; the preview’s benchmark position and the planned timing of the weights release are separate matters.
MODEL WATCH · OCTOBER 7, 2026
Mistral Large 4 Is Still Working To Catch The AI Frontier
Mistral’s API preview is live, with open weights planned for later in October. In the latest Artificial Analysis snapshot, its score trails several leading alternatives. That benchmark is a useful signal for model selection, but it cannot predict performance on every task.
01 / THE BENCHMARK SNAPSHOT
Several alternatives score higher
Artificial Analysis scores available October 7, 2026. Reasoning settings differ, so models were not evaluated with identical compute budgets.
Near two listed peers
Matches GPT-6 Luna at maximum reasoning effort and is one point below DeepSeek V4.1 Flash at maximum effort.
A dated aggregate measure
The index can help narrow options. It does not directly test reliability across every coding, research, or agent workflow.
Business outcomes
Scores do not establish task success, and company locations do not indicate where individual API requests are processed.
02 / THE RELEASE
A preview before open weights
Developers can use the API today. The planned weight release is a separate milestone, with its exact date and terms still unconfirmed in the supplied material.
Large scale, mixture of experts
Mistral describes Large 4 as its largest model to date: one trillion total parameters and 49 billion active parameters. It accepts text and images.
- Mistral says it was trained on the company’s own infrastructure in Europe.
- The company says development and improvement are continuing.
- Artificial Analysis reports context capacity of about 512,000 tokens.
Test the work you need it to do
A large context window and model size do not establish accurate reasoning across long inputs. Mistral’s coding and professional-work claims need evaluation on relevant tasks.
- Measure constraint following and output verification.
- Track errors across multi-step tool use and long-running tasks.
- Compare cost on matched workloads; supplied material lacks prices and calculation details.
Choose a real task
Use representative coding, research, or agent work.
Set matched conditions
Keep task, tools, budget, and success criteria consistent.
Verify the output
Check factual claims, constraints, and tool decisions.
Compare total cost
Include usage and supervision across the full workflow.
03 / EDITORIAL VIEW & ATTRIBUTION
Signals, with limits
Benchmark placement is one input to a decision. Firsthand impressions can inform questions to test, but they are not controlled comparative evidence.
The model was trained on Mistral’s own infrastructure in Europe, and the company says it is continuing to improve it.Mistral · October 6 announcement, as summarized by ThorstenMeyerAI.com
I would not choose it for demanding agentic work or long tasks when stronger models are available.Thorsten Meyer · editorial judgment
In my own use of this preview, I encountered hallucinations again.Thorsten Meyer · firsthand report
Meyer’s comments reflect his own use, not a controlled evaluation. They do not establish a comparative hallucination rate or prove that the model is less reliable than competitors.
04 / OPEN QUESTIONS
What the snapshot cannot settle
The available material leaves several practical questions unanswered. Treat the index as a dated reference point, not a final verdict.
- No independent, workload-specific results are supplied for coding, research, or long-running agent performance.
- The index score cannot determine whether the preview is dependable for a particular workflow.
- Performance may change before or after the planned weight release; the final release terms are unknown.
- Mistral’s capability claims are not accompanied here by detailed supporting results.
- Cost comparisons depend on workload, usage, and measurement method; underlying prices are not supplied.
- Benchmark scores are a dated snapshot and may change as models and evaluations are updated.
05 / KEY QUESTIONS
What developers can conclude now
The clearest confirmed picture as of October 7: API access is live, the index score is 38, and open weights are still ahead.
Is Mistral Large 4 publicly available?
It is available as a public preview API. Its weights are not yet publicly downloadable; Mistral says they are planned for later in October 2026.
How does it score against other models?
Artificial Analysis gives it 38 points. That trails several named US and Chinese models, matches GPT-6 Luna at maximum reasoning effort, and is one point below DeepSeek V4.1 Flash at maximum effort.
Does the benchmark prove it is unreliable?
No. An aggregate index score does not establish reliability on a particular task. Test the preview against your own requirements and verify its outputs.
What evidence would clarify its fit?
Independent evaluations on agentic coding, research, reliability, and matched-cost workloads—alongside the released weights—would give developers a clearer basis for comparison.
Benchmark Scores Shape Model Choices
For developers choosing a model for multi-step agent work, the result is a reason to test carefully rather than assume that a new release is competitive with the top of the market. An agent may plan, use tools and carry decisions across many steps; an unsupported assumption early in that chain can affect later work. Aggregate benchmark scores can help narrow choices, but they cannot establish whether a model will reliably follow constraints or verify its output in a specific workflow.
Thorsten Meyer, writing on ThorstenMeyerAI.com, says he would not choose the preview for demanding agentic work or long tasks when stronger-scoring alternatives are available. That is an editorial judgment, informed by the benchmark snapshot and his own use, rather than a controlled evaluation. He also reports encountering hallucinations, but says this reflects his experience and is not a comparative study. Readers should not treat it as proof that the model hallucinates more often than competitors.
The score comparison also gives a more limited picture than a simple ranking of national AI industries. Several US and Chinese models score above Mistral, while Cohere’s Canadian Command A+ scores below it. The evidence supports the narrower conclusion that this preview trails several named alternatives on this index. It does not show that every competitor is ahead or that benchmark differences translate directly into business outcomes.
As an affiliate, we earn on qualifying purchases.
A Preview Before Open Weights
The timing matters because Mistral Large 4 is currently available as a preview API, not as a publicly downloadable open-weight model. Mistral announced the API on October 6 and says weights are expected later in October. That means developers evaluating it now are assessing a service-accessible preview, while the future weight release remains a separate milestone. The model may change as Mistral continues its work.
Artificial Analysis reports a context capacity of about 512,000 tokens. A large context window allows more material to be supplied in a request; it does not, by itself, show that a model can reason accurately across that material. Mistral promotes the model for agentic coding and specialized professional tasks, but those capabilities need to be assessed on relevant tasks rather than inferred from model size or context capacity.
The source article also raises cost as a factor and says DeepSeek V4.1 Flash has approximately comparable benchmark intelligence at a much lower measured cost per task. The supplied material does not include the underlying prices or calculation details, so the size and conditions of that cost difference cannot be independently described here. Cost comparisons can depend on workload, usage and the measurement method.
“The model was trained on Mistral’s own infrastructure in Europe, and the company says it is continuing to improve it.”
— Mistral, in its October 6 announcement as summarized by ThorstenMeyerAI.com
AI model performance analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Preview Limits and Open Questions
The available material does not provide independent, workload-specific tests of Mistral Large 4’s coding, research or long-running agent performance. The Intelligence Index score cannot settle whether the preview is dependable for a particular user’s tasks. Meyer’s report of hallucinations is also anecdotal; no comparative rate or testing method is supplied.
It remains unclear whether the model’s performance will change before or after the planned weight release, what the final open-weight terms will be, and how its costs compare across matched workloads. Mistral’s claims about coding and professional tasks have not been substantiated with detailed results in the supplied material. The benchmark scores are a dated snapshot and may change as models or evaluations are updated.
As an affiliate, we earn on qualifying purchases.
Weights and Independent Testing
The next stated milestone is Mistral’s planned open-weight release later in October 2026. The company has not provided, in the supplied material, a confirmed release date or details of the final terms. Developers can evaluate the current API preview, but should distinguish their own task results from broad benchmark rankings and account for the supervision and verification their workflows require.
Further benchmark updates and independent tests on agentic coding, research, reliability and matched-cost workloads would help clarify where Large 4 is useful. Until those results and the weights are available, the clearest confirmed picture is limited: an API preview is live, its index score is 38, and the announced weights release is still ahead.
As an affiliate, we earn on qualifying purchases.
Key Questions
Is Mistral Large 4 publicly available?
As of October 7, 2026, it is available as a public preview API. Its weights are not yet publicly downloadable; Mistral says they are scheduled for release later in October.
How does Mistral Large 4 score against other models?
Artificial Analysis gives the preview an Intelligence Index score of 38. In the October 7 snapshot, that is below several named US and Chinese models, equal to GPT-6 Luna at maximum reasoning effort, and one point below DeepSeek V4.1 Flash at maximum effort. The index is not a direct prediction of task success.
Does the benchmark prove Mistral Large 4 is unreliable?
No. An aggregate score does not prove that the model will fail a particular task. The supplied source includes one reviewer’s report of hallucinations, but that is personal experience, not a controlled comparative study.
What are the model’s size and context capacity?
Mistral describes it as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. Artificial Analysis reports a context capacity of about 512,000 tokens. Neither size nor context capacity alone establishes reasoning accuracy.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
