The Upside And Agent Trade-Offs Of Mistral Large 4 Outside The US And China
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Upside And Agent Trade-Offs Of Mistral Large 4 Outside The US And China on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a substantial increase from the prior Mistral Large model but below current leading US and Chinese models. The source says its benchmark task cost and high output volume may make it a poor fit for some agent workflows; weights, licensing and final performance remain unsettled.

Mistral released Large 4 as a research public preview, positioning the French model as a high-performing option from outside the United States and China. On Artificial Analysis Intelligence Index v4.3.2, it scored 38.4, a sharp improvement on Mistral’s previous Large model but below the leading US and Chinese models in the source’s comparison.

The model has 1 trillion total parameters, with 49 billion active, and accepts text and images while generating text. Mistral lists a 512,000-token context window. The preview is available through Mistral’s API; the company has promised to release the weights by the end of October. Until then, Large 4 is proprietary, and its licence has not been published, according to the source.

At standard API rates, the source lists prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. Mistral is offering a 50% discount for the first two weeks. The source reports a cost of $1.13 per Artificial Analysis benchmark task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models scored 41.8 and 39.5, respectively, on the same index.

Large 4’s score is up from 9 for Mistral Large 3 on the same index version; Medium 3.5 scored 14, according to the source. That marks a major improvement in Mistral’s benchmark performance. It does not put Large 4 among the top models in the cited table: the leading US model scored 57.6, while the highest-listed Chinese model scored 44.8.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral has released Large 4 as a research public preview, posting a score of 38.4 on Artificial Analysis’s Intelligence Index while leaving its weights and licence for a later release.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Cost and Reliability in Agent Work

The results matter most for teams choosing models for multi-step agent tasks, where a system repeatedly calls a model to plan, use tools and act on intermediate results. The source notes that the Intelligence Index includes agent-focused benchmarks for knowledge work, software workflows and coding. A lower score does not translate directly into a specific failure rate, but it is relevant evidence for buyers evaluating demanding workflows.

The source also reports that Large 4 generated 200 million output tokens across the index evaluations, compared with a median of 81 million for comparable models. If that verbosity carries over to a customer’s use, it may add cost and latency across repeated agent steps. The author separately describes seeing confident false statements in hands-on testing; that is an attributed observation, not a published Artificial Analysis measurement.

For organisations outside the US and China, Mistral’s progress may offer another supplier with a European base. But the source’s comparison suggests buyers should weigh that consideration against benchmark performance, per-task price and error handling, rather than treating geography as a substitute for task-specific testing.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Large Jump From Mistral 3

The source compares model results using Artificial Analysis Intelligence Index v4.3.2, described as the current version, and says this allows like-for-like comparison. In its table, US models lead, with Claude Opus 5.5 at 57.6 and several other US models above 50. Chinese models listed include GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. Large 4’s 38.4 places it below those entries but above some older or lower-scoring models in the table.

The source frames Large 4 as the leading model from outside the US and China, while cautioning that this comparison is narrow: it says few other labs in that geographic group compete at the same frontier tier. It also reports that reinforcement learning is still underway, and Mistral says scores may change. The release therefore combines a significant jump over Mistral’s earlier results with a preview status and a benchmark position that remains behind the models at the top of the cited rankings.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com author

Amazon

large language model licensing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licence and Real-World Results

Several details remain open. The weights have not yet been released, and the licence is unpublished, so users cannot yet assess the terms or independently deploy the model from released weights. Mistral has promised weights by the end of October, but the source does not give a confirmed licence or specify whether the timeline could change.

The index score is also provisional in the sense that Mistral says reinforcement learning continues. The source does not establish how closely benchmark token use, latency or observed errors will match performance across different customer workloads. Its report of hallucinations comes from the author’s own tests, while the benchmark comparisons are attributed to Artificial Analysis; those forms of evidence should not be conflated.

The supplied source cuts off while comparing the cost of Large 4 with Gemini 4 Argon, so it does not provide the full comparison or a complete basis for assessing that model’s value. The listed task-cost figures are benchmark-specific, not a guarantee of what every customer will pay.

Amazon

AI token counting tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The October Weights Release

The next stated milestone is Mistral’s planned release of Large 4’s weights by the end of October. Publication of the weights and licence would clarify whether developers can run or adapt the model outside Mistral’s API and under what conditions.

Until then, buyers can test the API preview against their own tasks, tracking not just answer quality but also output-token use, latency, cost and whether errors compound over a sequence of actions. Mistral’s ongoing reinforcement learning may also affect later benchmark results. The available information does not specify a date for a final model or for updated independent evaluations.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. The result is higher than the 9 reported for Mistral Large 3 on the same index version, but below leading US and Chinese models in the cited table.

Can developers download Large 4 now?

Not yet, according to the source. Large 4 is available as a research public preview through Mistral’s API. Mistral has promised to release the weights by the end of October; the licence has not been published.

How does its benchmark task cost compare with some alternatives?

The source reports a cost of $1.13 per index task for Large 4, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. These figures refer to the benchmark evaluation and should not be treated as a fixed price for every customer workload.

Is Large 4 a proven choice for AI agents?

The source does not establish that. Its index includes agent-focused tasks, but benchmark results alone cannot predict performance in every workflow. The author also reports seeing confident false statements in hands-on testing; that observation is not an independent benchmark result.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How A New AI Firm Outmanaged Western Giants Against All Odds

A Chinese AI firm, Moonshot’s Kimi K3, outmanaged four Western frontier models in a live business simulation, winning against all odds.

Engineering Is Automated. Research Is the Residual.

Recent benchmarks show AI can now automate most AI engineering tasks, leaving research as the remaining challenge, with implications for AI development timelines.

Why AI Will Be Central To Your Tech In 2026

Exploring how artificial intelligence is set to become a core component of consumer and enterprise technology by 2026, transforming workflows and daily life.

8 Key AI Developments You Can’t Miss In 2026

Discover the eight key AI advancements in 2026 that are shaping technology, industry, and society, with confirmed updates and ongoing developments.