AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 And AI Agents: A Promising Model With A Key Limitation on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 earned a score of 38.4 on Artificial Analysis’ Intelligence Index v4.3.2, a sharp rise from its predecessor but below the leading US and Chinese models listed. Its benchmark cost, high output-token use and reported confident errors make its suitability for long-running AI agents uncertain.

Mistral released Large 4 as a Research Public Preview, and Artificial Analysis’ Intelligence Index v4.3.2 gives it a score of 38.4—a major jump over the company’s earlier models, but below the leading US and Chinese systems in the published comparison. The result makes Mistral a stronger European model contender, while raising practical questions about whether its cost and reliability suit long-running AI agents.

The source describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, text-and-image input, text output and a 512,000-token context window. Mistral has released it through its API as a Research Public Preview. The company says model weights are planned for the end of October, but the supplied material does not specify the year. Until those weights are available, the model is proprietary; its licence has not been published, according to the source.

On the current Artificial Analysis Intelligence Index, Large 4 scored 38.4. The source compares that with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version. In the listed results, Large 4 trails US flagships such as Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7. It also sits below several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. These are benchmark results, not a guarantee of performance on every task.

Mistral’s listed API rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million. The source says the company offered 50% off for the first two weeks. It reports a benchmark cost of $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Artificial Analysis’ task-cost figure is a benchmark comparison; actual costs depend on usage and workload.

At a glance
analysisWhen: Released yesterday, according to the su…
The developmentMistral released Large 4 as a research preview, with independent benchmark data showing substantial improvement over earlier Mistral models but gaps in capability and cost efficiency for agentic work.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Work Faces Cost and Reliability Tests

The index includes tests of agentic knowledge work, software workflows and coding, so its score is relevant to buyers evaluating models for multi-step tasks, not just conversational answers. Still, one composite benchmark cannot predict how a model will perform in a particular company’s tools, data or safeguards. The 38.4 result is evidence of progress, not a complete measure of operational fitness.

For agents, errors can carry forward: a fabricated detail used early in a workflow may shape later decisions or actions. The source author says hands-on testing found confident hallucinations in Large 4. That is an attributed observation, not a finding from the Artificial Analysis index. It warrants evaluation in the intended workflow before the model is trusted with consequential tasks.

Cost and output volume also matter because an agent may make repeated calls. Artificial Analysis data cited by the source show Large 4 generated 200 million output tokens across the index, versus a median of 81 million for comparable models. That is a benchmark observation, not a universal prediction of token use. But together with the stated per-task cost, it gives procurement teams a reason to measure completion rates, latency and total workflow expense rather than comparing token prices alone.

Amazon

AI development API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise From Mistral’s Prior Scores

The source frames the launch as a notable step for a European AI developer. On the same version of the Artificial Analysis Intelligence Index, Mistral Large 3 scored 9 and Medium 3.5 scored 14, compared with Large 4’s 38.4. That is a substantial improvement within the company’s lineup. It does not, by itself, establish that Large 4 matches the leaders or performs better across all uses.

The source also describes Large 4 as the most intelligent model outside the United States and China, based on the index comparison. That claim depends on the set of models and countries included; it should not be read as a broad measure of every AI system worldwide. The supplied benchmark table places the model behind multiple US and Chinese releases, including newer systems than some of the alternatives used in Mistral’s launch comparisons.

Mistral says reinforcement learning is still underway and that scores may change. That makes the published result a preview-stage snapshot, rather than a final assessment of a fixed model. The comparison is tied to Artificial Analysis Intelligence Index v4.3.2, as reported in the source.

“Reinforcement learning is still running, so scores may move.”

— Mistral, as reported in the supplied source

Amazon

large language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results Leave Open Questions

The model is still a Research Public Preview, and the source says reinforcement learning is ongoing. It does not provide a date with a year for the planned end-of-October weight release, nor does it give the licence terms. Buyers therefore cannot yet assess the published weights or their permitted uses based on the information supplied.

The source does not describe the author’s hands-on test methods, sample size or prompts, so the reported hallucinations should be treated as an attributed observation rather than a measured rate. Nor does the supplied material establish how Large 4 performs across specific agent frameworks, tool environments or safeguards. Benchmark rankings and task costs can change with model updates, test design and usage patterns.

Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Updated Benchmarks Ahead

The next stated milestone is Mistral’s planned release of Large 4 weights at the end of October; the source does not identify the year or confirm that the date is firm. Publication of the weights and licence would clarify whether developers can run the model themselves and under what terms.

Artificial Analysis scores may also change as Mistral’s reinforcement learning continues and the model is evaluated again. For organizations considering agents, the practical next step is to test the preview against their own tasks, tracking successful completion, unsupported claims, token use and total cost before assigning it to longer workflows.

Amazon

AI token usage monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral release?

Mistral released Large 4 as a Research Public Preview through its API. The source describes it as a one-trillion-parameter, text-and-image-input model with a 512,000-token context window.

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the supplied source. The result is above the cited scores for Mistral Large 3 and Medium 3.5, but below several US and Chinese models in the comparison.

Is Large 4 ready for AI agents?

The available evidence does not establish that it is ready for every agent workflow. The benchmark covers agent-related tasks, but the source also reports high output-token use and an author-observed instance of confident hallucination. Buyers should test it in their own setting.

When will Mistral release the model weights?

The source says Mistral plans to release the weights at the end of October, but it does not specify the year or confirm a firm date. It also says the licence had not been published at the time described.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Human Development Index & Statistics: Measuring Progress

Just when you think economic growth tells the whole story, the Human Development Index reveals a deeper, more comprehensive picture of true progress around the world.

Smart CCTV And AI Near-Miss Detection: Safer Warehouses Start Here

New AI-powered system analyzes existing warehouse CCTV feeds to identify near-misses, enhancing safety and reducing injuries.

A Peek Into Reddit’s Anti-spam Internals

Reddit has publicly shared details about its internal anti-spam systems, offering insights into how it detects and manages spam on the platform.

The Latest iPhone 18 News, Leaks, And Rumors: Release Date, Price Increases, And Color Options

Confirmed details on the iPhone 18 release date, expected price increases, and rumored features, based on recent leaks and rumors.