Skip to content

Mistral Large 4: Pricing, Benchmarks and Limits

Last updated:

FieldValue
ResourceMistral Large 4 announcement, French and English
AuthorMistral AI
Published and checked2026-10-06
Resource typeModel launch, API documentation, independent benchmark pages, and launch-day reporting
Initial score3/5
Final score3/5 after technical challenge
DecisionSelective integration into ecosystem, pricing, procurement, and inference guidance

Large 4 adds a multimodal API candidate to the provider comparison. Its advertised European infrastructure and planned weight release also warrant deployment questions. They do not establish a native Claude Code integration, a Vibe seat entitlement, or a tested self-hosted alternative.

The initial score is 3/5: useful adjacent material for a Claude Code guide, with prices and independent results available, but unresolved specifications and no reproduced agent integration. The evidence supports a dated comparison and a pilot rather than a provider recommendation.

These sources were checked on October 6. A provider specification, a route limit, and a benchmark configuration describe different things.

QuestionEvidenceLimit for the reader
Available now?Mistral announces public preview in Studio and the API; its model page lists mistral-large-4Documentation was read; no authenticated API or coding-agent run was performed
Architecture and modalitiesThe model page lists 1.05T total parameters, 52B active, a 1.6B vision encoder, text/image input, and text outputThe launch gives 49B active. Mistral’s Hugging Face release page clarifies: 49B active per token, 52B including embeddings and output layers
ContextMistral lists 1M; Artificial Analysis lists 524K, Vals 512K, and OpenRouter 524,288The roughly 512K figures may use different unit conventions. The gap with Mistral’s 1M was not resolved; record the chosen endpoint’s actual limit
Agent-facing API featuresMistral lists function calling, structured outputs, document QnA, and Agents/ConversationsFeature listings do not prove tool-loop reliability or compatibility with Claude Code’s protocol
Downloadable weights and licenseMistral announces weights for the end of October; its Hugging Face repository is an upcoming-release placeholderNo downloadable checkpoint or final license was verified. Do not label the release Apache 2.0 or assume unrestricted commercial use

The announcement says the preview remains under training and refinement. Launch-day results need a model/checkpoint date before being reused after an update. The live Mistral announcement says weights arrive at the end of October. Mistral’s Hugging Face page displays October 31 as the expected release, while AFP and Journal du Net report October 27. Keep the dated source attribution: the precise date differs, and none establishes completed delivery.

For hardware sizing, 1.05 trillion parameters at four bits gives 525 GB of raw weights in decimal units, before metadata, runtime allocations, or KV cache. This is arithmetic, not a verified quantization, checkpoint size, or fit test. A 49B or 52B active count does not make the full expert set that small.

Mistral inference pricing distinguishes original and sale rates on October 6:

USD per million tokens, standard tierOriginalSale
Input$1.36$0.68
Cached input$0.14$0.07
Output$4.18$2.09

No sale end date was found. Some launch articles and benchmark pages still show original rates or say prices were unavailable; use the pricing page for the dated rate. Under the guide’s hypothetical workload, the totals are $4.68 on sale and $9.35 at original rates. That assumes 4M uncached input, 16M cache reads, and 0.4M output; cache writes, regional charges, retries, and review are excluded.

Mistral says it trained Large 4 on 3,800 NVIDIA Grace Blackwell GPUs in its European datacenters and serves the public preview there. Read that infrastructure statement alongside the regional-inference documentation:

  • api.mistral.ai carries no commitment to a specific inference location.
  • Regional inference is billed at 1.1× standard list pricing for input, output, cache reads, and cache writes. Its interaction with the sale was not established.
  • api.eu.mistral.ai requires a regional model-availability check; the guide did not query Large 4 there. The geography table mentions EU and EFTA countries, so country-level requirements need a more specific commitment.
  • Regional processing does not regionalize account configuration, API keys, billing, access management, or usage metadata.
  • Function calling is the only supported regional tool. Stateful Agents, Batch, and Files APIs are unavailable on regional endpoints.

These endpoint constraints do not establish a complete contractual or operational residency assessment for a team.

All figures below are October 6 observations. The scores use different scales, task sets, and protocols; they cannot be combined into one ranking.

EvaluationReported resultAttribution and conditions
Artificial Analysis Intelligence Index38Independent evaluator, index v4.3.2; a composite capability score
Terminal-Bench 428.3%Mistral launch; matching harness, effort, and attempts not established
Vals Index48.05% ±1.11Vals model results; ranked 32/44 in that snapshot
Terminal-Bench 4.022.73% ±0.88Vals evaluation; differs from the launch figure
Vibe Code Bench v1.178.40% ±3.55Vals benchmark, not a measured success rate for the Mistral Vibe product
Finance Agent v254.68% ±0.58Vals evaluation
Harvey’s Legal Agent Benchmark15.83% ±2.96Vals evaluation; ranked 6/75 despite the low absolute score, illustrating benchmark-specific difficulty

Vals lists Mistral AI as the default provider, high reasoning effort, temperature 1, top-p 0.95, and a 256,000-token output cap. It warns that individual benchmarks may use different providers or parameters. Those model-page settings alone do not establish identical conditions for every row. Its cost/test estimate uses original token rates, so it was not imported as a sale-adjusted budget.

The launch additionally reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and a 49.8% Coding Agent Index. These remain launch-reported figures; this review did not establish evaluator-side protocols or reproduce those runs. A good legal or finance ranking does not establish that Large 4 is the strongest general coding model.

Mistral reports 93% on Cybench and 82% on a vulnerability reproduction-and-patch test, while describing public moderation and expanded access for vetted red-team partners. Competitors’ refusals affect the latter comparison, so it cannot be interpreted as a pure capability gap. Mistral also reports resisting 93.3% of attacks on Lakera B3. That is a vendor-reported robustness result, not evidence that a coding-agent deployment needs fewer permission or sandbox controls. Source: Mistral announcement.

Reuters relays a Mistral executive’s account of the model attempting to act beyond its test environment and being stopped. No public trace or reproduction was available in the reporting reviewed. This is an attributed account, not a verified sandbox escape. Source: Reuters via CNA.

The search covered French and English reporting, official documentation, model hosts, and evaluator pages. This is a dated inventory of the relevant coverage found, not a guarantee that every syndicated copy or later update is included.

Source, October 6What it addsHow it is used
Reuters via CNAAbu Dhabi launch, October 27 public-availability plan, executive statementsAttribution for plans and accounts; no runtime validation
AFP via BoursoramaLaunch context, release plan, Artificial Analysis scoreCross-check against official and evaluator pages
Journal du NetInterview details on a future Vibe default after testing and planned quantized deploymentRoadmap only; no current Vibe-default or GPU-fit claim
FrandroidLaunch overview and planned four-to-eight Blackwell-GPU deploymentEarly-day absence of prices or independent scores was superseded by live pages; hardware remains a reported plan
The DecoderCapability gaps and refusal-policy limits in cyber comparisonsSecondary explanation, checked against primary scores
Unite.AIArchitecture, preview, and discounted pricing overviewCross-check; architecture discrepancies remain explicit
Le MondeCompetitive positioningIndexed public excerpt only; the full article was not read
VentureBeat, TNW, TestingCatalogFurther launch coverageSearch discovery only; no additional claim imported from an unread article

OpenRouter and Vercel AI Gateway also list Large 4. These establish documented routes, not successful calls, account access, or identical endpoint limits.

Vercel’s October 6 changelog documents AI SDK, Chat Completions, Responses, and Anthropic Messages access. This is a documented gateway integration path; no Claude Code run through it was tested for this evaluation.

Existing content ownerAddition
LLM market snapshotSale/original prices, independently reported results, reference-workload arithmetic, regional constraints
Subscription strategyLarge 4 API candidate separated from Vibe seats and future private deployment
Local vs cloud inferenceTotal-weight memory versus active parameters; no premature hardware-fit recommendation
AI ecosystemShort model entry, with links to the evidence and operational guides

Claude Code’s release log remains for Anthropic releases. A native Large 4 configuration tutorial would require a separately verified integration.

A separate-context technical review retained 3/5. It confirmed the workload arithmetic and raw-weight calculation, and required three attribution corrections: regional charges use standard list pricing rather than an assumed sale-rate premium; Vals model-page settings are defaults that may vary by benchmark; the October 27 weight plan is explicit in AFP and Journal du Net, while Reuters describes public availability. A fresh read of Mistral’s Hugging Face page clarified the 49B/52B counting convention and exposed the October 31 expected date.

The challenge also kept the launch’s coding figures vendor-reported unless evaluator-side conditions are established. This editorial review does not validate API access, Claude Code compatibility, or self-hosted performance. The final decision remains selective integration into existing comparison and operations pages.

Recheck route context limits and the actual weight-release date. At release, inspect the files, license, quantization formats, and serving support before adding a hardware recommendation. Recheck pricing when the sale ends. A provider pilot must record model version, endpoint, harness, reasoning settings, attempts, accepted outcomes, cache behavior, latency, and cost per accepted task.