TL;DR: Mistral Large 4 is a 1.05T-parameter mixture-of-experts model from Mistral AI, in public preview since October 6, 2026. It takes text and images and returns text. It leads on defensive security, visual grounding, and legal agent benchmarks, and is mid-pack on general intelligence. Before you wire it in, three details matter: thinking is on by default and reasoning_effort doesn't turn it off through AIHubMix, the preview can change without notice, and the per-task cost that can be higher than its per-token price suggests.
The facts you'll need in code
| Field | Value |
|---|---|
| Model ID (AIHubMix) | mistral-large-4-0 |
| Model ID (Mistral) | mistral-large-4, alias mistral-large-4-0 |
| Input | text + up to 100 images per request |
| Output | text only |
| Context | 1M tokens |
| Max output | 262,144 tokens, reasoning included |
| Thinking (AIHubMix) | on by default, off via reasoning.effort = none |
| Thinking in response | reasoning_details field |
| Price per 1M tokens | $1.36 input, $0.14 cached, $4.18 output |
Specs follow Mistral's model card. The open weights are due at the end of October 2026, and the license hasn't been announced.
What it's actually good at
Numbers come from Mistral's launch post (self-reported) and from independent runs by Artificial Analysis and Vals AI. All are preview-stage results.
| Task | Large 4 | Reference |
|---|---|---|
| CyberGym-E2E (reproduce + patch a vuln) | 82%, #1 | MiMo-V2.6-Pro 79% |
| DIOR-RSVG (grounding in aerial images) | 73% | GPT-6 Astra 68% |
| Harvey Legal Agent | 15.8%, #6 of 76 | best open model |
| Terminal-Bench 4 (Vals) | 22.7%, #21 of 45 | Claude Opus 5.5 65.2% |
| AA Intelligence Index | 38 | GLM-5.3 45 |
The security result needs context. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero on CyberGym-E2E because they refuse the task. Large 4 does the work. Mistral also reports that it refuses genuinely harmful cyber prompts more often than any other open model it tested.
So use it for defensive security tooling, object localisation in satellite, aerial, or engineering images, and legal or finance research agents. For terminal-heavy coding agents, test the alternatives first.
Prerequisites
- An AIHubMix key in
AIHUBMIX_API_KEY(see the quick start). - The OpenAI Python SDK. AIHubMix serves Large 4 on its OpenAI-compatible Chat Completions endpoint at
https://aihubmix.com/v1.
How thinking behaves through AIHubMix
On Mistral's own API, reasoning_effort takes "none" or "high", and "high" turns content into a list of chunks. Through AIHubMix, the behaviour is different. These results come from about 30 requests sent on October 9, 2026, with the prompt "Which is larger, 2*31 or 3*20?":
| Request | Thinks? | Reasoning tokens |
|---|---|---|
| reasoning_effort omitted | yes | 1,492 |
| reasoning_effort = none | yes | 427 |
| reasoning_effort = high | yes | 764 |
| reasoning.effort = none (body field) | no | 0 |
Each row is a single sample, so the token counts only show that thinking happened, not how much each setting costs. "low" and "medium" were accepted too, and also thought.
What that means in code:
-
contentis always a plain string, also during streaming, where it is""while the model thinks. - The thinking arrives in
reasoning_details, a list of{"type": "reasoning.text", "text": ...}entries, not inreasoning_content. With the OpenAI SDK, read it frommessage.model_extra. - To turn thinking off, send
"reasoning": {"effort": "none"}in the request body. - Thinking tokens are counted in
completion_tokensand billed as output.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
def ask(messages, think=True, max_tokens=16000):
"""Call Large 4 and return (choice, thinking). Thinking is on unless think=False."""
response = client.chat.completions.create(
model="mistral-large-4-0",
messages=messages,
max_tokens=max_tokens, # thinking counts against this
extra_body=None if think else {"reasoning": {"effort": "none"}},
)
choice = response.choices[0]
details = (choice.message.model_extra or {}).get("reasoning_details") or []
thinking = "".join(d.get("text", "") for d in details)
return choice, thinking
messages = [{
"role": "user",
"content": "Write a Sigma rule that flags PowerShell launched by Office apps.",
}]
choice, thinking = ask(messages)
print(choice.message.content)
print(choice.finish_reason, len(thinking), "characters of thinking")
# Multi-turn: append the assistant message as returned, reasoning_details included.
messages.append(choice.message.model_dump(exclude_none=True))
messages.append({"role": "user", "content": "Exclude signed corporate add-ins."})
choice, _ = ask(messages)
print(choice.message.content)
Mistral's reasoning guide says to replay the thinking on the next turn, and warns that dropping it "significantly degrades output quality". Through AIHubMix, a replayed message with reasoning_details was accepted every time. The replayed thinking is billed as input on the next turn: in one test, the follow-up request had 948 prompt tokens with the thinking replayed and 57 without it.
Failure modes and fixes
You pay for thinking you didn't ask for
reasoning_effort="none" doesn't stop thinking through AIHubMix. In testing, one short Sigma-rule request with "none" produced 2,059 reasoning tokens. Send "reasoning": {"effort": "none"} in the body instead, as ask(..., think=False) does.
Empty answer, finish_reason == "length"
The thinking used up max_tokens before the answer started. In testing, a 300-token cap on a small logic puzzle returned content as None, with 280 of the 300 tokens spent on thinking. Give the cap headroom, and don't call string methods on content without checking it first.
The thinking is missing from your logs
Your code reads reasoning_content, the field many gateways use. Through AIHubMix, Large 4 returns reasoning_details instead, in both normal and streaming responses.
AttributeError: 'list' object has no attribute ... when calling Mistral directly
On Mistral's own API, reasoning_effort="high" returns content as a list of chunks. LiteLLM, Roo Code, and openlit have all had to patch their Mistral handling for this. The AIHubMix route returns a string, so it doesn't hit this.
Outputs drift week to week
That's expected in a preview. Mistral's lifecycle policy allows silent updates to preview models and gives one month's notice before retirement. Pin a small regression set and rerun it on a schedule.
Cost: check per task, not per token
At $1.36 and $4.18 per million tokens, Large 4 is cheaper per token than GPT-6.1 Sol at $2 and $10. On Vals AI's index, though, one test cost $13.78 with Large 4 and $3.24 with Sol, and Sol scored 61.2% against 48.1%. On Vals's Finance Agent v2, the order flips: $1.20 per test for Large 4 against $6.82 for GPT-6 Astra, at similar scores. Before you commit, measure on your own tasks.
Other things measured, and what wasn't
The same October 9 tests also found:
-
Caching: a 3,927-token prompt had 3,052 tokens served from cache on the repeat, with or without
prompt_cache_key. Hits aren't guaranteed: one repeat without the key missed. -
Images: a 400×300 PNG cost about 176 input tokens. Asked for boxes scaled 0 to 1000, the model returned
box_2dcoordinates within about 50 units of the true edges. - Context: a prompt of 559,988 tokens was accepted, and the gateway's error for a larger one stated a limit of 1,048,576 tokens.
Not verified: whether "low", "medium", and "high" change how much the model thinks (single samples showed no consistent pattern), the benchmark figures above, the active parameter count (Mistral has said both 49B and 52B), and the license for the open weights.
Try it on the AIHubMix model page.
Sources
- Introducing Mistral Large 4 (Mistral AI)
- Reasoning (Mistral Docs)
- Mistral Large 4.0 on AIHubMix
- Mistral Large 4 analysis (Artificial Analysis)
- Mistral Large 4 benchmarks (Vals AI)
Top comments (0)