DEV Community

Cover image for Mistral Large 4 for Developers: What It Is and What to Check Before You Call It
AIHubMix
AIHubMix

Posted on

Mistral Large 4 for Developers: What It Is and What to Check Before You Call It

TL;DR: Mistral Large 4 is a 1.05T-parameter mixture-of-experts model from Mistral AI, in public preview since October 6, 2026. It takes text and images and returns text. It leads on defensive security, visual grounding, and legal agent benchmarks, and is mid-pack on general intelligence. Before you wire it in, three details matter: thinking is on by default and reasoning_effort doesn't turn it off through AIHubMix, the preview can change without notice, and the per-task cost that can be higher than its per-token price suggests.

The facts you'll need in code

Field Value
Model ID (AIHubMix) mistral-large-4-0
Model ID (Mistral) mistral-large-4, alias mistral-large-4-0
Input text + up to 100 images per request
Output text only
Context 1M tokens
Max output 262,144 tokens, reasoning included
Thinking (AIHubMix) on by default, off via reasoning.effort = none
Thinking in response reasoning_details field
Price per 1M tokens $1.36 input, $0.14 cached, $4.18 output

Specs follow Mistral's model card. The open weights are due at the end of October 2026, and the license hasn't been announced.

What it's actually good at

Numbers come from Mistral's launch post (self-reported) and from independent runs by Artificial Analysis and Vals AI. All are preview-stage results.

Task Large 4 Reference
CyberGym-E2E (reproduce + patch a vuln) 82%, #1 MiMo-V2.6-Pro 79%
DIOR-RSVG (grounding in aerial images) 73% GPT-6 Astra 68%
Harvey Legal Agent 15.8%, #6 of 76 best open model
Terminal-Bench 4 (Vals) 22.7%, #21 of 45 Claude Opus 5.5 65.2%
AA Intelligence Index 38 GLM-5.3 45

The security result needs context. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero on CyberGym-E2E because they refuse the task. Large 4 does the work. Mistral also reports that it refuses genuinely harmful cyber prompts more often than any other open model it tested.

So use it for defensive security tooling, object localisation in satellite, aerial, or engineering images, and legal or finance research agents. For terminal-heavy coding agents, test the alternatives first.

Prerequisites

  • An AIHubMix key in AIHUBMIX_API_KEY (see the quick start).
  • The OpenAI Python SDK. AIHubMix serves Large 4 on its OpenAI-compatible Chat Completions endpoint at https://aihubmix.com/v1.

How thinking behaves through AIHubMix

On Mistral's own API, reasoning_effort takes "none" or "high", and "high" turns content into a list of chunks. Through AIHubMix, the behaviour is different. These results come from about 30 requests sent on October 9, 2026, with the prompt "Which is larger, 2*31 or 3*20?":

Request Thinks? Reasoning tokens
reasoning_effort omitted yes 1,492
reasoning_effort = none yes 427
reasoning_effort = high yes 764
reasoning.effort = none (body field) no 0

Each row is a single sample, so the token counts only show that thinking happened, not how much each setting costs. "low" and "medium" were accepted too, and also thought.

What that means in code:

  • content is always a plain string, also during streaming, where it is "" while the model thinks.
  • The thinking arrives in reasoning_details, a list of {"type": "reasoning.text", "text": ...} entries, not in reasoning_content. With the OpenAI SDK, read it from message.model_extra.
  • To turn thinking off, send "reasoning": {"effort": "none"} in the request body.
  • Thinking tokens are counted in completion_tokens and billed as output.
import os

from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)


def ask(messages, think=True, max_tokens=16000):
    """Call Large 4 and return (choice, thinking). Thinking is on unless think=False."""
    response = client.chat.completions.create(
        model="mistral-large-4-0",
        messages=messages,
        max_tokens=max_tokens,  # thinking counts against this
        extra_body=None if think else {"reasoning": {"effort": "none"}},
    )
    choice = response.choices[0]
    details = (choice.message.model_extra or {}).get("reasoning_details") or []
    thinking = "".join(d.get("text", "") for d in details)
    return choice, thinking


messages = [{
    "role": "user",
    "content": "Write a Sigma rule that flags PowerShell launched by Office apps.",
}]
choice, thinking = ask(messages)
print(choice.message.content)
print(choice.finish_reason, len(thinking), "characters of thinking")

# Multi-turn: append the assistant message as returned, reasoning_details included.
messages.append(choice.message.model_dump(exclude_none=True))
messages.append({"role": "user", "content": "Exclude signed corporate add-ins."})
choice, _ = ask(messages)
print(choice.message.content)
Enter fullscreen mode Exit fullscreen mode

Mistral's reasoning guide says to replay the thinking on the next turn, and warns that dropping it "significantly degrades output quality". Through AIHubMix, a replayed message with reasoning_details was accepted every time. The replayed thinking is billed as input on the next turn: in one test, the follow-up request had 948 prompt tokens with the thinking replayed and 57 without it.

Failure modes and fixes

You pay for thinking you didn't ask for
reasoning_effort="none" doesn't stop thinking through AIHubMix. In testing, one short Sigma-rule request with "none" produced 2,059 reasoning tokens. Send "reasoning": {"effort": "none"} in the body instead, as ask(..., think=False) does.

Empty answer, finish_reason == "length"
The thinking used up max_tokens before the answer started. In testing, a 300-token cap on a small logic puzzle returned content as None, with 280 of the 300 tokens spent on thinking. Give the cap headroom, and don't call string methods on content without checking it first.

The thinking is missing from your logs
Your code reads reasoning_content, the field many gateways use. Through AIHubMix, Large 4 returns reasoning_details instead, in both normal and streaming responses.

AttributeError: 'list' object has no attribute ... when calling Mistral directly
On Mistral's own API, reasoning_effort="high" returns content as a list of chunks. LiteLLM, Roo Code, and openlit have all had to patch their Mistral handling for this. The AIHubMix route returns a string, so it doesn't hit this.

Outputs drift week to week
That's expected in a preview. Mistral's lifecycle policy allows silent updates to preview models and gives one month's notice before retirement. Pin a small regression set and rerun it on a schedule.

Cost: check per task, not per token

At $1.36 and $4.18 per million tokens, Large 4 is cheaper per token than GPT-6.1 Sol at $2 and $10. On Vals AI's index, though, one test cost $13.78 with Large 4 and $3.24 with Sol, and Sol scored 61.2% against 48.1%. On Vals's Finance Agent v2, the order flips: $1.20 per test for Large 4 against $6.82 for GPT-6 Astra, at similar scores. Before you commit, measure on your own tasks.

Other things measured, and what wasn't

The same October 9 tests also found:

  • Caching: a 3,927-token prompt had 3,052 tokens served from cache on the repeat, with or without prompt_cache_key. Hits aren't guaranteed: one repeat without the key missed.
  • Images: a 400×300 PNG cost about 176 input tokens. Asked for boxes scaled 0 to 1000, the model returned box_2d coordinates within about 50 units of the true edges.
  • Context: a prompt of 559,988 tokens was accepted, and the gateway's error for a larger one stated a limit of 1,048,576 tokens.

Not verified: whether "low", "medium", and "high" change how much the model thinks (single samples showed no consistent pattern), the benchmark figures above, the active parameter count (Mistral has said both 49B and 52B), and the license for the open weights.

Try it on the AIHubMix model page.

Sources

Top comments (0)