DEV Community

Ashraf
Ashraf

Posted on

Mistral Large 4: 1T Parameters, 1M Context, and the Open-Weight Gap Just Got Smaller

Mistral Large 4 landed on October 6 and sat at 1,790 points and 1,000+ comments on Hacker News. That's not hype noise. Let's strip the launch post down to what matters if you ship software.

The specs

  • ~1T total parameters, mixture-of-experts. The announcement says 49B active; the API docs say 52B active / 1.05T total plus a 1.6B vision encoder. Pick your number. Either way, you pay compute for ~5% of the weights per token.
  • 1M token context (per the API docs).
  • Natively multimodal, 160+ languages claimed.
  • Hybrid instruct + reasoning in one model. No separate "thinking" SKU.
  • Trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.
  • API: mistral-large-4, public preview. Function calling, structured outputs, batch, agents, and built-in tools are supported.

Pricing

$1.36 / M input, $4.18 / M output.

Mistral Large 3 was listed at $2 / $5. A commenter on HN read this as Mistral discounting to stay competitive with DeepSeek's cheap tier. That's their interpretation, not something Mistral says. But the direction is clear: frontier-ish open-weight models are getting cheaper, fast.

Benchmarks (vendor-reported)

Area Result
DeepSWE v1.1 61.7%
Terminal-Bench 4.0 28.3%
SWE-Atlas-QnA 59.4%
Cybench 93%
Vulnerability reproduction 82%
AutomationBench 59.9%
Human coding eval (Surge AI) 3.74 / 5

The one number I'd actually trust is the honest one: in the Surge AI human coding eval, ML4 scored 3.74 vs 4.22 for Claude Opus 5. Second of five models, but a real gap behind the leader. Mistral published that. Credit where due.

Everything else is vendor-run. Treat it like any launch-day chart.

What the launch post doesn't tell you

  1. The license. The docs call it "open-weight." Neither the announcement nor the docs I read spell out the terms. Apache 2.0 and a restricted community license are very different things for your product. Weights are promised for the end of October. Read the license before you plan around it.
  2. Can you run it? 1T parameters is ~2 TB at 16-bit, ~1 TB at 8-bit, even with MoE. "Open weights" does not mean "runs on your box." Realistically this is multi-node, or you wait for quantizations. HN had the same question.
  3. Terminal-Bench 4.0 at 28.3%. Agentic terminal work is where models still fall over. If your use case is autonomous ops, test it yourself.

Try it

# /// script
# dependencies = ["mistralai"]
# ///
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.chat.complete(
    model="mistral-large-4",
    messages=[{"role": "user", "content": "Review this diff for race conditions: ..."}],
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

Run it with uv run script.py. Check the model identifier in the docs, since it's still a public preview and the exact name may shift.

My take

The interesting part isn't that Mistral beat anyone. It's the shape of the market. A European lab, on its own hardware, ships a 1M-context multimodal MoE at roughly two-thirds of frontier closed-model pricing, and the one independent human eval puts it close, not equal. As one HN commenter put it, since Chinese labs publish their research, it would have been odd if Mistral didn't catch up.

That's the story. The moat on raw model quality keeps shrinking, so the decision moves to license, latency, data residency, and cost per task. Mistral is betting on the European-hosting angle. For regulated workloads that's a real feature.

What I'd do this week:

  • Run your own eval set against mistral-large-4. Not theirs.
  • Compare cost per successful task, not per token.
  • Hold off on any architecture decision until the license text is public.

Sources: Mistral announcement, Mistral docs, HN thread.

Top comments (0)