If you last checked in on open-weight models a year ago, the picture has inverted. The clearest takeaway from the first half of 2026 is that choosing an open model is no longer a single-model decision. It is a portfolio decision, made across a crowded field of frontier-class releases documented in Digital Applied's H1 2026 retrospective. Here is what actually shipped, who shipped it, and what "open" turned out to mean.
Ten weeks, five Chinese labs
Between April and June 2026, five Chinese labs shipped open-weight frontier models back to back: DeepSeek V4, Kimi K2.6, GLM-5.1, Qwen 3.6, and MiniMax M3, according to Sunset Browser's roundup. April alone was absurd across the whole industry. DeepSeek V4, Qwen 3.5-Omni, Gemma 4, Meta Muse Spark, GPT-6, and Claude Opus 4.7 all landed within a single calendar month (Fazm).
DeepSeek V4 set the pace
DeepSeek V4 shipped on 24 April 2026 under an MIT licence in two tiers. V4-Pro sits at 1.6T parameters with 49B active; V4-Flash at 284B with 13B active. Both carry a 1M-token context window and 384K maximum output (DataMy). Flash is the one you can actually run: a Mixture-of-Experts build aimed at coding, tool use, and agentic workflows, downloadable through LM Studio. Moonshot's Kimi K3 is similarly a download away on Hugging Face, deployable with a standard vLLM server.
Four labs, four meanings of "open"
Openness itself is uneven. Tracking four flagship open-weight releases since mid-June, the lag between a model appearing on a public API and its weights being published ran zero days for DeepSeek V4 Flash and GLM-5.2, and eleven days for Moonshot's Kimi K3 (Alephant). Alibaba has institutionalised the split: since 2026 it runs a deliberate two-track strategy, open weights for the community and on-premise deployment, with proprietary Plus and Max tiers sold through Alibaba Cloud (Tongyis). Even Thinking Machines joined the movement, releasing its Inkling weights openly on Hugging Face. That is a real shift in transparency and licensing, though it comes with notable restrictions (Checkingmarket).
Google came back, and everyone else kept coming
Google DeepMind's Gemma 4 arrived as a four-model open-weight family combining Thinking Mode reasoning, native multimodality, 256K context windows, and Apache 2.0 licensing. It benchmarked 89.2% on AIME 2026 with an 86.4% score on a second benchmark (LinkedIn).
Zhipu's GLM line kept the pressure on. GLM-5.2, released on 13 June 2026, brought MIT-licensed weights, a 1M-token context window, and top-of-open-source coding scores (Kie.ai). Its successor GLM-5.3 Flash is flagged as one of the most important open-weight releases to watch in late 2026, though reviewers caution it is not an uncensored chat model and not the easiest local model to fit on a consumer GPU (AirMore).
Elsewhere, the releases kept stacking up. Tencent's Hy4 preview landed on 28 August 2026 as an open-weight model you can download and run on your own hardware (LLM Stats). Nvidia shipped Nemotron 3.5 Lightning on 11 August 2026, a free open-weight model aimed at autonomous agent workloads (Andrew.ooo). Mistral confirmed a sparse Mixture-of-Experts family entering early access in July 2026, its boldest open-weight bet yet, backed by a €4bn EU data-centre buildout (AI TechConnect).
Switzerland's Apertus got a more sober verdict one year on: "When it comes to agentic capabilities, despite considerable progress, Apertus is not yet at the level of other open-weight models," Binaghi told SwissInfo.
The policy fight went public
On 27 July 2026, Dario Amodei published a direct rebuttal to industry chatter suggesting Anthropic secretly wants open models banned. The clarification reshapes how founders, policymakers, and builders should read the open-weights debate heading into late 2026 (Kalinga.ai).
Choosing between them
The comparison infrastructure has matured alongside the models. BenchLM maintains direct comparison tables across GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open models. HypeBench runs a comparison of DeepSeek V4 Pro, GLM-5.3, Qwen 3.8 Max, and Kimi K3 with workload tests and access checks. Artificial Analysis computes its Intelligence Index from output tokens per task divided by speed, weighted across benchmarks.
Older standbys still hold their niches. Qwen 3.5 leads most open-source benchmarks under Apache 2.0 licensing, DeepSeek R1 dominates math and reasoning with a 97.3 on AIME, and Llama 3.3 70B remains the safe, well-supported generalist (Toolhalla). A Turkish roundup ranks Llama 4, Qwen 3.6, and DeepSeek V4 by benchmarks, hardware requirements, and true cost (GetAIPerks). Capability maturity is uneven down the stack: Mistral Large 2 and Qwen 2.5 call tools natively, Llama 3.1 weakly, DeepSeek V3 at a mid level (Omeronal).
The loose ends
Meta is the big one. Muse Spark shipped inside that crowded April window, but nothing in the coverage here establishes whether Meta's 2026 models are open-weight at all. Until that question gets an answer, Llama's open-source standing rests on older releases like Llama 4 and 3.3.
The second loose end is smaller but worth flagging. A few headline numbers arrived without their backstory: which benchmark sits behind Gemma 4's 86.4% score, and what the full weight-lag figures look like for the four-release tracker beyond the three models named. A number that cannot say where it came from cannot tell you much.
The quiet side of open
There is a whole side of this story the headlines skip, and it falls into four clusters: the genuinely open camp (IBM Granite 4.x, Hugging Face's SmolLM3, Ai2's OLMo line, Stanford's Marin, OpenEuroLLM, Sarvam, plus GLM-5.3's full-weights release), open multimodal, video, and robotics models, two major Chinese labs (Baidu and ByteDance), and the open-weights policy track. Two of those blind spots, IBM's Granite and Hugging Face's SmolLM3, have documented substance behind them, and both get a closer look below.
Granite 4.x, the omission with the most substance
Granite is the omission with the most substance behind it, and its release profile looks nothing like the hype cycle. Granite 4.0 runs a hybrid Mamba and transformer architecture and ships with signed model weights and documented training data, positioned explicitly for sensitive use in healthcare and the public sector (Geeky Gadgets). Granite 4.1 is the Apache 2.0 dense family in 3B, 8B, and 30B sizes, and the 3B runs from roughly 2 GB of memory at Q4 quantization, with hardware-fit guidance for each size (The AI Bench). Context stretches to 131k across the generation: the 4.1 8B supports up to 131,072 tokens (free.ai), and the 4.1 30B sits on Artificial Analysis' leaderboard of 250+ models at 131k (Artificial Analysis). Granite 4.2 8B is IBM's dense reasoning model, served on OpenRouter at $0.10 per million input tokens and $0.15 per million output tokens (OpenRouter). The 1B and 3B variants are IBM's first mixture-of-experts Granite models, designed for low-latency use (Ollama), with a lineage that runs back through Granite 3 MoE, which activated about 3B of a 10 to 15B total per token (RunLocalAI).
Want to run it yourself? Granite runs locally through Ollama on modest hardware (LinkedIn), a free-tier granite-4.0-micro endpoint is available through a single OpenAI-compatible API (UnoRouter), and the brand stretches into small specialists: a 258M-parameter Granite Docling model with a hosted demo (Free2AI Tools) and a 278M multilingual embedding model open-sourced on GitHub (Toolify).
A permissive licence, a hybrid MoE family, signed weights, and documented data. That is the release profile regulated industries and the public sector actually ask for, and it is a different story than the weights-only headline labs.
SmolLM3, small on purpose
The other documented blind spot is small on purpose. Hugging Face's Smol Models team shipped SmolLM3, a compact open-weight 3B model built around efficiency, multilingual reach, and long-context reasoning (Pure Neo). It carries a 128,000-token context window and dual reasoning modes, and it is reported to outperform larger models at that footprint (CTOL, ScaleByTech). Where it runs is the point: local execution (ML Hive) and personal devices (Techzine), not someone else's cloud. Its predecessor SmolLM2 was trained on 11 trillion tokens, including custom math, code, and instruction datasets (Analytics India Magazine).
What makes it fully open is the training set. SmolLM-Corpus is itself published, combining Cosmopedia v2 (quality synthetic educational content) and FineWeb-Edu among its three main components (Noze), and other locally-runnable models are now being trained on it too (The Unwind AI).
For everything else, 2026's story is already clear: open weights stopped being the alternative and became the market.
Top comments (0)