๐ค๐ป AI Daily Digest โ July 20, 2026
OpenAI GPT-5.6 Series Goes Public with Three Tiered Models
OpenAI made the GPT-5.6 series publicly available on July 9, rolling out three distinct models โ Sol, Terra, and Luna โ across ChatGPT, Codex, and the OpenAI API. The flagship Sol introduces two new capability tiers: "max" allocates additional inference time for exploring alternative solutions and self-correcting approaches, while "ultra" coordinates four parallel agent instances to tackle complex multi-step tasks with higher token consumption. Terra is positioned as the balanced daily-work model, and Luna as the fastest, most cost-efficient option. The series ships with what OpenAI calls its most robust safety deployment yet, including extensive evaluations in biological and cybersecurity domains documented in the accompanying system card.
Separately, OpenAI introduced GPT-Live on July 8, a new generation of voice models built on a native real-time architecture. The model supports natural interruptible conversation, pause comprehension, real-time translation, and dictation, while seamlessly orchestrating backend models like GPT-5.5 for complex reasoning and web search. Over 150 million people now use ChatGPT Voice and related speech features weekly, according to OpenAI. โ OpenAI Newsroom ยท Xinhua
๐ OpenAI GPT-5.6 Announcement ยท GPT-Live Announcement ยท Xinhua Coverage
Anthropic Claude Fable 5 Returns as Sonnet 5 Targets Agentic Coding
Anthropic restored access to Claude Fable 5 on July 1 after US export controls imposed on June 12 were lifted. The model and its sibling Mythos 5 were restricted when the government became aware of a report in which Amazon researchers found a method of bypassing Fable 5's safeguards. Anthropic worked with the US government to train an improved safety classifier that blocks the reported technique in over 99% of cases. The company also released Claude Sonnet 5 on June 30, a more powerful mid-size model designed for agentic tasks. Sonnet 5 can make plans, use tools like browsers and terminals, and run autonomously at a level that previously required larger models. It is now the default model for Free and Pro Claude users, priced at $2 per million input tokens through August 31. โ Anthropic Newsroom
๐ Fable 5 Redeployment ยท Claude Sonnet 5
Meta Muse Spark 1.1 Enters AI Coding Wars, Llama API Shuts Down
Meta released Muse Spark 1.1 on July 12, its most powerful agent model now specifically targeting agentic coding. CEO Mark Zuckerberg emerged from a three-year social media hiatus to personally announce the model on X. According to Meta AI head Alexandr Wang, Muse Spark 1.1 is "currently the strongest model in agentic tasks and coding." The model excels at multi-application computer use workflows, navigating unfamiliar interfaces with minimal human intervention and maintaining context across long sessions.
In a related strategic shift, Meta officially shut down the Llama API service on July 6, ending its 14-month experiment in selling API access. The company is pivoting to a dual-track strategy: open-source Llama continues for the community while closed-source Muse powers Meta's proprietary ecosystem across WhatsApp, Instagram, and Facebook. Meta also confirmed it is training a larger model codenamed "Watermelon" that has reportedly matched GPT-5.5 on key benchmarks. โ Meta AI ยท The Verge
๐ Muse Spark 1.1 Announcement ยท Llama API Shutdown
Microsoft Quietly Replaces OpenAI Models with In-House MAI in Excel and Outlook
Microsoft has begun replacing third-party AI models from OpenAI and Anthropic with its self-developed MAI model family within Excel and Outlook. The new MAI-Thinking 1 model has demonstrated performance matching Claude Opus 4.8 on coding benchmarks, according to internal testing cited by Bloomberg. Tens of thousands of AI prompt requests in these two flagship Office applications are now handled entirely by Microsoft's own models each week. The shift is driven by cost pressures and data residency requirements. Mustafa Suleyman, Microsoft's AI CEO, is leading the effort to reduce dependency on premium third-party APIs as OpenAI's discounted partnership window narrows. While MAI's overall share remains modest, the Excel and Outlook deployments represent a beachhead that could expand across Copilot, Azure AI, and the entire Microsoft 365 ecosystem. โ Bloomberg ยท Microsoft
๐ Bloomberg via 163.com ยท Microsoft AI Blog
Mistral Leanstral 1.5 Achieves Perfect Formal Verification Score
Mistral AI released Leanstral 1.5 on July 2, a 119B-parameter Mixture-of-Experts model (6B active per token) specialized for Lean 4 formal verification. Released under Apache 2.0, the model saturates miniF2F at 100% on both validation and test sets, solves 587 out of 672 PutnamBench problems, and achieves new state-of-the-art results on FATE-H (87%) and FATE-X (34%) graduate algebra benchmarks โ all at an estimated $4 per solved problem, compared to $300+ for competing high-budget provers. The model's test-time scaling behavior is remarkable: performance on PutnamBench rises monotonically from 44 problems at 50K tokens to 587 at 4M tokens per attempt. Beyond pure math, Leanstral identified 11 genuine bugs across 57 open-source Rust repositories, including an integer overflow in a varint decoding library that could cause silent data corruption. โ Mistral AI ยท HuggingFace
๐ Mistral Leanstral 1.5 ยท HuggingFace Model
NVIDIA and HuggingFace Launch Open Data for Agents Initiative
NVIDIA and HuggingFace jointly announced the Open Data for Agents initiative, publishing over 10 trillion pre-training tokens and millions of post-training samples specifically designed for building AI agents. The release includes region-specific synthetic personas and an interactive Nemotron Post-Training v3 Prompt Atlas, enabling organizations to fine-tune agent models without exposing proprietary data. In a parallel move, the partners also integrated NVIDIA's Isaac GR00T 1.7 reasoning vision-language-action model and the Isaac Teleop data-collection framework into HuggingFace's LeRobot ecosystem, dramatically reducing the barrier to entry for robotics AI training and deployment. The combined effect standardizes post-training, evaluation, and deployment for humanoid robotics inside a widely used open-source stack. Cosmos 3 integration is planned next. โ NVIDIA Blog ยท HuggingFace
๐ NVIDIA Blog ยท HuggingFace Blog
LLMs as a Jury: Cross-Model Consensus Outperforms Trained Reward Models
A new paper on arXiv (2607.10139) proposes a surprisingly simple yet effective method for selecting correct reasoning chains: instead of training reward models or relying on single-model self-consistency, use the agreement signal across a panel of independently trained models. Dubbed "LLMs as a Jury," the approach treats the panel's consensus pattern โ not any model's score of another โ as the verification signal. Across seven benchmarks, this cross-model consensus selects correct answers better than self-consistency and far better than a model scoring its own candidates. On competition math, it closes the entire gap to an oracle selector. The mechanism is error decorrelation: independently trained models err differently, so their wrong answers scatter while the correct one accumulates agreement. The authors derive a parameter-free law that predicts consensus accuracy from three panel statistics to a mean absolute error of 0.03. Against four trained verifiers, the free LLM-jury matches the strongest inside their math training domain and is the top selector outside it. โ arXiv ยท Ning Liu
๐ arXiv:2607.10139
Next digest arrives tomorrow. Follow KD Agentic for daily AI intelligence.
Top comments (0)