DEV Community

AI Maker
AI Maker

Posted on

AI Roundup (Sat Aug 15): DeepSeek Goes GA, Grok 4.6 Ships Agent-First, Google Halves Gemini Flash Pricing

DeepSeek V4 Pro 0813 goes GA — open weights, 1M context, cheap

After nearly four months in preview, DeepSeek shipped the general-availability build of V4 Pro 0813 (Aug 13). It's a 1.65-trillion-parameter mixture-of-experts model released under the MIT license, with weights on Hugging Face, a 1M-token context window, and up to 384K output tokens.

  • Price is the story: ~$0.435 / 1M input and $0.87 / 1M output tokens via DeepSeek and OpenRouter — roughly an order of magnitude below closed frontier APIs.
  • Benchmarks: Artificial Analysis Intelligence Index 53; #2 on SWE-bench Verified (96.40%), the highest-scoring open-weight model on that board, ahead of Kimi K3. Strong on GPQA Diamond (93%) and agentic tool use (Toolathlon-Verified 74.1).
  • Availability: live on the DeepSeek app, API, and OpenRouter, with a lighter V4-Flash sibling shipped alongside it.

Read: it puts "open weights + strong + cheap" back together — the exact combination squeezing Western frontier labs on cost.

Grok 4.6 — xAI's agent-first flagship

xAI (now trading as SpaceXAI after the SpaceX acquisition) released Grok 4.6 on Aug 12 — not a bigger model, but a post-training upgrade on the same 1.5T base as Grok 4.5.

  • Artificial Analysis Intelligence Index 61, tying GPT-5.6 Sol and one point behind Claude Fable 5 — among the top handful of frontier models.
  • Built for long-running agents: better self-verification across many steps, a new "xhigh" reasoning setting, and supplemental training on anonymized Cursor workflow data.
  • Pricing flat at $2 / $6 per 1M tokens (a 2x fast variant is available). Shipped same-day inside Cursor — which SpaceXAI acquired — and Grok Build.

The signal: current foundations still have post-training headroom, and xAI is spending it on multi-step reliability rather than scale.

Gemini 3.7 Flash — Google halves the workhorse price

Google released Gemini 3.7 Flash on Aug 13, just three weeks after 3.6 Flash, doubling down on coding and agent workflows.

  • Big coding jumps: DeepSWE v1.1 65.3% (up from 49.0%), FrontierCode 1.1 Main 43.6% (up from 34.4%), AutomationBench 30.4% (up from 17.0%).
  • Half price: introductory $0.75 / $3.75 per 1M tokens through Dec 31, 2026 (then doubles). Same 1M context window, 65K output.
  • Fastest in class: Artificial Analysis ranks it #1 of 186 models on output speed at ~340 tok/s. API and enterprise only — no open weights.

The cadence is the moat: a new Flash every three weeks, each cheaper and better, making "latest model" a moving target.


Daily AI briefing at AI Nexus Daily.

Top comments (0)