DEV Community

HIROKI II
HIROKI II

Posted on

AI Daily Digest — August 14, 2026: Claude Raises Riemann Bound to 67.2%, OpenAI Ultrafast 14x, DeepSeek V4-Pro GA

Cover

An unreleased Claude raises the Riemann zeta lower bound from 41.6% to 67.2%

Anthropic published a research note on August 13 describing what happened when a staff member told an unreleased research version of Claude to "take a real stab" at the Riemann hypothesis. It failed at the main conjecture, as expected for a problem open since 1859, but it did something unexpected: it raised the proven lower bound for the fraction of zeta-function zeros that lie on the critical line from 41.6% to 67.2%. That is a jump of 25.6 percentage points, against a prior pace where mathematicians took roughly five decades to go from about 33% to 41.6%, and about 37 years for the last 0.8 points. Anthropic's two in-house mathematicians reviewed the argument, external experts Brian Conrey and Dan Goldston examined the paper on short notice, and Claude produced a Lean formalization that passes standard tooling.

The way it got there is the part I find hard to look away from. Claude generated and discarded about 650 ideas first, all of which failed. Then it ran two Claude Code sessions totaling 31 million output tokens, coordinating roughly 60 subagents over a day and a half. Between them they executed about 2,400 shell commands, wrote hundreds of Python scripts, ran thousands of numerical checks against known zeta zeros, downloaded 54 arXiv papers to check for prior work, and refereed one another's reasoning. The actual math, according to Anthropic, combines recent work by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh (which lets Montgomery's techniques run without assuming the hypothesis) with a 2000 result by Bombieri, handled through a Weil quadratic form where on-line and off-line zeros sit in positive- and negative-definite subspaces. The technique is a new combination of existing tools, not a new tool.

I would not read this as "AI solved a piece of the Riemann hypothesis," and neither does Anthropic, which explicitly says the approach is unlikely to reach a full proof. The paper is not peer-reviewed. What it is: the first AI-produced result in a hard branch of mathematics that survived review by two experts in the field and passed formal verification, generated by a swarm of subagents instead of a single prompt. My honest reaction is mixed. As an agent-orchestration demo it is the most impressive thing I've seen this month. As mathematics, the interesting test is replication — whether the 67.2% holds up under independent scrutiny, and whether the same swarm pattern generalizes to problems where the result is not just a bound that can be checked by Lean.

— Anthropic · 新浪财经 (via 腾讯)
🔗 Anthropic (Learning more about Claude's mathematical capabilities) · 新浪财经 (黎曼猜想证明为何如此困难) · dig.watch

OpenAI previews Ultrafast: GPT-5.6 Sol at 14x speed, powered by Cerebras

OpenAI rolled out a preview mode called Ultrafast on August 13 that runs its flagship GPT-5.6 Sol up to 14x faster than standard processing, delivering up to 750 output tokens per second, according to a company blog post reported by TechCrunch. The framing matters as much as the number. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the company wrote. "Ultrafast points to progress in a new direction: more useful work per second." OpenAI lists incident response, customer service and support, financial market analysis, and e-commerce as target workflows, and 9to5Mac adds voice, developer agents, financial research, and security response. The company says its own developers have used it to analyze logs and traces during incidents and to compress research cycles that previously ran overnight into several iterations during the workday.

The interesting technical detail is the hardware. Ultrafast is powered by OpenAI's partnership with Cerebras, the wafer-scale chip company, which means the speed story is really an inference-infrastructure story. Latency has become the procurement gatekeeper in enterprise AI: for a contact-center script, a code-review pipeline, or a trader's workflow, a model that responds slowly might as well not respond. Anthropic has a fast mode for Claude, but TechCrunch reports it does not match the claimed speed here. The preview is limited to a small group of customers, with access expanding "as capacity grows" — which is the honest caveat. If Ultrafast depends on specialized Cerebras capacity, flipping it on for everyone is an infrastructure problem, not a software toggle.

I keep coming back to the strategic read. OpenAI is competing on tempo now, packaging inference speed as a product feature aimed at the enterprise seat where decisions get made. That pairs with the same-day news that its revenue chief, Denise Dresser, is leaving after less than a year, replaced by Wiz president Dali Rajic, two days after longtime executive Brad Lightcap departed. The product story and the org story point the same direction: OpenAI is pushing hard on velocity and distribution while the commercial bench churns, with an IPO and a well-funded Anthropic on the other side of the table. Ultrafast being 14x faster does not tell you whether the enterprise can actually buy it at scale — that is the number I will be watching.

— OpenAI (via TechCrunch) · 9to5Mac
🔗 TechCrunch (OpenAI introduces 'Ultrafast') · 9to5Mac · Barron's (via TradingView, CRO change)

Meta approved and ran AI-generated CSAM ads for nine months, WIRED reports

WIRED reported on August 5 that researchers at the Tech Transparency Project found more than 50 paid image and video ads containing AI-generated child sexual abuse material in Meta's public ad library, running across Facebook, Instagram, Messenger, and Threads between November 2025 and August 2026. Some ads reached several thousand accounts; one reached 2,563 accounts in Europe, targeted at users in the US, UK, and more than a dozen European countries. The ads were not user posts that slipped through after the fact — they went through Meta's ad review pipeline, which the company says is "primarily" automated, and were monetized. Several linked to AI nudify apps, including MaskAI, which Apple removed from the App Store after WIRED contacted the company. "These ads made no effort to mask the images or hide what they were promoting," TTP director Katie Paul told WIRED. "These are ads that were reviewed, approved, and allowed to run by Meta, never encountering interference while the company collected the ad dollars."

Meta removed the ads after WIRED reached out, and a spokesperson said "sexual exploitation is horrific" and that the company removed over 36 million pieces of child sexual exploitation content last year. But the response also conceded the detection failure: many of the ads predate "new AI technology we launched recently to better detect and block violating ads at upload." Researchers found about 30 more ads hours before publication, some published after WIRED first asked Meta about the original batch. This is the second paid-ad CSAM finding in weeks, after a July BBC investigation showed Instagram serving ads in India that directed users to Telegram channels selling illegal material.

I want to be careful with my own reaction here, because the details are genuinely vile and the stakes are not abstract. The technical point that matters for the AI story: generative tools have made CSAM production cheap and variable, and classifiers trained on past examples keep missing new variations, which is exactly the arms-race dynamic the industry warned about. The regulatory point: Spain has already directed prosecutors to investigate Meta, X, and TikTok over synthetic CSAM, and this is the second incident in weeks, which tends to move enforcement faster than outrage. What I cannot shake is the timeline — nine months of ads running while ad revenue was collected, removed only after an external watchdog and a reporter pushed. That is a process failure as much as a model failure, and no better ad classifier fixes the first one.

— WIRED (Tech Transparency Project) · MediaNama
🔗 WIRED (Meta Ran Ads That Contained AI-Generated CSAM Imagery) · MediaNama · IBTimes SG

OneDayAgent: a long-horizon harness from Zhejiang and Ant that sets a new agent record

Researchers from Zhejiang University's Zhang Ningyu group and Ant Group published OneDayAgent on arXiv (August 4), a harness built for open-ended, long-horizon agent requests — the kind that span work, study, and life in a single instruction: research a topic on the web, then edit a local deliverable, then produce a deck, while holding onto the original constraints the whole way. The paper names the three failure modes it targets: goal drift (the agent forgets early requirements as context accumulates), state loss (information gathered in one environment fails to transfer to the next), and context overflow. OneDayAgent turns the request into a managed execution process built on three capabilities: task decomposition into bounded subtasks, execution memory that compresses observations and checkpoints state under context pressure, and verification-and-repair that re-aligns the final deliverable with the original intent.

On AgentIF-OneDay, a benchmark of 104 real tasks, OneDayAgent with a GLM-5.2 backend scored 0.821, a new state of the art and the top result across all task types, domains, and rubric dimensions. The more interesting claim is generality: the same harness runs on five backend LLMs from three model families without any backend-specific tuning, even though different models produce different execution styles under the same workflow. The code, data, and trajectories are open-sourced. The news picked up in Chinese tech media on August 13, a day after another agent benchmark paper, which tells you how crowded this lane has become.

The framing — "harness" rather than "model" — is the same word DeepSeek used this week for its own agentic push, and I think that convergence is the actual story. The market has figured out that raw model capability is commoditizing; what differentiates agents now is the orchestration layer around the model. OneDayAgent's contribution is narrower than that: it shows a single harness design can manage multiple interacting failure modes at once, which prior work treated one at a time. The honest limits: 104 tasks is a small benchmark, GLM-5.2 is the backend that produced the headline score, and AgentIF-OneDay tasks, while "real," are still constructed. What I find genuinely notable is that the harness generalizes across model families without tuning — that is the property that makes a harness worth shipping, not a demo.

— arXiv · 腾讯新闻 (via TechWeb)
🔗 arXiv:2608.05013 (OneDayAgent) · TechWeb (浙大蚂蚁联手打造全能AI助理) · 腾讯新闻

DeepSeek ships V4-Pro-0813 with real agent gains and a peak/off-peak price switch

DeepSeek replaced its preview with the official V4-Pro-0813 on the API on the night of August 12, announced via its official WeChat account and API docs on August 13. The model name stays the same; the capability jump is not subtle. Agent-focused benchmarks released by DeepSeek: DeepSWE (software engineering) went from 12.8 to 62.7, DSBench-Hard (data science) from 31.1 to 67.2, Terminal-Bench 2.1 from 72.1 to 87.9 (against 88 for Anthropic's Fable 5), CyberGym from 52.7 to 83.3 (slightly above Fable 5's 83.1), and HLE-with-tools from 48.2 to 60.0. The context window is 1 million tokens, max output 384K, with thinking and non-thinking modes and three thinking strengths (low / high / max). The API now natively supports OpenAI's Responses API format and ships a one-click config script for Codex.

The pricing move is the other half of the story. DeepSeek announced peak/off-peak pricing effective August 17, 2026: during peak hours (9:00–12:00 and 14:00–18:00 Beijing time) V4-Pro lists at ¥9 per million input tokens on cache miss and ¥27 per million output; off-peak is half. Cached input is ¥0.3 per million at peak. SemiAnalysis congratulated DeepSeek and said the model "massively beats Nemotron 3 Ultra on agentic tasks"; 猎豹 CEO 傅盛 put the price-performance bluntly: performance near the top models at roughly one-fifty-seventh of Fable 5's price. xAI's Grok 4.6 landed the same day (separate story below), which made August 12–13 a genuinely crowded 48 hours for agentic models.

The question nobody has fully answered is whether the price war is ending, not continuing. Morgan Stanley's research this week noted Chinese API prices have been rising over the past year while US closed models keep cutting, so the gap is narrowing — DeepSeek's own announcement says "we will update API pricing" as the full V4 family goes GA, and the ¥27 peak output rate is well above V4-Flash's levels. I read the peak/off-peak mechanism as DeepSeek admitting its inference costs are no longer negligible: you do not build time-of-day pricing into a product whose marginal cost is zero. That is the tell that the era of absurdly cheap Chinese frontier inference is shading into an era of merely cheap frontier inference — still a big gap vs. US pricing, but a gap that is closing from both sides.

— DeepSeek (官方 API 文档) · 新京报 · 21世纪经济报道
🔗 DeepSeek 官方 API 文档 · 新京报 (DeepSeek-V4-Pro正式版上线) · Global Times (DeepSeek launches V4-Pro) · 21世纪经济报道 (DeepSeek重大更新)

Tencent's Q2: AI capex up 176% to ¥52.8B, free cash flow turns negative

Tencent reported Q2 2026 results on August 12: revenue ¥204.8 billion (+11% YoY), gross profit ¥118.4 billion (+13%), Non-IFRS operating profit ¥75.6 billion (+9%). The number that matters for the AI story is capex: ¥52.8 billion, up 176% year over year and 65% quarter over quarter. First-half capex of ¥84.7 billion already exceeds all of 2025's ¥79.2 billion. The company says it is "substantially increasing compute procurement" to support Hy model upgrades, WorkBuddy and CodeBuddy inference, WeChat AI initiatives, and cloud demand. That spending dragged free cash flow to negative ¥13.8 billion for the quarter (operating cash flow of ¥52.7 billion minus ¥59.3 billion in capex payments and other items). Tencent notes the operating cash flow includes large AI-related prepayments, and that excluding compute prepayments, FCF would have been +¥37.6 billion — the prepayment swing is roughly ¥51.4 billion.

The AI products are still in investment phase, and the company is explicit about it. Excluding new AI products (Hy, Yuanbao, CodeBuddy, WorkBuddy, Xiaowei), Non-IFRS operating profit would have been ¥86.1 billion, up 19%; including them, the AI line items took about ¥10.5 billion off operating profit in the quarter. Tencent says Hy3 has ranked top-three globally in token consumption on OpenRouter since launch, WorkBuddy is showing "rapid user growth" with healthy retention, and the WeChat agent Xiaowei is in expanded gray-scale testing. Management framed the quarter as building a "new AI-empowered Tencent" across intelligence, applications, and infrastructure layers.

The context that makes this a big story rather than a single company's earnings: Tencent is late but heavy in the hyperscaler capex race, and it is not alone. Alphabet's Q2 capex hit $44.9 billion, roughly doubling year over year, with free cash flow turning negative for the first time (-$5.9 billion); Alibaba's fiscal Q4 FCF also went negative. Three of the world's largest platforms are now betting that compute procurement converts into revenue faster than depreciation catches up with their income statements. I find Tencent's numbers the most honest read on that bet so far: ¥51.4 billion of prepayments, booked in one quarter, for hardware that will take years to monetize. The market's response was telling — Tencent shares fell over 4% on the report despite revenue beating. The bull case is that AI is already lifting ads and cloud; the bear case is that you cannot tell yet whether the ¥52.8 billion buys durable advantage or just a seat at the table.

— Tencent (财报) · 北京日报客户端 · 网易
🔗 Tencent Q2 2026 业绩公告 (PDF) · 北京日报 (腾讯二季度营收增11%) · 网易 (腾讯Q2资本开支527.8亿)

xAI releases Grok 4.6, matching GPT-5.6 Sol on the AA Intelligence Index

xAI shipped Grok 4.6 on August 12, an upgrade aimed squarely at long-running agents and "more ambitious interactive and visual work." The headline: it scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max and one point behind Claude Fable 5 Max's 62. On GDPVal-AA v2 it hits 1753 Elo, on CursorBench 3.2 it posts 69.9% (the highest of the four models compared), and DeepSWE v1.1 jumps from 54% to 65.9%. The training story is a longer supplemental run with curated model-generated reasoning data, Grok 4.5 regenerating the SFT trajectories, and agentic RL across domains including kernel optimization, web development, and CAD. xAI says it shipped its widest-ever pre-deployment test suite, with safeguards calibrated to the expanded capabilities.

Pricing is where the positioning shows. Grok 4.6 lists at $2/$6 per million tokens under 200K prompt tokens, with a long-context tier that doubles to $4/$12 once a prompt crosses 200K — and the higher rate applies to the whole request, a detail most comparisons skip. Priority Processing is a 2x lane on all token types, not a separate model. It is available via the xAI API, Cursor, and Grok Build (2x included usage for the first week), plus OpenRouter, Vercel, and Cloudflare, with a 500K context window and a February 1, 2026 knowledge cutoff. This is the second xAI model release in five days, after Grok Imagine Image 2.0, and the first from the family since Grok 4.5 hit Cursor on July 9.

What I take from the launch is less about Grok itself than the state of the agent-model market. The interesting numbers are the deltas: DeepSWE +11.9, APEX-Agents +10.4, Terminal-Bench +10.3 over Grok 4.5, while Terminal-Bench in absolute terms still sits at 26%, well behind GPT-5.6 Sol's 34.6% and Fable 5's 34.1%. So this is a real step up in agentic coding, and it is still not a sweep. Together with DeepSeek's V4-Pro-0813 landing the same night, the pattern is unmistakable: the competition has fully moved from "who answers better in chat" to "who completes long multi-step work at a price and latency the customer can live with." xAI's answer is parity-plus-price; the honest caveat is that its vendor-reported benchmarks come with the usual "trust us" discount until independent evals confirm them.

— xAI (官方博客) · LLM Stats
🔗 xAI (Introducing Grok 4.6) · LLM Stats (Grok 4.6 release, benchmarks and agent loops) · developersdigest (Grok 4.6 release guide)

Top comments (0)