Salesforce posted its best quarter in years — and Claudeforce is the counterpunch to the SaaSpocalypse
Salesforce reported Q2 FY27 (ended July 31) on August 26 with revenue of $11.35 billion, up 11% year over year, adjusted EPS of $5.90 beating consensus by $2.63, and free cash flow of $1.1 billion, up 81%. cRPO grew 14% in constant currency to $33.5 billion, and management raised full-year revenue guidance to $46.1–46.4 billion. The AI numbers are the story: Agentforce ARR passed $1.5 billion, up over 240% year over year, and combined Agentforce plus Data 360 ARR reached nearly $3.9 billion, up over 210%. Customers generated 3.2 billion Agentic Work Units in the quarter, up 97% quarter over quarter (7 billion cumulative), and accounts with agents in production grew 70% sequentially. Slack posted its fastest quarterly net-new-annual-order-value growth since acquisition, with Slackbot passing 1 million active users, up 150% quarter over quarter. Data 360 ingested 104 trillion records in the quarter, up 355%.
The quarter is a direct counterargument to the February selloff that erased roughly $285 billion from SaaS valuations after Claude Cowork launched; Salesforce fell 26% in that stretch. Benioff called the fear "the SaaSpocalypse" and pushed back on the call: "This is not the SaaSpocalypse," arguing customers deploy AI through software platforms rather than replacing them. The day of the earnings, Salesforce and Anthropic announced Claudeforce, making Claude the default model across Slack AI, Slackbot, Agentforce Coworker, and Salesforce's internal engineering tools, with a "Salesforce in Claude" plugin carrying 37 prebuilt sales skills, GA in September. My read: the record is real, but monetization still leans on premium-edition upgrades — only about 5% of sales/service knowledge workers have upgraded, and roughly half of Agentforce bookings came from customers replenishing usage credits. The question worth watching is whether Agentforce ARR keeps compounding from real agent workloads, or flattens once the easy per-seat-to-premium migration is done. The $2.6 billion investment gain, tied partly to Salesforce's Anthropic stake, also flatters the bottom line, which is worth remembering when comparing EPS.
— Salesforce (official) · Nasdaq · SaaS Sentinel
🔗 Salesforce Q2 FY27 press release (SEC exhibit) · Nasdaq earnings call highlights · SaaS Sentinel on the record quarter
Anthropic locked in $45 billion of compute from Nscale — six years, 460 MW, Vera Rubin
Bloomberg reported Wednesday that Anthropic agreed to spend $45 billion over six years renting AI compute capacity at Nscale's data center campus in West Virginia, about 460 megawatts, with NVIDIA Vera Rubin systems expected online late next year. At $7.5 billion a year on average, this is a utility-style commitment, not a cloud contract: pure capacity rental, no equity and no ownership of the facility. Nscale is a two-year-old British neocloud that is itself preparing for an IPO (reports this week suggested it hopes to raise around $3 billion), and it already signed Microsoft for 1.35 GW at the same Monarch Compute Campus, also on Vera Rubin NVL72 hardware. Anthropic declined to comment; the deal follows a string of capacity moves — renting the full compute of SpaceX's Colossus 1 (220,000+ NVIDIA processors), a $10 billion deal with Volta in Norway, $5 billion with AMD, and more than 10 GW committed from cloud providers including a $200 billion agreement with Google.
The context is the IPO. Anthropic is preparing for a listing that will need to justify a reported valuation around $965 billion, and projected 2028 revenue of roughly $190–200 billion against a current run rate around $47 billion. That math only works if inference demand, especially from Claude Code and agentic workloads, keeps compounding. The risk on the other side is hardware generation risk: the capacity comes online on Vera Rubin late 2027, and if that architecture slips or the 2028 generation shifts, a six-year lockup at these numbers gets awkward. Read it together with the NVIDIA guarantee story below and the picture is a web of interlocking obligations: NVIDIA underwrites the campus, Anthropic rents it, and everyone's capex feeds back into NVIDIA revenue. What I keep turning over is whether six-year compute commitments behave more like utilities or more like the long-term chip agreements that burned hardware buyers in past downturns.
— Bloomberg (via Data Center Dynamics) · Inside AI · 财联社/经济参考报
🔗 Data Center Dynamics: Anthropic signs $45bn Nscale agreement · Inside AI on the capacity deal · 经济参考报 on the deal details
NVIDIA put its balance sheet behind OpenAI's Ohio campus — up to $105 billion in guarantees
NVIDIA's Q2 FY27 CFO commentary, filed August 26, formalizes what was announced August 17: NVIDIA has entered guarantees covering land, power, and shell buildout for about 4.25 GW at SB Energy's PORTS-Pike campus in Ohio, which will exclusively host NVIDIA infrastructure under 20-year leases to OpenAI. The guarantee cap is $105 billion, obligations phase in as data centers become ready for service (first expected in fiscal 2029), and exposure declines as OpenAI fulfills lease payments. NVIDIA also invested $1.5 billion in SB Energy and retains an option to support roughly 3.8 additional GW as the site scales. The filing notes each generation of NVIDIA infrastructure at the site could represent about 1.5 million GPUs and $150–200 billion in NVIDIA revenue. Total guarantee exposure on the books: $108.5 billion, including a $3.5 billion pool for other AI cloud partners.
This is the second act of a story Huang framed in March, when he said NVIDIA's $30 billion OpenAI investment "might be the last time" it buys shares because OpenAI is going public — but the credit window stays open. CFO Colette Kress preempted the criticism on the call: "We recognize the scale of this support, and we know some will call this circular financing. We see it differently," describing it as low-risk, high-reward. Technically the structure is a residual-value guarantee: NVIDIA pays only if OpenAI defaults, and OpenAI must reimburse every dollar NVIDIA actually disburses; the guarantee expires early if OpenAI reaches a satisfactory credit rating. The market's read is more mixed — when a $250 billion figure surfaced in July, NVIDIA's credit default swap spread widened from 0.40% to 0.82%, and one analyst argued the final $105 billion number "read as less demand and not less risk." NVIDIA reported $96.2 billion in revenue for the quarter, $89 billion of it data centers, with guidance of $108 billion next quarter and about 70% growth expected next fiscal year. My read: NVIDIA's moat is quietly migrating from silicon to capital — it now underwrites the demand for its own chips — and the question is whether that leverage cuts the other way in a downturn, when the guarantee and the revenue both compress at once.
— NVIDIA SEC filing (official) · Nasdaq · Certified Strategic
🔗 NVIDIA Q2 FY27 CFO commentary (SEC) · Nasdaq on the $105B guarantee structure · Certified Strategic on the circular-financing debate
Z.ai opened GLM-5.3's weights — every gain comes from post-training, and the cyber capability doubled
Z.ai released GLM-5.3 as open weights at midnight JST on August 29 (August 28 in the US), two weeks later than originally promised because the model's cybersecurity capabilities went through an extra safety review. It is a 744B-parameter mixture-of-experts model with 40B active per token, a 1M context window and 128K max output, and it shares the same 743B base checkpoint as GLM-5.2 — the claimed gains come entirely from post-training. On Z.ai's reported numbers: Terminal-Bench 3.0 jumped from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and the in-house Z.ai Code Bench improved about 50% while generating fewer output tokens. The cyber numbers are the sharpest: CyberGym 84.5 (up from 77.2, open-source SOTA) and ExploitBench 54.4, more than double GLM-5.2's 24.4, on a benchmark where Fable 5 scores 78 and GPT-5.6 Sol 76.5. Z.ai says work with security teams produced 2,436 vulnerability findings across 269 open-source projects, 1,097 of them critical or high severity.
The unusual claim is that tuning for agentic coding produced the vulnerability-exploitation jump as a side effect rather than an explicit optimization target. The licensing is also worth reading carefully: the GLM-5.3 license lets anyone copy, modify, distribute, sell, and fine-tune the weights, with one carve-out — organizations with more than $10 billion in annual revenue offering the model as an external service must pass a security review. That is a governance mechanism aimed at hyperscale providers for a dual-use model, and it is a different posture from the MIT-licensed GLM-5.3-Flash (320B/18B, released August 26), which was identified as the mystery "Ox Alpha" model that showed up on OpenRouter with a 1M-token window and no named developer. Serving GLM-5.3 is data-center territory — 756 GB of shards, FP8 tensors, vLLM/SGLang/OpenRouter day-0 support (vLLM measured 537.6 tokens/s/user on NVFP4). Artificial Analysis puts its Intelligence Index at 60 and Agentic Index at 59, above Claude Opus 4.8 on the first. My read: the post-training-only gain is the claim to scrutinize — if a 743B base can be lifted this much by RL and verification pipelines, then the race is increasingly about data and post-training infrastructure, not just base-model scale, and Chinese open labs just set the price for that argument.
— Z.ai (official) · GIGAZINE · AlphaSignal
🔗 Z.ai GLM-5.3 weights on Hugging Face · GIGAZINE on the open-weight release · AlphaSignal on the post-training gains
DeepMind ran the first double-blind evaluation of a frontier model — neither side could peek
Google DeepMind announced August 27 the world's first double-blind evaluation of a proprietary frontier-class model, testing a Gemini Flash Lite against confidential benchmarks held by outside evaluators. The partners — Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons — kept their test prompts sealed, and Google kept the model weights sealed. The mechanism: the model and the prompts meet inside Confidential Space, part of Google Cloud's Confidential Computing portfolio, running on an NVIDIA H100 80 GB Confidential GPU with Intel TDX host-memory encryption. Both parties run remote attestation before sending anything in, and OpenMined's PySyft lets each side review and approve the code running inside the enclave, with sensitive evaluation portions blocked from external network calls. Neither party can extract the other's data; the reported performance overhead is under 5%.
This attacks benchmark contamination, the structural problem where a model that has seen the test questions measures memorization rather than capability — one cited analysis found signs of leakage in roughly half of 31 models examined. The old trade was weights for prompts: either the evaluator handed over its test questions (risking the provider seeing them) or the provider handed over the weights (risking its IP). Double-blind evaluation replaces that trade with cryptography, and the July 2026 Singapore Consensus on Global AI Safety Research Priorities, spanning 13 countries and 100 contributors, explicitly called out the absence of this kind of infrastructure. My read: the score of this pilot matters less than the plumbing. MLCommons running MLPerf makes this look like a standards push rather than a one-off demo, and national AISIs are the natural first customers — they want to grade the most capable systems without trusting the grader or the student. The honest limits: double-blind tells you the score was clean, not what the score means, and one small model is a pilot, not a regime. The open question is whether the other frontier labs sign up to be tested the same way.
— Google DeepMind (official) · Tech Times · AI Chat Daily
🔗 DeepMind: Piloting the world's first double-blind AI evaluations · Tech Times on the cryptographic enclave · AI Chat Daily on the standards angle
Waymo published 10 AI lessons from 200 million driverless miles — and cameras alone aren't enough
Waymo's vice president of onboard software, Srikanth Thirumalai, published August 26 on the Waypoint blog ten lessons drawn from more than 200 million fully autonomous miles, the company's answer to the two most argued questions in autonomous driving. First: multimodal sensors are indispensable. The latest Ojai van carries 13 cameras, four lidar units, six radars, and microphones, and Waymo argues the 200 million miles show cameras alone cannot deliver safe Level 4 operation — lidar builds 3D geometry to millimeter precision, cameras read semantics like signs and light colors, radar tracks velocity and sees through rain, fog, and dust. Second: HD maps are a powerful "prior," treated like memory rather than a live input, so onboard compute focuses on what is new — a temporary stop sign, an unplanned detour — while an AI-driven mapping system keeps the reference layer updated. The other lessons consolidate the architecture: fewer, larger foundation models with teacher-student pairs instead of a "modular spaghetti" of single-task detectors, and an independent onboard validation layer that checks every proposed trajectory against physics constraints and traffic laws, because pure end-to-end architectures that map raw pixels to steering "run the risk of black box failures." Waymo runs tens of billions of simulated miles against its real fleet, and a "Critic" system continuously reviews driving behavior for engineers.
The post is aimed directly at the camera-only approach, and the timing is deliberate: Tesla is preparing to deploy its purpose-built Cybercab more widely, and a proposed New Jersey framework would require multiple sensors on robotaxis. Waymo also made the point that supervised driver-assist improvement is a different problem from full autonomy — "a false summit," in his words. It pairs with Waymo's custom silicon: a 5nm ASIC announced this week that converts raw sensor data and delivers more than 1,000 trillion operations per second, now powering the Ojai robotaxi in Phoenix, Los Angeles, and San Francisco. The counterargument remains cost and scale: Tesla says vision plus massive data scales cheaper, and 13 cameras plus four lidars is expensive hardware to ship at fleet volume. My read: Waymo is betting that safety-critical machines need redundancy and a validation layer you can audit, and the 200-million-mile dataset is the evidence it will keep putting in front of regulators. The question is whether that architecture wins commercially outside the US cities where it already operates, starting with Munich by end-2027.
— Waymo (official) · IoT Tech News · EVMagz
🔗 Waymo: 10 AI Lessons from Driving 200+ Million Fully Autonomous Miles · IoT Tech News on the architecture · EVMagz on the multisensor defense
Fireworks says DeepSeek V4 Pro undercuts Fable 5 by 3x per solved task — and takes the security work closed models refuse
Fireworks AI published August 26 a benchmark post arguing its hosted DeepSeek V4 Pro (0813) beats Anthropic's Claude Fable 5 on SWE-Bench and LiveCodeBench while costing about a third per solved task. The sharper claim is behavioral: across 840 traced adversarial security runs, the model recorded zero refusals and zero output-length truncations, meaning it takes legitimate security-review work that closed models decline. Fireworks' framing on X was blunt: "Your closed model refused a task it should have done. Ours didn't." The model itself, released by DeepSeek on August 13 as the official V4 Pro (superseding the preview, with a DSpark speculative-decoding module), is served on Fireworks at $1.32 per million input tokens and $3.96 output, with a 1M context window and 131K max output; Fireworks also offers SFT, DPO, and RFT training on top of it.
The angle here is security-agent economics, which is timely given last week's OpenAI/Hugging Face incident and the industry-wide tightening of agent boundaries. Fireworks is deliberately selling cost-per-solved-task rather than raw benchmark score, and it argues the model's wins land on tasks Fable 5 misses, making V4 Pro a better routing partner even where K3 scores higher on CyberGym. The caveats are real: these are Fireworks' own measurements, "zero refusals" measures willingness, not competence, and a closed model's refusal is often policy rather than capability. The deeper question the post raises is whether an enterprise that wants a security agent that never refuses has thought through what it is asking for — the same capability that finds vulnerabilities in open-source code is the capability that chains them into exploits, and the GLM-5.3 story above shows the same pattern emerging on the open side.
— Fireworks AI (official) · PulseAugur · Requesty
🔗 Fireworks: DeepSeek V4 Pro tops SWE-Bench, cuts cost per task by 3x · PulseAugur on the comparison · Requesty pricing/benchmarks for deepseek-v4-pro-0813

Top comments (0)