I've been staring at Meta's Q2 numbers all week, and I still can't quite shake the feeling that we're watching something real happen — the kind of thing people will look back on and say "yeah, that was the moment."
Let me walk through what actually caught my eye.
Meta burned through nearly all its cash in one quarter.
Not figuratively. Not "spent a lot." The company posted $60.8 billion in revenue — up 28% year-over-year, beat expectations — and nobody cared. Because on the other side of that ledger, capital expenditures hit $31.1 billion in a single quarter. Free cash flow collapsed from $10.9 billion (Q2 2024) to $784 million. That's a 91% drop. Shares fell 10% in after-hours trading.
The part that bothers me is not the spending itself. Every hyperscaler is spending like this — Google's Alphabet just reported its first-ever negative free cash flow of $5.9 billion. But Google has Cloud ($24.8B, growing 82%). Amazon has AWS. Microsoft has Azure. Meta has ads, and that's it.
Zuckerberg acknowledged this on the earnings call, saying they're getting "a lot of offers for compute at a significant premium." Which is basically a polite way of saying "we're sitting on a goldmine of compute but haven't figured out how to sell it yet." The Meta Compute initiative is still in its infancy. Reality Labs lost another $4.62 billion. And the company just announced a $14 billion data center venture with BlackRock in El Paso — where BlackRock owns 80% and Meta leases the capacity back.
That's not a growth story. That's a cash flow management story. And I'm not sure the market is buying it.
ByteDance's Seed LLM is going into cars now.
This one flew under the radar but I think it matters. OMODA (the Chery-backed brand) just launched a "Super AI Cockpit" in Southeast Asia, and it's running ByteDance's Seed LLM under the hood. The debut model is the OMODA 4, unveiled at an event in Indonesia.
Here's what's interesting: they're running 10 AI agents inside the car — an Orchestrator agent for intent decomposition, a Vehicle Control Agent that handles up to 7 voice commands in a single utterance, and a whole ecosystem covering navigation, vehicle usage guidance, conversational companionship, and real-time info. The cloud-edge architecture means it works even when you lose signal, which honestly is where most in-car AI systems fall apart.
Three-stage OTA evolution is planned, so the car gets smarter over time rather than shipping frozen at launch. ByteDance's Seed model handles the natural language side — multi-round conversations, ambiguous intent, even generating AI wallpapers on the fly.
To be fair, this is still early. The real test is whether these 10 agents actually work together smoothly in practice, or if it ends up feeling like 10 separate apps glued into a dashboard. But the move itself says something about where LLMs are heading: off the screen and into the physical world, fast.
Want more VRAM for local LLMs? Someone will solder it for you.
A company called GPU Solutions is offering hardware VRAM upgrades for consumer graphics cards. They've demonstrated an RTX 2080 Ti going from 11GB to 22GB — enough to run decent-sized local models alongside AAA games.
The catch is that it's not a simple swap. You need micro-soldering, a custom VBIOS, and often memory strap resistor modifications on the PCB. Driver support can be spotty — earlier RTX 3070 16GB mods didn't work with newer drivers because the software simply didn't recognize the variant.
They're based in the UAE, charging AED 50 (~$13.60) for pickup and delivery, with 12 days turnaround and a 90-day warranty. Pricing for the actual upgrade work isn't listed, and with the current DRAM shortage, I'd expect a premium.
Still, the fact that this exists at all tells you something about the demand for local AI compute. People want to run LLMs at home without dropping $3,000 on an RTX 5090. And if soldering is what it takes, some folks will pay for it.
Two very different approaches to keeping AI agents honest.
Factify and Patronus AI are both selling LLM guardrails, but they're coming at the problem from opposite directions. Factify is all about document governance — certifying that the information an agent acts on is version-controlled, approved, and auditable. Think of it as compliance infrastructure for AI agents. Patronus AI, on the other hand, stress-tests agents by building simulated digital environments and measuring how they perform.
Which one matters more depends on what you're building. If you're deploying an agent that handles customer contracts or insurance claims, you probably want Factify's approach — you need to know exactly which policy version was in effect when the agent made a decision. If you're fine-tuning a model for open-ended tasks, Patronus's simulated testing might be more useful.
The truth is, both approaches will probably end up being necessary. Agentic AI is moving fast enough that the safety tooling is going to have to keep up, and right now it's still playing catch-up.
Also worth a quick mention: someone paired a local LLM with Obsidian on their phone and called it a productivity boost worth having. The whole stack runs offline. It's not fancy, but it works. I've been running llama.cpp on a MacBook for months now and honestly, the local LLM experience has gotten dramatically better just in the last year. If you haven't tried it since the early days, give it another shot.
That's it for this week. The Meta story is going to keep unraveling — that cash flow number is not something you walk back from in one quarter. ByteDance putting its LLM in cars is a signal worth watching. And the DIY GPU mod scene for AI is just getting started.
If you found any of this useful, pass it along. Always appreciate the conversation.
Built something practical? Check out PayCalc — a straightforward tool that does one thing well.

Top comments (0)