AMD takes direct aim at Nvidia
What's going on?
AMD unveiled Helios - a rack packing 72 MI455X GPUs and 31TB of memory with Microsoft, Meta, OpenAI, and Oracle already signed on as customers.
Why it matters?
Nvidia still owns ~95% of the data centre GPU market. A credible rival with big-name backing means real competition, and eventually better pricing for anyone running AI at scale.
What we think at CometChat
More memory per rack sounds like plumbing until you're the one serving a million-token context to real users. The infrastructure race is quietly deciding what AI features you can actually afford to ship. We're watching closely.
Nvidia lands $500B in Asia AI deals
What's going on?
Jensen Huang spent two weeks in Japan and South Korea, striking partnerships to wire AI into factories, robots, and a 2-gigawatt data centre with SK Group.
Why it matters?
This is AI leaving the chatbox and moving into the physical world - factories, machines, sovereign infrastructure. The compute buildout underneath everything just got a lot bigger.
What we think at CometChat
Robots and factories still need to talk - to each other, to operators, to the humans in the loop. As physical AI scales, the coordination layer stops being a nice-to-have. That's the boring part that decides whether any of it works.
Alibaba open-sources its biggest model yet
What's going on?
Alibaba dropped Qwen3.8-Max - 2.4 trillion parameters, competitive with the best from OpenAI and Anthropic, and open-sourced with weights coming next week.
Why it matters?
A frontier-class model you can download and run yourself changes the math. Less vendor lock-in, more control, and a serious option for agentic coding and research.
What we think at CometChat
Every capable open model raises the same quiet question: who's operating this in production? The model is the easy part. Wiring it into real conversations, agents, and users is where the actual work starts and where it usually breaks.
Qualcomm buys Modular to take on CUDA
What's going on?
Qualcomm closed its $3.9B acquisition of Modular Chris Lattner's AI software company betting on a vendor-neutral layer that runs AI across any chip, not just NVIDIA's.
Why it matters?
CUDA lock-in has kept teams tied to one vendor for years. A real cross-silicon software layer means more choice on where and how cheaply you run AI.
What we think at CometChat
The interesting fights are moving up the stack, from hardware to the software that makes it usable. Same story with chat and agents: the chip matters less than the layer that turns it into something people actually talk to.
AMD undercuts Nvidia on inference cost
What's going on?
New benchmarks show AMD's MI355X running the 2.8-trillion-parameter Kimi K3 model at under half the per-token cost of Nvidia's B300.
Why it matters?
Nvidia still wins on raw throughput, but AMD wins on tokens per dollar. For anyone running big models in production, that math adds up fast.
What we think at CometChat
Inference cost is the boring part nobody plans for right up until your AI features hit real usage. When chat and agents scale, the bill scales with them. Cheaper tokens mean more room to actually ship the good stuff.
Top comments (0)