DEV Community

Damian Dixon
Damian Dixon

Posted on

Fists up: 9 live B200 and A100 jobs, 4 suppliers, 3 hits taken, every one refunded.

Today we ran live B200 and A100 jobs through Kilawatt Cloud's agent-exec API, the same path an AI agent uses to provision compute. Nothing simulated. Every job launched a real machine at a real supplier, and every machine was terminated and checked against the supplier's own records afterward.


Image 1

The results: 6 of 9 launched jobs reached healthy

  • B200: both runs passed. RunPod was healthy in 0.5 minutes and Vast.ai in 1.6.
  • A100 (1 card): RunPod 0.2 min, Lambda 2.9, Hyperstack 3.5. Vast.ai never became healthy.
  • A100 (2 cards): RunPod passed in 1.0 min. Vast.ai failed twice.
  • Skipped: a 2-card A100 on Hyperstack, which our spend guard blocked because Hyperstack prices two cards as a whole machine.

Image 2

What the run showed

  1. B200 capacity is real and reachable through an API. Two live B200 jobs passed health checks on different suppliers.
  2. Supplier speed varies a lot. The same A100 took anywhere from 0.2 to 3.5 minutes to become healthy depending on who supplied it. That gap is what a single-supplier setup makes you absorb.
  3. We never substitute a card. Only RunPod and Vast.ai had B200 in stock. Lambda and Hyperstack reported no stock, so requests went only to suppliers that could deliver the exact card asked for.
  4. Failed machines cost the customer nothing. All three Vast.ai machines that never finished booting were refunded in full, automatically.
  5. Cleanup works. Every machine was confirmed gone against the supplier's own records.

Image 3

The parts that aren't flattering

We're publishing these because a verification report that hides its problems isn't worth much.

  • Vast.ai boot reliability is still an open problem. Its A100 machines failed to boot on three launches, and it has failed on several runs across the day. Customers are always refunded, but we haven't solved the reliability itself yet.
  • Shutdown timing. During the run, one RunPod A100 finished its booking but wasn't shut down until the next 15-minute check. Customers were only billed for the booked time. We've since moved to a 1-minute shutdown check, and a live test afterward stopped a machine 6 seconds after its booking ended.
  • An extra Vast.ai launch. A supplier-pinning step didn't take effect on the first try, so one RunPod job landed on Vast.ai. It was a valid launch with no card substituted, and it was refunded. We've disclosed it in the report.
  • Vast.ai's marketplace moves fast. Its B200 price changed between quote time and launch. The report shows what customers were actually billed.

How we ran it: each job was booked for 10 minutes using a temporary API key on an internal test account, sent to the production endpoint, with the GPU pinned to an exact card. Other suppliers were switched off around each launch to test one at a time, then switched back on. All temporary keys were revoked afterward.

We'll keep running these and publishing the results, including the failures. Every verification report, with charts and raw numbers, is here:


Image 4

👉 https://www.kilawattcloud.dev/super-intelligence

If you're building agents that need to provision GPUs on their own, what would you want to see in a report like this?

Top comments (0)