Ran real provisioning tests against our own infra. Verified results:
• RunPod: hard caps every pod at 8 GPUs (confirmed via direct API — 9 and 16-GPU requests rejected)
• Vast.ai: real on-demand capacity up to 16 GPUs (workstation-class: A4000/A5000/A6000/3090/4090/A40) and up to 4 GPUs (B200), confirmed with live rentals
• Ran 3 simultaneous 9-12 GPU deployments across Japan, Romania, and Canada — all held concurrently; one request failed cleanly when it collided with another over the only 12-GPU machine on the marketplace (no charge)
We're kilawattcloud.dev — multi-provider GPU orchestration.
Happy to answer questions on methodology in the comments.
Curious what other GPU brokers' actual concurrency limits look like under real load — most don't publish this.


Top comments (0)