Short answer (October 10, 2026): spot GPUs are the same hardware as on-demand GPUs at a deep discount (AWS and Azure advertise up to 90% off, GCP up to 91%), but the cloud can take the machine back with 30 seconds to 2 minutes of notice, and anything not saved somewhere durable is gone. Spot is safe when your job checkpoints to storage that outlives the machine, saves again when the notice arrives, resumes on start and can fall back when spot capacity runs out; without that, a few reclaims can make spot cost more than on-demand.
Below: what fails on a reclaim, how much warning each platform gives, when spot really wins, and a copy-paste setup.
Spot vs on-demand at a glance
| On-demand | Spot (interruptible) | |
|---|---|---|
| Hourly price | List price | Discounted, often 60% to 90% below list |
| Runs until | You stop it | You stop it, or the cloud needs the capacity back |
| Warning before a reclaim | None needed | 30 seconds to 2 minutes |
| Capacity when you ask | Usually there | Can be empty for hours |
| Best for | Real-time inference, jobs that cannot checkpoint | Checkpointed training, batch inference, sweeps, evals |
How much notice each platform gives
From each platform's own docs, read October 10, 2026.
| Platform | Listed discount | Notice before reclaim | How your code finds out | What happens to the machine |
|---|---|---|---|---|
| AWS Spot Instances | Up to 90% | 2 minutes |
spot/instance-action in instance metadata returns JSON (404 until then) |
Stopped, hibernated or terminated |
| GCP Spot VMs | Up to 91% | Up to 30 s shutdown by default; an optional 120 s notice period before it |
instance/preempted metadata flips to TRUE, then a soft-off triggers your shutdown script |
Stopped or deleted; GPU Spot VMs are also preempted for maintenance |
| Azure Spot VMs | Up to 90% | 30 s minimum | A Preempt event in Scheduled Events |
Deallocated (disks kept and billed) or deleted with its disks |
| Modal | n/a | Grace period after an interrupt signal | Your exit handler | GPU Functions are always preemptible; restarted on the same input |
| Nodus | n/a | Urgent checkpoint request on reclaim; SIGTERM, then SIGKILL 30 s later on a stop |
Events socket, or a SIGTERM handler |
New attempt on another machine with your state directory restored |
Sources: AWS interruption notices, GCP Spot VMs, Azure Spot VMs, Modal preemption.
What actually breaks when a spot GPU is reclaimed
Most spot guides stop at "use checkpoints". These failures still bite teams who do.
| What breaks | What you see | Fix |
|---|---|---|
| Work since the last save is lost | Run resumes from an old step, or from step 0 | Save on a cadence, and save again when the notice arrives |
| Half-written checkpoint |
torch.load fails or loads a corrupt file on resume |
Write to a temp file, then os.replace it; keep the last two |
| Checkpoint on the machine's local disk | Nothing to resume from; local SSDs and deleted VMs take the data with them | Write to object storage or a volume that outlives the machine |
| Only the weights were saved | Loss spikes or repeated samples after resume | Also save optimizer, LR scheduler, AMP scaler, RNG states, epoch and sampler offset |
| One rank of a multi-node job is reclaimed | Every other rank hangs in an NCCL collective until it times out | Restart the whole gang from a sharded torch.distributed.checkpoint
|
| Script exits 0 after its emergency save | The scheduler marks the run finished, and it never resumes | Exit non-zero after a notice-triggered save |
| Spot pool is empty on relaunch | Job sits pending for hours | Accept several GPU types or regions, or fall back to on-demand |
| A restart loop that never makes progress | You pay for boots that crash before the first save | Cap retries by progress, not by count |
The cost math: when does spot actually win?
Spot is billed for everything the job does, including the work it loses:
billed hours = useful hours x (1 + save overhead) + reclaims x (lost work + restart time)
spot wins while billed hours / useful hours < 1 / (1 - discount)
At 60% off, spot still breaks even when you pay for 2.5x the useful hours. At 30% off, the limit is 1.43x, so restarts eat the discount fast.
A worked example with illustrative rates: 8 GPUs, 20 hours of useful training, $4.00 per GPU-hour on-demand, spot at 60% off ($1.60), 3 reclaims, a 15 minute restart (new machine, image pull, checkpoint load) and a 1 minute save every 30 minutes (3.3% overhead).
| Setup | Billed hours per GPU | Cost for 8 GPUs |
|---|---|---|
| On-demand, no reclaims | 20.0 | $640 |
| Spot, save every 30 min (about 15 min lost per reclaim) | 20.67 + 3 x 0.5 = 22.2 | $284 |
| Spot, save every 30 min and on notice | 20.67 + 3 x 0.25 = 21.4 | $274 |
| Spot, no checkpoints (reclaimed after 7 h, 12 h and 5 h) | 24 + 20 + 3 x 0.25 = 44.75 | $573 |
Checkpointed spot saves about 56%. Spot without checkpoints saves 10% while gambling the whole run, and a fourth reclaim more than about 5 hours into an attempt would push it past on-demand. For real inputs, the Nodus pricing page lists an H100 SXM at $2.60/hr and a B200 at $5.50/hr (as of October 10, 2026), and AWS's Spot Instance Advisor reports interruption frequency in bands from under 5% to over 20% per month.
A notice watcher that works on AWS, GCP and Azure
Run this alongside your training loop (replace train_step and save_checkpoint with your own). A background thread polls each cloud's metadata endpoint and also catches SIGTERM, which Kubernetes and Nodus send before killing a container.
import signal, sys, threading, time, urllib.request
stop = threading.Event()
def _get(url, headers=None, method="GET"):
try:
req = urllib.request.Request(url, headers=headers or {}, method=method)
with urllib.request.urlopen(req, timeout=1) as r:
return r.read().decode()
except Exception:
return None # no notice yet, or a different cloud
def notice_pending():
# AWS: IMDSv2 token, then instance-action exists only after a notice
token = _get("http://169.254.169.254/latest/api/token",
{"X-aws-ec2-metadata-token-ttl-seconds": "300"}, "PUT")
if token and _get("http://169.254.169.254/latest/meta-data/spot/instance-action",
{"X-aws-ec2-metadata-token": token}):
return True
# GCP: preempted flips to TRUE on reclaim
if _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
{"Metadata-Flavor": "Google"}) == "TRUE":
return True
# Azure: a Preempt entry in Scheduled Events
events = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
{"Metadata": "true"})
return bool(events) and '"Preempt"' in events
def watch(every=5):
while not stop.is_set():
if notice_pending():
stop.set()
time.sleep(every)
def on_sigterm(*_):
stop.set()
signal.signal(signal.SIGTERM, on_sigterm)
threading.Thread(target=watch, daemon=True).start()
for step in range(start_step, total_steps):
train_step()
if step % save_every == 0 or stop.is_set():
save_checkpoint(step) # atomic write to durable storage
if stop.is_set():
sys.exit(143) # non-zero: interrupted, not finished
Two limits. A full save must fit inside the notice (30 seconds on GCP and Azure by default), so keep regular saves frequent enough that a missed emergency save costs little. Inside Docker on AWS, the IMDSv2 token has a hop limit of 1 by default; raise it to 2 (aws ec2 modify-instance-metadata-options --http-put-response-hop-limit 2) or the watcher never sees the notice. The checkpointing walkthrough covers the save and resume code itself.
Making spot safe on Nodus
Nodus is an AI compute platform that runs each job on the cheapest capacity that will finish it, with checkpoints and recovery built in. Interruptible capacity is opt-in per job:
pip install nodus-compute
nodus login
nodus get gpus --interruptible -o wide # interruptible offerings, startup time, interruption rate
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 --dry-run -- python train.py
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 -d -- python train.py
nodus get attempts -l nodus.dev/job=<name> # one row per machine, with reasons such as Preempted
--dry-run prints the expected cost and start time without launching. allow uses interruptible capacity only when it is cheaper to finish; prefer (the bare --interruptible) favors it with a 10% allowance; never keeps the job off it. On a reclaim, Nodus requests an urgent checkpoint, prepares a replacement in parallel, starts it only once the old machine is provably gone, and restores /nodus/state. A run with no progress twice in a row stops with NoProgress instead of looping, and --max-cost suspends the job with a checkpoint at the cap. Your code still decides what goes into the state directory and loads it on start. Multi-node gang checkpoints are Beta. New accounts get a $30 starter grant once the email is verified, valid for 30 days.
FAQ
How much notice do I get before a spot GPU is reclaimed? Two minutes on AWS, 30 seconds by default on GCP (with an optional 120 second notice period) and at least 30 seconds on Azure, as of October 10, 2026.
How often do spot GPUs get interrupted? It varies by GPU, region and week. AWS publishes monthly interruption bands per instance pool; nodus get gpus -o wide shows a measured rate per interruptible offering.
Can I serve inference on spot GPUs? Batch and offline inference, yes. For real-time traffic, keep an on-demand floor and use spot only for overflow.
Is spot cheaper than reserved capacity? For bursty, checkpointed work, usually. If GPUs stay busy most hours for months, a reservation can beat spot with no reclaim risk; Nodus offers Reserved Compute in Beta.
What is the difference between spot and preemptible VMs on GCP? Preemptible VMs are the older model with a 24 hour maximum runtime. Spot VMs have no maximum runtime unless you set one.
More guides like this on the Nodus Compute blog.
Top comments (1)
tr.ee/dev-to