DEV Community

Viswaretas Kotra
Viswaretas Kotra

Posted on

Spot vs on-demand GPUs: what breaks, and how do you make spot safe?

Short answer (October 10, 2026): spot GPUs are the same hardware as on-demand GPUs at a deep discount (AWS and Azure advertise up to 90% off, GCP up to 91%), but the cloud can take the machine back with 30 seconds to 2 minutes of notice, and anything not saved somewhere durable is gone. Spot is safe when your job checkpoints to storage that outlives the machine, saves again when the notice arrives, resumes on start and can fall back when spot capacity runs out; without that, a few reclaims can make spot cost more than on-demand.

Below: what fails on a reclaim, how much warning each platform gives, when spot really wins, and a copy-paste setup.

Spot vs on-demand at a glance

On-demand Spot (interruptible)
Hourly price List price Discounted, often 60% to 90% below list
Runs until You stop it You stop it, or the cloud needs the capacity back
Warning before a reclaim None needed 30 seconds to 2 minutes
Capacity when you ask Usually there Can be empty for hours
Best for Real-time inference, jobs that cannot checkpoint Checkpointed training, batch inference, sweeps, evals

How much notice each platform gives

From each platform's own docs, read October 10, 2026.

Platform Listed discount Notice before reclaim How your code finds out What happens to the machine
AWS Spot Instances Up to 90% 2 minutes spot/instance-action in instance metadata returns JSON (404 until then) Stopped, hibernated or terminated
GCP Spot VMs Up to 91% Up to 30 s shutdown by default; an optional 120 s notice period before it instance/preempted metadata flips to TRUE, then a soft-off triggers your shutdown script Stopped or deleted; GPU Spot VMs are also preempted for maintenance
Azure Spot VMs Up to 90% 30 s minimum A Preempt event in Scheduled Events Deallocated (disks kept and billed) or deleted with its disks
Modal n/a Grace period after an interrupt signal Your exit handler GPU Functions are always preemptible; restarted on the same input
Nodus n/a Urgent checkpoint request on reclaim; SIGTERM, then SIGKILL 30 s later on a stop Events socket, or a SIGTERM handler New attempt on another machine with your state directory restored

Sources: AWS interruption notices, GCP Spot VMs, Azure Spot VMs, Modal preemption.

What actually breaks when a spot GPU is reclaimed

Most spot guides stop at "use checkpoints". These failures still bite teams who do.

What breaks What you see Fix
Work since the last save is lost Run resumes from an old step, or from step 0 Save on a cadence, and save again when the notice arrives
Half-written checkpoint torch.load fails or loads a corrupt file on resume Write to a temp file, then os.replace it; keep the last two
Checkpoint on the machine's local disk Nothing to resume from; local SSDs and deleted VMs take the data with them Write to object storage or a volume that outlives the machine
Only the weights were saved Loss spikes or repeated samples after resume Also save optimizer, LR scheduler, AMP scaler, RNG states, epoch and sampler offset
One rank of a multi-node job is reclaimed Every other rank hangs in an NCCL collective until it times out Restart the whole gang from a sharded torch.distributed.checkpoint
Script exits 0 after its emergency save The scheduler marks the run finished, and it never resumes Exit non-zero after a notice-triggered save
Spot pool is empty on relaunch Job sits pending for hours Accept several GPU types or regions, or fall back to on-demand
A restart loop that never makes progress You pay for boots that crash before the first save Cap retries by progress, not by count

The cost math: when does spot actually win?

Spot is billed for everything the job does, including the work it loses:

billed hours = useful hours x (1 + save overhead) + reclaims x (lost work + restart time)
spot wins while billed hours / useful hours < 1 / (1 - discount)
Enter fullscreen mode Exit fullscreen mode

At 60% off, spot still breaks even when you pay for 2.5x the useful hours. At 30% off, the limit is 1.43x, so restarts eat the discount fast.

A worked example with illustrative rates: 8 GPUs, 20 hours of useful training, $4.00 per GPU-hour on-demand, spot at 60% off ($1.60), 3 reclaims, a 15 minute restart (new machine, image pull, checkpoint load) and a 1 minute save every 30 minutes (3.3% overhead).

Setup Billed hours per GPU Cost for 8 GPUs
On-demand, no reclaims 20.0 $640
Spot, save every 30 min (about 15 min lost per reclaim) 20.67 + 3 x 0.5 = 22.2 $284
Spot, save every 30 min and on notice 20.67 + 3 x 0.25 = 21.4 $274
Spot, no checkpoints (reclaimed after 7 h, 12 h and 5 h) 24 + 20 + 3 x 0.25 = 44.75 $573

Checkpointed spot saves about 56%. Spot without checkpoints saves 10% while gambling the whole run, and a fourth reclaim more than about 5 hours into an attempt would push it past on-demand. For real inputs, the Nodus pricing page lists an H100 SXM at $2.60/hr and a B200 at $5.50/hr (as of October 10, 2026), and AWS's Spot Instance Advisor reports interruption frequency in bands from under 5% to over 20% per month.

A notice watcher that works on AWS, GCP and Azure

Run this alongside your training loop (replace train_step and save_checkpoint with your own). A background thread polls each cloud's metadata endpoint and also catches SIGTERM, which Kubernetes and Nodus send before killing a container.

import signal, sys, threading, time, urllib.request

stop = threading.Event()

def _get(url, headers=None, method="GET"):
    try:
        req = urllib.request.Request(url, headers=headers or {}, method=method)
        with urllib.request.urlopen(req, timeout=1) as r:
            return r.read().decode()
    except Exception:
        return None  # no notice yet, or a different cloud

def notice_pending():
    # AWS: IMDSv2 token, then instance-action exists only after a notice
    token = _get("http://169.254.169.254/latest/api/token",
                 {"X-aws-ec2-metadata-token-ttl-seconds": "300"}, "PUT")
    if token and _get("http://169.254.169.254/latest/meta-data/spot/instance-action",
                      {"X-aws-ec2-metadata-token": token}):
        return True
    # GCP: preempted flips to TRUE on reclaim
    if _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
            {"Metadata-Flavor": "Google"}) == "TRUE":
        return True
    # Azure: a Preempt entry in Scheduled Events
    events = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
                  {"Metadata": "true"})
    return bool(events) and '"Preempt"' in events

def watch(every=5):
    while not stop.is_set():
        if notice_pending():
            stop.set()
        time.sleep(every)

def on_sigterm(*_):
    stop.set()

signal.signal(signal.SIGTERM, on_sigterm)
threading.Thread(target=watch, daemon=True).start()

for step in range(start_step, total_steps):
    train_step()
    if step % save_every == 0 or stop.is_set():
        save_checkpoint(step)   # atomic write to durable storage
    if stop.is_set():
        sys.exit(143)           # non-zero: interrupted, not finished
Enter fullscreen mode Exit fullscreen mode

Two limits. A full save must fit inside the notice (30 seconds on GCP and Azure by default), so keep regular saves frequent enough that a missed emergency save costs little. Inside Docker on AWS, the IMDSv2 token has a hop limit of 1 by default; raise it to 2 (aws ec2 modify-instance-metadata-options --http-put-response-hop-limit 2) or the watcher never sees the notice. The checkpointing walkthrough covers the save and resume code itself.

Making spot safe on Nodus

Nodus is an AI compute platform that runs each job on the cheapest capacity that will finish it, with checkpoints and recovery built in. Interruptible capacity is opt-in per job:

pip install nodus-compute
nodus login
nodus get gpus --interruptible -o wide    # interruptible offerings, startup time, interruption rate
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 --dry-run -- python train.py
nodus run --gpu H100 --interruptible=allow --checkpoint /nodus/state --max-cost 25 -d -- python train.py
nodus get attempts -l nodus.dev/job=<name>  # one row per machine, with reasons such as Preempted
Enter fullscreen mode Exit fullscreen mode

--dry-run prints the expected cost and start time without launching. allow uses interruptible capacity only when it is cheaper to finish; prefer (the bare --interruptible) favors it with a 10% allowance; never keeps the job off it. On a reclaim, Nodus requests an urgent checkpoint, prepares a replacement in parallel, starts it only once the old machine is provably gone, and restores /nodus/state. A run with no progress twice in a row stops with NoProgress instead of looping, and --max-cost suspends the job with a checkpoint at the cap. Your code still decides what goes into the state directory and loads it on start. Multi-node gang checkpoints are Beta. New accounts get a $30 starter grant once the email is verified, valid for 30 days.

FAQ

How much notice do I get before a spot GPU is reclaimed? Two minutes on AWS, 30 seconds by default on GCP (with an optional 120 second notice period) and at least 30 seconds on Azure, as of October 10, 2026.

How often do spot GPUs get interrupted? It varies by GPU, region and week. AWS publishes monthly interruption bands per instance pool; nodus get gpus -o wide shows a measured rate per interruptible offering.

Can I serve inference on spot GPUs? Batch and offline inference, yes. For real-time traffic, keep an on-demand floor and use spot only for overflow.

Is spot cheaper than reserved capacity? For bursty, checkpointed work, usually. If GPUs stay busy most hours for months, a reservation can beat spot with no reclaim risk; Nodus offers Reserved Compute in Beta.

What is the difference between spot and preemptible VMs on GCP? Preemptible VMs are the older model with a 24 hour maximum runtime. Spot VMs have no maximum runtime unless you set one.

More guides like this on the Nodus Compute blog.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to