TL;DR We left a Vast.ai instance stopped after a run and assumed the meter had stopped with it. The GPU charge had stopped, but its disk kept billing at about $0.0167/hour, or $0.40/day for our 90 GB disk. The balance eventually went negative; we could not restart the instance to retrieve a trained result. Our rule now is simple: download and verify the output, then destroy the instance. We also launch a detached watchdog that attempts destruction after a deadline.
The bill we missed
We are a small content team renting GPUs on Vast.ai for AI video work. Renting works well for jobs that need more memory than our local machines have, especially when we can queue many clips into one session. But a rented machine has more than one meter. GPU time is the obvious one; attached storage is easier to forget.
Our mistake was leaving an instance in the STOPPED state for days. Its 90 GB disk continued to cost roughly $0.0167 per hour, about $0.40 per day. We had not downloaded the trained result. By the time we returned, our balance was negative, and we could not start the stopped instance to retrieve it. We never downloaded that result.
The distinction now shapes our shutdown procedure:
| Action | What we expect from it | When we use it |
|---|---|---|
| Stop | End the running GPU session while retaining the instance and its storage | Only when we deliberately need to retain that disk |
| Destroy | Remove the instance after its output is safe elsewhere | At the end of a completed batch |
Stopping can be useful when you intend to resume with the same disk. It is a poor substitute for cleanup. Before choosing it, we ask whether the data on that disk is worth the ongoing storage charge and whether we have a clear plan to return.
Finish the job before ending the rental
Our shutdown sequence begins with the output, not the instance command. We copy results back, back them up, and verify that the copy is complete. In one batch, we checked that 5,342 files and 1.379 GiB matched locally before destroying the instance.
A successful copy command alone is not the whole check. We inspect the destination file count and total size, and we make sure the files we need are actually present. Only then do we release the rental. This order matters because destroying the instance also removes the convenient path back to anything left on its disk.
The Vast CLI provides a copy command, but the source and destination depend on where you keep your results. A session might use the CLI to copy from the instance to a local directory:
vastai copy "$REMOTE_SOURCE" "$LOCAL_DESTINATION"
Set those variables to the actual source and destination for your session. We also back up the local result to a cloud remote and verify that copy. The extra check costs a little attention; losing an output after paying for its generation costs much more.
Destroy explicitly, then check
Once the output is safe, we destroy the instance rather than leave it stopped. The CLI command we used asks for y/N, so unattended scripts need to answer the prompt. This form works with the CLI behavior we encountered:
echo y | vastai destroy instance "$INSTANCE_ID"
Newer CLI versions also have a -y option. We use the piped answer in the watchdog below because it matches the prompt behavior we verified.
Do not read “the command ran” as “the account is clean.” A network problem, CLI error, or stale listing can obscure what happened. We had a particularly confusing case where the old v0 instance listing was deprecated and returned an error that looked like “0 instances.” An error is not evidence that an instance is gone.
For a final check, we query the v1 instances endpoint directly. Keep the API key in an environment variable; do not paste it into a script or print it in a log:
: "${VAST_API_KEY:?Set VAST_API_KEY in your environment}"
curl -fsS \
-H "Authorization: Bearer ${VAST_API_KEY}" \
"https://console.vast.ai/api/v1/instances/" |
python -m json.tool
Inspect the response for the instance ID you destroyed. Also check that the request succeeded and that you received a valid response before treating an absent ID as confirmation. If the ID remains, investigate and retry destruction. We make this check after the session, even when the destroy command reports success.
A deadline that survives the job script
Copying and cleanup still rely on a controller reaching the end of its run. A crash can leave the GPU meter running. We therefore launch a separate, detached watchdog when the instance starts. It waits for a deadline, then tries to destroy the instance up to three times.
Here is the Bash script. Save it somewhere on the machine controlling the Vast session as watchdog.sh:
#!/usr/bin/env bash
set -u
instance_id=${1:?Pass an instance ID}
deadline_seconds=${2:?Pass a deadline in seconds}
retry_wait_seconds=${3:?Pass a retry wait in seconds}
case "$deadline_seconds:$retry_wait_seconds" in
*[!0-9:]* | :* | *:) exit 2 ;;
esac
sleep "$deadline_seconds"
for attempt in 1 2 3; do
if echo y | vastai destroy instance "$instance_id"; then
exit 0
fi
if [ "$attempt" -lt 3 ]; then
sleep "$retry_wait_seconds"
fi
done
exit 1
Launch it after you have an instance ID and have chosen a deadline and retry interval:
nohup bash watchdog.sh \
"$INSTANCE_ID" \
"$RUN_SECONDS" \
"$RETRY_WAIT_SECONDS" \
>./vast-watchdog.log 2>&1 </dev/null &
The deadline is a backstop, not the normal end of a job. We still copy, verify, and destroy promptly when the batch completes. Choose RUN_SECONDS from the offer’s hourly cost and the maximum spend you are willing to leave exposed. We used a $1 cost cap for our workflow. Allow for setup time as well as rendering, and remember that a timer based on GPU hours does not account for every possible charge, including inbound traffic or storage retained after a stop.
A detached process can survive the controlling shell exiting. It cannot help if the machine running that process goes down, its credentials stop working, or Vast cannot be reached. Those limits are why we also check the v1 instance list after cleanup. Keep the watchdog log until that check is complete; it can show whether the deadline fired and whether a destroy attempt failed.
If your normal job finishes before the deadline, destroy the instance yourself after verification. The watchdog may later attempt to destroy an ID that is already gone. That is less concerning than leaving a live instance unattended, but the final API check remains the authoritative step in our procedure.
Prevent a setup mistake from becoming a cleanup bill
The best watchdog still leaves room for avoidable costs during setup. We now inspect offers before launching. In particular, we check inet_down_cost: an earlier RTX PRO 6000 Blackwell test incurred $0.64 of inbound traffic billing while pulling 245 GB of model weights. The GPU portion was $1.47; the total was $2.18. For subsequent searches, we look for free inbound traffic where practical.
We also check reliability, driver and CUDA compatibility, and disk space. An A100 offer in Sweden had an old driver, version 535, that did not fit our image. Large model downloads need generous disk provision; we use at least 150 GB for those sessions. A smaller disk may suit smaller weights, but it is a bad place to discover that a download will not fit.
For repeatable setup, we built a public Docker image through GitHub Actions and used an on-start script to download only the weight group needed for the run. The selected group comes from DL_GROUPS, such as DL_GROUPS=h3. That keeps the instance setup tied to the job we intend to run. We still batch many jobs into one session: setup took about 9–11 minutes in our lip-sync runs, which is a large share of a short rental.
If you are setting up a Vast.ai account for this workflow, this is our referral link. Check the current offer details and pricing yourself before renting; they are more useful to your budget than any price from one of our past sessions.
Make “done” mean no instance remains
Our practical end state has three parts: the results are local, the backup matches, and the instance is absent from the v1 list. A stopped instance satisfies none of those by itself. The watchdog covers a missed deadline; it does not replace checking the data or the account afterward.
What it cost us
The stopped 90 GB disk billed about $0.0167/hour, or $0.40/day, while we left it for days. Our balance went negative, we could not restart that instance, and we never downloaded its trained result. In a separate Blackwell test, $0.64 of a $2.18 total came from inbound traffic for 245 GB of weights; GPU time was $1.47. For a later A100 image-to-video session, 82 hero clips cost $3.59 over about 3 hours, including setup. These are our observed charges, not promises for another offer or workload.
The storage loss changed a small habit with a large consequence. We now treat “download, verify, destroy, verify again” as part of the job itself. The render is finished only when its files are safe and the rental is gone.
Top comments (0)