TL;DR: Start with the workload’s GPU and memory needs, then search for offers with reliability at least 0.97, free inbound traffic, enough disk, and a driver compatible with your container. Sort by hourly price, inspect the raw offer data, and choose the cheapest offer that passes every check. A low hourly rate is useful only if the instance can run the job and receive its model weights without a surprise charge.
Start with the workload
This is article 1 of our eight-part series on renting GPUs through Vast.ai for real AI work. We are a small content team, and our jobs tend to involve downloading model weights, running batches, and copying the results home. That makes offer selection a practical engineering decision rather than a contest to find the lowest number on the search page.
Before searching, we write down what the job requires:
- A GPU with enough memory for the model and settings we plan to use.
- A container image whose CUDA requirements the host driver can support.
- Enough disk for the image, weights, inputs, temporary files, and outputs.
- Acceptable reliability and network charges.
- The connectivity our workflow needs for setup and file transfer.
For a large model download, we provision at least 150 GB of disk. A 40 GB disk can be enough for smaller workloads, but it is a poor default when the weights alone are substantial. Our MiniMax H3 setup downloaded 72 GB of weights. Disk space was part of making that run predictable.
GPU memory needs also depend on the exact workflow. LatentSync 1.6 at 512 px needed more memory than a shared local GPU had available to us, while LatentSync 1.5 at 256 px ran locally. We would not infer that every lip-sync job needs the same rented GPU. Pick the workload first; then search for hardware that fits it.
If you are setting up an account to follow along, this is our Vast.ai referral link.
Search for offers with the CLI
The basic search command is:
vastai search offers \
'reliability >= 0.97 inet_down_cost = 0 disk_space >= 150' \
-o 'dph+' \
--raw
The quoted expression filters offers. reliability >= 0.97 is our starting threshold. inet_down_cost = 0 excludes offers that charge for inbound traffic. disk_space >= 150 leaves room for a large download. The -o 'dph+' option puts lower hourly prices first, and --raw returns data that we can inspect or pass to a script.
These are starting filters, not a complete compatibility test. An offer can pass them and still have the wrong GPU, an unsuitable driver, or insufficient connectivity for your setup. Search results also change as machines become available or get rented. Treat an offer ID as a candidate to inspect, not as a permanent recommendation.
The CLI accepts fields including gpu_name, num_gpus, disk_space, cuda_vers, cuda_max_good, driver_version, direct_port_count, inet_down, inet_down_cost, rentable, verified, and geolocation. GPU names in query expressions use underscores where the displayed names contain spaces. We use the fields relevant to the job and then examine the returned records before creating an instance.
Here is how the filters map to the problems we actually hit:
| Filter | Problem it prevents |
|---|---|
reliability >= 0.97 |
Hosts that drop out mid-run |
inet_down_cost = 0 |
A separate bill for downloading model weights |
disk_space >= 150 |
Running out of room while pulling weights |
driver_version / cuda_max_good
|
A host driver too old for the container image |
-o 'dph+' |
Paying more than needed among offers that pass everything else |
Check inbound traffic before hourly price
Inbound traffic was our clearest lesson in reading offers carefully. In an earlier test, we rented an RTX PRO 6000 Blackwell 96GB at $1.625 per hour. The 56-minute session cost $2.18: $1.47 for GPU time and $0.64 for inbound traffic after downloading 245 GB of model weights.
That download charge was material compared with the compute charge. A different offer with a slightly higher hourly rate and free inbound traffic could have been cheaper for a weight-heavy setup. We now check inet_down_cost alongside the hourly price, before we commit to an offer.
The amount you download depends on what is already present in your image and which weights the job needs. Our bootstrap image downloads only the selected weight groups at startup. That reduces unnecessary transfers, but it does not remove the need to inspect the offer’s network pricing.
When comparing offers, estimate the whole session: setup and download time, processing time, disk, and any traffic charges. The hourly GPU figure describes only one part of that bill.
Check driver, CUDA, and the image together
Our container image and the host driver must work together. The raw offer fields driver_version, cuda_vers, and cuda_max_good help us evaluate that before renting. We compare them with the requirements of the image we intend to launch.
We learned this by encountering an A100 offer in Sweden with driver 535 that did not fit our planned image. The GPU name and memory were appealing, but they did not make the software stack compatible. We moved on to another offer rather than treating the GPU model as proof that the job would run.
For our LatentSync 1.6 setup, we used pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime on a Vast RTX 3090. The exact image matters when checking an offer: changing the image may change the CUDA and driver requirements. Write down the image you plan to use, read its requirements, and compare them with the offer’s raw fields. If a field is absent or unclear, inspect the offer further before creating an instance.
The same applies to ports and location. Check direct_port_count if your setup depends on direct connectivity. Use geolocation when location matters to your workflow. We do not assume that a familiar GPU name tells us anything about those fields.
Read raw results as data
Raw JSON makes it easier to review several offers consistently. We first save a sorted search result:
vastai search offers \
'reliability >= 0.97 inet_down_cost = 0 disk_space >= 150' \
-o 'dph+' \
--raw > offers.json
python pick_offer.py < offers.json
Here is a small pick_offer.py helper. It takes the first offer that passes the repeated checks because the CLI search was sorted by ascending hourly price. Set WANTED_GPU if you want to restrict the result to a GPU name. Set MIN_CUDA or MIN_DRIVER only after checking the requirements of your chosen image.
import json
import os
import re
import sys
def number(value):
try:
return float(value)
except (TypeError, ValueError):
return None
def version(value):
parts = re.findall(r"\d+", str(value))
return tuple(int(part) for part in parts)
def meets_version(actual, minimum):
return bool(actual) and version(actual) >= version(minimum)
payload = json.load(sys.stdin)
offers = payload if isinstance(payload, list) else payload.get("offers")
if not isinstance(offers, list):
raise SystemExit("Expected a list of offers; inspect the raw JSON format.")
wanted_gpu = os.environ.get("WANTED_GPU", "").lower().replace("_", " ")
min_cuda = os.environ.get("MIN_CUDA")
min_driver = os.environ.get("MIN_DRIVER")
for offer in offers:
reliability = number(offer.get("reliability"))
inbound_cost = number(offer.get("inet_down_cost"))
disk = number(offer.get("disk_space"))
gpu = str(offer.get("gpu_name", "")).lower().replace("_", " ")
if reliability is None or reliability < 0.97:
continue
if inbound_cost != 0 or disk is None or disk < 150:
continue
if wanted_gpu and wanted_gpu not in gpu:
continue
if min_cuda and not meets_version(offer.get("cuda_max_good"), min_cuda):
continue
if min_driver and not meets_version(offer.get("driver_version"), min_driver):
continue
print(json.dumps({
"id": offer.get("id"),
"gpu_name": offer.get("gpu_name"),
"reliability": reliability,
"inet_down_cost": inbound_cost,
"disk_space": disk,
"cuda_max_good": offer.get("cuda_max_good"),
"driver_version": offer.get("driver_version"),
"geolocation": offer.get("geolocation"),
}, indent=2))
break
else:
raise SystemExit("No matching offer in these search results.")
This helper is intentionally a shortlist tool. It relies on the order supplied by -o 'dph+'; it does not calculate a total job cost or prove that an image will launch. Inspect the selected record, confirm its GPU memory and connectivity, and check the image’s driver requirements before creating an instance. If your CLI’s raw response has a different JSON shape, inspect it and adjust the offers extraction rather than silently choosing from incomplete data.
What happened when we selected carefully
For MiniMax H3 image-to-video, we used an A100 SXM4 80GB offer in Czechia at $1.19 per GPU hour, about $1.25 per hour with disk. The job produced 82 hero clips in a session costing $3.59, including setup. Weight download and setup were part of the session, so the per-clip processing time alone would understate its true cost.
For LatentSync 1.6, we used an RTX 3090 offer in the UAE at $0.27 per hour with free inbound traffic. Setup took about 9 minutes. That is why we batch jobs into one session: repeated setup can dominate a short run. In one queued batch self-test, setup took 11 minutes of a 15.9-minute run.
Our selection checklist now has a final operational step. After copying outputs back, we verify the backup and destroy the instance. A stopped instance continues billing for storage. We once left one stopped for days, lost access when the balance went negative, and never downloaded the trained result. A good offer cannot compensate for an unfinished shutdown procedure.
What it cost us
- MiniMax H3 on an A100 SXM4 80GB: $1.19 per GPU hour, about $1.25 with disk; 82 hero clips cost $3.59 for a session of about 3 hours, including setup.
- Earlier RTX PRO 6000 Blackwell test: $2.18 over 56 minutes, including $1.47 for GPU time and $0.64 for inbound traffic after 245 GB of weight downloads.
- LatentSync 1.6 on an RTX 3090 with free inbound: $0.27 per hour; one clip took 238 seconds and cost about $0.09.
- A stopped instance with a 90 GB disk billed about $0.0167 per hour, or about $0.40 per day, until it was destroyed.
If you want to follow along with the rest of the series, our referral link is here. The cheapest useful offer is the one that satisfies the workload’s requirements at the lowest session cost. For us, reliability, inbound traffic, disk, and driver compatibility are the first checks. Hourly price decides among the offers left standing.
Top comments (0)