DEV Community

Big Mazzy
Big Mazzy

Posted on Originally published at serverrental.store

Renting a GPU Server for AI Video Generation: What We Learned (Costs, Pitfalls, Checklist)

💡 Need a GPU server for this? You can rent NVIDIA GPU servers by the hour at Vast.ai — pay-as-you-go, no long-term contract.

Recently we rented a GPU for what should have been a two-hour experiment. It ended with a surprise line on the invoice: $0.64 of a $2.18 total was just inbound traffic for downloading model weights. A few days later a stopped instance we forgot about kept charging us about $0.40 per day. Here is what we learned running AI video generation on rented GPUs, so you can skip the expensive lessons.

Why rent instead of using an API

For image-to-video, API pricing adds up fast. One hosted API we compared charges roughly $0.08 per second of video, so a 5-second clip costs about $0.40. On a rented GPU we measured about $0.03 per clip, roughly 10-20x cheaper once you batch work. The catch: you handle setup, weights and cleanup yourself.

What we actually ran

Job Model GPU Approx. price Result
Image-to-video, 5 s clips MiniMax H3 (turbo 8-step LoRA) A100 80GB ~$1.2/h ~86 s per clip, ~$0.03
Text-to-video test MiniMax H3 turbo 96GB workstation GPU ~$1.6/h ~51 s per 5 s clip
Depth-guided video Wan 2.1 VACE 14B 96GB workstation GPU ~$1.6/h ~2 min per clip
Lip-sync LatentSync 1.6 RTX 3090 ~$0.27/h ~238 s per clip, ~$0.09

Notes on quality: H3 kept identity (face, clothes, art style) well from a first frame, and using the same image as first and last frame gave clean loops. Wan looked more "CGI" in our tests. Lip-sync on the cheap 3090 was perfectly fine, so don't overpay for a big card on small models.

Pitfall 1: inbound traffic is not always free

Video model weights are huge. One of our test stacks pulled about 245 GB. Some marketplace hosts bill inbound traffic per GB, and that turned into 30% of our bill. Filter for hosts where inbound cost is zero before you click rent.

With the Vast.ai CLI you can check this in the offer search output (look at the inbound cost column):

vastai search offers 'reliability>0.97 disk_space>=150 gpu_name=A100_SXM4 inet_down_cost<=0.001' -o 'dph'
Enter fullscreen mode Exit fullscreen mode

Pitfall 2: stopped is not free

A stopped instance still bills for its disk. Ours held a 90 GB volume and cost about $0.0167/h, roughly $0.40 per day, until it drained the account balance. Worse, a stopped instance can't always be restarted when the host GPU is taken or your balance is negative.

The rule we follow now: destroy, don't stop. Download results first.

# copy results out, then destroy (the CLI asks y/N, so pipe the answer)
vastai copy <instance_id>:/workspace/out ./out
echo y | vastai destroy instance <instance_id>
# verify nothing is left running or stopped
vastai show instances
Enter fullscreen mode Exit fullscreen mode

Pitfall 3: weak hosts

  • Reliability: pick hosts with at least 97% reliability. Below that we saw failed pulls and dropped sessions.
  • CUDA driver: a cheap A100 with an old driver (535) could not run the newer CUDA builds we needed. Check the driver version in the offer.
  • Disk: allocate at least 150 GB. Weights plus the model cache plus outputs fill 100 GB quickly.

Pitfall 4: paying to wait

Setup time is billed too. Our weights download took about 8 minutes for a 72 GB model group, and lip-sync environment setup took about 9 minutes. So:

  • Prepare a Docker image with your dependencies and a startup script that fetches only the weights you need.
  • Batch all jobs into one session. Build your job list (prompts, first frames, audio) locally, upload once, run the whole queue, download once.
  • Prepare the input images at the exact target aspect ratio. Our model stretched a first frame that didn't match the canvas.

Cost example

Item Cost
A100 80GB, 3 hours including setup ~$3.6
~80 clips at ~86 s each (about 2 hours of the session) included above
Effective cost per clip including setup ~$0.045
Lip-sync session (3090, ~9 min setup + a test clip) under $0.15

Checklist

  1. Filter offers: reliability >= 97%, disk >= 150 GB, recent CUDA driver, inbound traffic free.
  2. Use a prebuilt image and an on-start script; download only the weights you need.
  3. Prepare all inputs locally and queue the whole batch in one session.
  4. Test one clip first, then run the full batch.
  5. Download outputs and back them up somewhere else.
  6. Destroy the instance (not stop) and confirm the instance list is empty.
  7. Check the account balance the next day.

Marketplace vs managed GPU cloud

Summary

Renting a GPU for video generation can cut your per-clip cost by an order of magnitude versus APIs, but only if you avoid the hidden costs: paid inbound traffic, stopped-but-billing instances and idle setup time. Choose reliable hosts, batch your work, back up the results and destroy the instance when done.


Ready to try it? Rent a GPU server on Vast.ai and follow the checklist above.

Top comments (0)