TL;DR: We ran LatentSync 1.6 on a rented RTX 3090 at $0.27/h. Our 720×1280 talking-head clip with eight seconds of TTS audio took 238 seconds to process, using the 512-pixel stage, 20 inference steps, guidance scale 1.5, and DeepCache. The recorded cost was about $0.09 for that clip. Setup took roughly nine minutes, so a queue that keeps the GPU busy matters as much as inference speed. Download every result before destroying the instance: a stopped instance still bills for storage.
Why we rented the GPU
We are a small content team using rented GPUs for AI video work. LatentSync 1.6 gave us a useful way to align an existing talking-head video with speech audio, but its 512-pixel stage needed more VRAM than the shared local GPU we had available. LatentSync 1.5’s 256-pixel stage ran locally; for 1.6, we used a Vast.ai RTX 3090 in the UAE.
That is a practical distinction when planning a batch. A model that starts locally at a smaller resolution does not tell you whether its larger configuration will fit alongside everything else sharing that GPU. Renting let us give the run a dedicated card and return to local work when the batch finished.
If you need a Vast account, this is our referral link. Check the live offer details before renting; the hourly price is only one part of the bill.
Choose an offer for the whole run
Our offer was $0.27/h and had free inbound traffic. The latter mattered because model setup involves downloading weights. In an earlier video-model test, inbound traffic billing was substantial enough to change the economics of the session. We now check inet_down_cost before accepting an offer.
We also filter for reliability of at least 0.97, enough disk for the image, repository, inputs, weights, and outputs, and a driver compatible with the container’s CUDA version. For larger model downloads, we provision at least 150 GB rather than trying to fit everything into a small disk. Check cuda_max_good and driver_version; we have encountered an otherwise attractive A100 offer whose old driver did not fit the intended environment.
The following is the Vast CLI pattern we use to inspect offers and create an instance. Replace OFFER_ID with an offer you have checked. Keep your API key in the CLI’s configured environment or credentials, never in a script or terminal transcript you plan to share.
vastai search offers 'gpu_name=RTX_3090 num_gpus=1 reliability>=0.97 inet_down_cost=0 disk_space>=150 rentable=true' -o 'dph+' --raw
vastai create instance OFFER_ID \
--image pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime \
--disk 150 \
--label latentsync-batch \
--ssh --direct
vastai show instances --raw
vastai ssh-url ID
An offer can look suitable in search results and still deserve inspection. Check the reported price, inbound traffic cost, disk, driver, CUDA compatibility, reliability, and whether it is actually rentable before creating the instance. We use --raw when we want to inspect fields without relying on a formatted table.
Build the environment once
For this run, the starting image was pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime. We installed build-essential because the insightface dependency needed g++, then installed the LatentSync repository’s requirements.txt. We downloaded the weights from ByteDance/LatentSync-1.6 on Hugging Face, including whisper/tiny.pt and latentsync_unet.pt. The weights we needed were about 5 GB.
Setup took roughly nine minutes. That is long enough to affect the cost of a one-clip experiment, and it is time you pay again if you destroy the instance and start a new one. We prepare input videos and audio before renting, then submit as many ready jobs as we can in the same session.
We keep the repository and its inference environment separate from the batch queue. This makes it easier to see whether a problem came from a missing dependency, a particular input pair, or the queue runner. Before starting a batch, we run one known input through the installed repository’s documented inference command and confirm that its output plays. We then put that working command behind a small adapter that accepts three paths: input video, input audio, and output video.
We avoid copying an inference command from an unrelated LatentSync version into the queue script. The script below manages files and failures; the adapter holds the invocation verified against the version actually installed on the instance.
The settings we used
Our test paired a 720×1280 talking-head video with eight seconds of TTS audio. We used LatentSync 1.6’s stage2_512 configuration, 20 inference steps, guidance scale 1.5, and DeepCache. Processing took 238 seconds for the clip.
Those are the settings and timing from our run, rather than a promise for every input. Video length, source frames, preprocessing, and the installed environment can all affect a run. In particular, decide how your adapter handles a mismatch between video duration and audio duration before sending a large batch. Inspect the first output for framing, mouth motion, and the beginning and end of speech.
| Item | Our run |
|---|---|
| GPU and location | Vast RTX 3090, UAE |
| Rental rate | $0.27/h; free inbound traffic |
| Input | 720×1280 talking-head clip and 8 s TTS audio |
| Model settings | LatentSync 1.6, stage2_512, 20 steps, guidance scale 1.5, DeepCache |
| Processing time | 238 s per clip |
| Recorded clip cost | About $0.09 |
| Environment setup | About 9 min |
The recorded clip cost includes more than the 238 seconds spent processing. Do not estimate a session by multiplying inference time alone by the hourly rate: setup, transfers, inspection, and any idle time also sit on the meter.
A directory queue we can inspect
Our simplest batch layout uses matching filename stems. Put take01.mp4 and take01.wav in queue/; the runner sends them to the adapter and expects output/take01.mp4. On success, it moves both inputs to done/. On failure, it leaves the pair in queue/ for inspection and records the error in failed/.
Save the runner as run_queue.py. Here it is. It processes one job at a time, which keeps both GPU use and failure handling easy to follow. The adapter must be an executable file whose three arguments are the video input path, audio input path, and output path. Configure that adapter with the inference command and settings you verified in your installed LatentSync checkout.
#!/usr/bin/env python3
import argparse
import shutil
import subprocess
import time
from pathlib import Path
parser = argparse.ArgumentParser()
parser.add_argument("--adapter", type=Path, required=True)
parser.add_argument("--root", type=Path, default=Path("/workspace/lipsync"))
args = parser.parse_args()
root = args.root
queue = root / "queue"
done = root / "done"
output = root / "output"
failed = root / "failed"
for directory in (queue, done, output, failed):
directory.mkdir(parents=True, exist_ok=True)
if not args.adapter.is_file():
parser.error(f"adapter does not exist: {args.adapter}")
for video in sorted(queue.glob("*.mp4")):
audio = queue / f"{video.stem}.wav"
final = output / video.name
partial = output / f"{video.stem}.partial.mp4"
if not audio.is_file():
print(f"SKIP {video.name}: matching WAV is missing", flush=True)
continue
if final.exists():
print(f"SKIP {video.name}: output already exists", flush=True)
continue
partial.unlink(missing_ok=True)
started = time.monotonic()
print(f"START {video.name}", flush=True)
with (failed / f"{video.stem}.log").open("w") as log:
result = subprocess.run(
[str(args.adapter), str(video), str(audio), str(partial)],
stdout=log,
stderr=subprocess.STDOUT,
check=False,
)
if result.returncode != 0 or not partial.is_file():
partial.unlink(missing_ok=True)
print(f"FAILED {video.name}; see failed/{video.stem}.log",
flush=True)
continue
partial.replace(final)
shutil.move(str(video), str(done / video.name))
shutil.move(str(audio), str(done / audio.name))
(failed / f"{video.stem}.log").unlink(missing_ok=True)
elapsed = time.monotonic() - started
print(f"DONE {video.name} in {elapsed:.1f}s", flush=True)
Run it with the adapter you tested:
python3 run_queue.py \
--adapter /workspace/run_latentsync_job.sh \
--root /workspace/lipsync
The .partial.mp4 name keeps an interrupted output from looking complete. A failed pair stays in the queue, so inspect its log and input before rerunning the script. A successful output is skipped on later runs. This is deliberately a directory queue, not a job service; it is enough when one instance is processing one prepared batch.
Finish the session, not just the inference
A queued self-test took 15.9 minutes in total. About 11 minutes of that was setup, and the session cost $0.07. That is why we prepare a batch before starting the meter. Once the environment is ready, each additional queued job uses the setup we have already paid for.
We copy results back, back them up, and verify file counts and sizes before destroying the instance. In a separate backup check, our local and cloud copies matched at 5,342 files and 1.379 GiB. A transfer command completing is less convincing than checking what arrived.
This lesson came from a mistake: we once left an instance stopped for days. Its 90 GB disk continued billing at about $0.0167/h, or about $0.40/day. The balance went negative, the stopped instance could not be started, and we never downloaded the trained result. We now download first and destroy the instance afterward. For longer sessions we also set a cost cap and a detached deadline watchdog that retries destroy if the controlling process crashes.
The CLI’s destroy command asks for confirmation. After destroying, verify the instance list is empty. We encountered an old, deprecated v0 listing that returned an error resembling “0 instances,” so an error is not proof of cleanup; the v1 instances endpoint at https://console.vast.ai/api/v1/instances/ is another way to verify state using a bearer token kept out of logs.
What it cost us
The RTX 3090 offer was $0.27/h with free inbound traffic. Environment setup took about 9 minutes. Our 720×1280 clip with 8 seconds of TTS audio took 238 seconds to process and had a recorded cost of about $0.09. A queued self-test took 15.9 minutes overall, including 11 minutes of setup, and cost $0.07. A forgotten stopped instance later cost about $0.0167/h for its 90 GB disk until the balance went negative.
For us, the useful pattern is to prepare inputs, verify one LatentSync invocation, drain a simple queue, check the downloaded outputs, and destroy the rental. The GPU rate makes the experiment accessible; finishing that last step determines what it actually costs.
Top comments (0)