📝 Originally published (in Japanese) at forge.workstyle.tech.
On the same machine (1 GPU, 12GB VRAM), two Claude Code sessions were running separate tasks concurrently.
- This one: Generating lip-sync caricature videos using InfiniteTalk (10–12 minutes per video)
- The other: Estimating hand poses from real-life video footage (WiLoR)
One day, the lip-sync task for 6 videos was still pending after an hour and a minute of waiting.
Looks Normal on the Surface
I checked nvidia-smi.
Memory Usage Near Limit
Utilization 100%
The GPU is running at full capacity. It looks like the slowness is due to heavy processing.
I tried to break it down by process.
nvidia-smi --query-compute-apps=pid,used_memory
→ used_memory: [N/A]
In this environment (WSL2), per-process usage isn't visible. I have no idea who is using how much.
It Was Happening on the Other Side Too
In the other session, I asked about the speed of hand estimation.
Normal 5ms/frame VRAM 2.6GB Utilization 51% Detection present
Conflicted 431,000ms/frame VRAM 11.9GB Utilization 100% Zero detections
Seven minutes per frame. And zero detections. The result was that when VRAM was running low, the hand estimation system was "successfully" returning the outcome that no hands were found. No errors were thrown. In 40 minutes, it only processed two frames, and both were empty.
Our lip-sync generation was also suffering, with VAE decode times dropping to 758 seconds per iteration. Under normal conditions, processing around 300 frames takes 10 to 12 minutes.
For both tasks, the GPU indicated it was running at "full load" the entire time it was exhausted, while virtually no progress was being made.
Exhaustion Looks Deceptively Healthy
Normally, when things run slow, you check resource utilization. If it's low, something is bottlenecked; if it's high, it's just chewing through a heavy workload.
Running out of VRAM is the exact opposite. Memory and utilization both peg at 100%, making the system look like it's firing on all cylinders. To make matters worse, one of them fails by returning a completely normal-looking "zero detections." You'd never catch it just by looking at the metrics.
How to Distinguish
Comparison was made based on time required per unit, rather than display.
- Lip sync: 300 frames taking 10-12 minutes is normal. If it takes 1 hour, it's abnormal
- Hand estimation: 5ms per frame is normal. 431 seconds is 80,000 times slower
By passing this data along with the start time to the other party, it's possible to determine "when it got stuck".
Agreement
We established the following rules for our sessions:
- Before starting a long GPU task, notify the other person (content, estimated VRAM usage, and duration).
- Wait while the other person is using the resources. Once finished, let them know "I'm free."
- If you suspect a bottleneck, share the processing time per unit and the start time rather than just the display status.
From then on, the other person started notifying me in advance: "I'm going to use the GPU for a few minutes. VRAM will be 1-2GB, and it should take about 5-10 minutes." I also started reaching out before running batch processes of lip-sync generations.
Summary
- VRAM exhaustion can look "healthy" because both usage percentage and memory are maxed out.
- Depending on the inference process, failures might be returned normally as empty results.
- To distinguish issues, don't look at displays—look at the processing time per unit.
- If you are sharing a GPU, notify others before using it. If you're suspicious, pass the processing time and start timestamp.

Top comments (0)