DEV Community

orca_forge
orca_forge

Posted on Originally published at forge.workstyle.tech

When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM

📝 Originally published (in Japanese) at forge.workstyle.tech.

On the same machine (1 GPU, 12GB VRAM), two Claude Code sessions were running separate tasks concurrently.

  • This one: Generating lip-sync caricature videos using InfiniteTalk (10–12 minutes per video)
  • The other: Estimating hand poses from real-life video footage (WiLoR)

One day, the lip-sync task for 6 videos was still pending after an hour and a minute of waiting.

Looks Normal on the Surface

I checked nvidia-smi.

Memory Usage  Near Limit
Utilization   100%
Enter fullscreen mode Exit fullscreen mode

The GPU is running at full capacity. It looks like the slowness is due to heavy processing.

I tried to break it down by process.

nvidia-smi --query-compute-apps=pid,used_memory
→ used_memory: [N/A]
Enter fullscreen mode Exit fullscreen mode

In this environment (WSL2), per-process usage isn't visible. I have no idea who is using how much.

It Was Happening on the Other Side Too

In the other session, I asked about the speed of hand estimation.

Normal   5ms/frame    VRAM 2.6GB   Utilization 51%   Detection present
Conflicted 431,000ms/frame  VRAM 11.9GB  Utilization 100%  Zero detections
Enter fullscreen mode Exit fullscreen mode

Seven minutes per frame. And zero detections. The result was that when VRAM was running low, the hand estimation system was "successfully" returning the outcome that no hands were found. No errors were thrown. In 40 minutes, it only processed two frames, and both were empty.

Our lip-sync generation was also suffering, with VAE decode times dropping to 758 seconds per iteration. Under normal conditions, processing around 300 frames takes 10 to 12 minutes.

For both tasks, the GPU indicated it was running at "full load" the entire time it was exhausted, while virtually no progress was being made.

Comparison between normal and contested states. Hand pose estimation went from 5ms to 431,000ms per frame, while VRAM usage and GPU utilization were higher during contention

Exhaustion Looks Deceptively Healthy

Normally, when things run slow, you check resource utilization. If it's low, something is bottlenecked; if it's high, it's just chewing through a heavy workload.

Running out of VRAM is the exact opposite. Memory and utilization both peg at 100%, making the system look like it's firing on all cylinders. To make matters worse, one of them fails by returning a completely normal-looking "zero detections." You'd never catch it just by looking at the metrics.

How to Distinguish

Comparison was made based on time required per unit, rather than display.

  • Lip sync: 300 frames taking 10-12 minutes is normal. If it takes 1 hour, it's abnormal
  • Hand estimation: 5ms per frame is normal. 431 seconds is 80,000 times slower

By passing this data along with the start time to the other party, it's possible to determine "when it got stuck".

Agreement

We established the following rules for our sessions:

  • Before starting a long GPU task, notify the other person (content, estimated VRAM usage, and duration).
  • Wait while the other person is using the resources. Once finished, let them know "I'm free."
  • If you suspect a bottleneck, share the processing time per unit and the start time rather than just the display status.

From then on, the other person started notifying me in advance: "I'm going to use the GPU for a few minutes. VRAM will be 1-2GB, and it should take about 5-10 minutes." I also started reaching out before running batch processes of lip-sync generations.

Summary

  • VRAM exhaustion can look "healthy" because both usage percentage and memory are maxed out.
  • Depending on the inference process, failures might be returned normally as empty results.
  • To distinguish issues, don't look at displays—look at the processing time per unit.
  • If you are sharing a GPU, notify others before using it. If you're suspicious, pass the processing time and start timestamp.

Top comments (0)