DEV Community

Sevan Safarian
Sevan Safarian

Posted on Originally published at fring.sevan.zone AI-assisted

On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32

I run the AI for a small side project, Fring, a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do.

It wasn't. The model was loaded in fp16, because "use half precision to save VRAM" is the default advice everywhere. On the P40 (GP102, compute capability 6.1), fp16 runs at roughly 1/64 of fp32 throughput. Same inputs, same seed:

s / diffusion step VRAM 20-step render
fp16 63.8 6.6 GB 1241 s
fp32 4.7 9.8 GB 85 s

The images were practically identical (mean difference 0.77/255). The card never needed the VRAM savings: the model used 6.4 GB out of 23.

Two other traps from the same month

  • The Dockerfile installed a CPU-only torch wheel, inherited from an older 4 GB card. nvidia-smi saw the GPU while torch.cuda.is_available() returned False, and renders took hours. Check inside the container, not on the host.
  • An unpinned mediapipe upgrade removed mediapipe.solutions. The health check only tested the import, so it kept reporting pose estimation as healthy while the service silently fell back to a centred paste.

What I changed

The service now picks fp16 only when the GPU actually accelerates it (Volta and newer, or P100), or when fp32 would not fit in VRAM. An environment variable can force either.

The general lesson for old datacenter cards: measure before applying the usual advice.

Originally published on Fring's guides.

Top comments (0)