I run the AI for a small side project, Fring, a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do.
It wasn't. The model was loaded in fp16, because "use half precision to save VRAM" is the default advice everywhere. On the P40 (GP102, compute capability 6.1), fp16 runs at roughly 1/64 of fp32 throughput. Same inputs, same seed:
| s / diffusion step | VRAM | 20-step render | |
|---|---|---|---|
| fp16 | 63.8 | 6.6 GB | 1241 s |
| fp32 | 4.7 | 9.8 GB | 85 s |
The images were practically identical (mean difference 0.77/255). The card never needed the VRAM savings: the model used 6.4 GB out of 23.
Two other traps from the same month
- The Dockerfile installed a CPU-only torch wheel, inherited from an older 4 GB card.
nvidia-smisaw the GPU whiletorch.cuda.is_available()returnedFalse, and renders took hours. Check inside the container, not on the host. - An unpinned
mediapipeupgrade removedmediapipe.solutions. The health check only tested the import, so it kept reporting pose estimation as healthy while the service silently fell back to a centred paste.
What I changed
The service now picks fp16 only when the GPU actually accelerates it (Volta and newer, or P100), or when fp32 would not fit in VRAM. An environment variable can force either.
The general lesson for old datacenter cards: measure before applying the usual advice.
Originally published on Fring's guides.
Top comments (0)