DEV Community

Francisco Booth
Francisco Booth

Posted on

My GPU Training Job Ran for 20 Hours and Produced Nothing. There Was No Error Message.

I was fine-tuning the Exolio AI classifier that detects machine-generated texts. The classifier can be seen at huggingface.co/FranciscoBooth1/exolio-ai-detector.

Between January and April 2026, I ran dozens of training jobs using Vast.ai. Cost of runs: approximately £1,000.

All of the eventually identified failed runs had one thing in common: none of them gave me an obvious sign that they had failed.


The run that ran all night and gave no output

For Exolio v15, I was switching from DeBERTa-v3-large to ModernBERT. I rented a machine, uploaded the training script, installed dependencies, and initiated the run.

Everything in the terminal looked exactly as expected:


Uploading training script...
Installing Python dependencies (this may take several minutes)...
[Dependencies installed successfully]
Starting training in tmux session 'ssh_tmux'...
Training started successfully! Instance will auto-destroy when complete.


The script returned to my local terminal and training started running in tmux on the remote machine. There was no way to monitor what happened inside that tmux session. Everything looked good. I went to sleep.

The next day I tried to SSH back in:


ssh: connect to host [instance IP] port 55157: Connection refused


The instance was gone. The log files disappeared along with it. I checked my HuggingFace repository to see if there were any results:


Last modified: 2026-04-12 01:20:33+00:00
.gitattributes
config.json
deberta_3class_latest.zip
model.safetensors
tokenizer.json
training_args.bin


There was no modernbert_latest.zip. The last modification time was April 12 which was the previous DeBERTa run. No new files had been added.

That was the only sign that something had gone wrong. Not an error message. Not a bad benchmark score. Absence of new files.

I could not recover the logs the instance auto-destroyed and took everything with it so I cannot identify the root cause. Twenty hours had passed, the instance was gone, and no new model had appeared in my HuggingFace repository.

I ran a training job, nothing obviously failed, and after twenty hours the infrastructure disappeared and there was no artifact left. I was not able to tell exactly why.


Another failure: training on CPU for hours without knowing

Earlier in the same period I rented what seemed to be a good machine and started fine-tuning DeBERTa. The loss function was printing. The iteration counter was counting. Everything was going as expected.

Except the speed was wrong. I saw approximately 24 seconds per iteration. On working instances the same workload had been around 0.4 seconds per iteration.

I ran nvidia-smi. GPU utilisation: 0%.

My script resolved the device to CPU because torch.cuda.is_available() returned False no exception thrown, just CPU. The machine had CUDA installed but my PyTorch environment could not see the GPU. Training continued on CPU with no errors while I was paying full GPU rates.

I stopped the run. Rented a different machine. Restarted.


The pattern

In both cases there was no error message that I could act on. In the first case there was no way to know anything had gone wrong until I tried to SSH back in. In the second case I noticed a strange iteration speed and stopped the job.

In the second case the problem was visible before python train.py. torch.cuda.is_available() returning False is visible before the job starts. GPU utilisation at 0% is visible within seconds of launch.

Additionally and not as an explanation of the overnight run, since I had no logs there are other silent failure modes that are also detectable before the training job starts. At least some of the conditions surrounding these failures were detectable before python train.py if I had known to check them:

Whether PyTorch could see the GPU: torch.cuda.is_available() returns False before the job starts
Whether the HuggingFace cache was configured to use a potentially ephemeral path: an environment variable check
Whether the checkpoint output directory was under a known instance-local prefix: a path check

None of these checks were complicated. I just did not perform them systematically on each new machine.


What I built

After enough of these I decided to build a pre-flight checker. One command before each training job:


pip install computefence && computefence doctor


Or without installing the package:


uvx computefence doctor


Real output from uvx computefence doctor on my local Mac — on a real GPU pod the PyTorch and volume checks run against actual hardware:


ComputeFence v0.2.6 — Pre-flight diagnostic
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

5 WARNINGS · 0 BLOCKERS · 1 PASSED

Environment
✓ Python 3.14.3
⚠ PyTorch not found — skipping GPU checks
Fix: pip install torch torchvision torchaudio

Storage
⚠ HF_HOME is not set. HuggingFace may cache to instance-local storage
that is not persistent across pod termination.
Fix: Set HF_HOME to a persistent volume path before training
⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
Fix: Mount a network volume at /workspace before launching your job
⚠ Root disk (/) — 12.1 GB free of 460.4 GB (below 20 GB)
Fix: Free up disk space or move checkpoints to a larger volume

Dataset
⚠ No dataset path provided — skipping dataset checks

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
5 warning(s) found. Review before launching.


Checks:

1) GPU and CUDA visibility - fails if torch.cuda.is_available() returns False; reports relevant CUDA and PyTorch version information and warns on detected version skew

2) HuggingFace cache path — warns if HF_HOME is not set or points to a path known to be instance-local on supported providers.

3) Accelerate configuration — warns if there is an inconsistency between the configured number of processes and the GPUs visible in the environment.

4) Disk headroom — warns when less than 20 GB is available; blocks if less than 5 GB is available.

5) Checkpoint output directory — warns if the path is under known instance-local prefixes that may not survive pod termination.

6) Dataset — optionally checks for duplicates and conflicting labels with --dataset.

Checks that it does not do: training script correctness, model architecture, learning rate safety, runtime behaviour during training, or GPU utilisation during training. ComputeFence is a pre-flight check. It does not monitor the running job.

A passing check does not guarantee that the training job will succeed.

The pre-flight check took under 30 seconds on my test pod.


The ask

ComputeFence is free, open source, and MIT licensed.

If you run training jobs on Vast.ai, RunPod, or any other rented GPU, try uvx computefence doctor before your next run. If it finds something interesting, paste the output in the comments below.

GitHub: github.com/Francisco-Booth/ComputeFence
Install: pip install computefence

Top comments (0)