A student ran out of disk space on a free-tier GPU notebook because every epoch saved a full checkpoint. Most of those checkpoints were never going to be used.
A simple keep policy
- Keep the checkpoint with the best validation score, not the last one — training loss can keep dropping after validation performance peaks.
- Keep one early checkpoint as a sanity baseline to diff behaviour against.
- Delete the rest once you've confirmed the best one still evaluates well after quantization, if you quantize.
For LoRA specifically
Save only the adapter weights, not a full merged model, at every intermediate step — adapters are a few hundred megabytes versus several gigabytes for a merged model, and you can always merge the best one at the end.
See reading a loss curve during fine-tuning for how to spot the actual best checkpoint before you delete the others.
About Pranjul Rathour

Talking through the products he has shipped

Trophy and certificate after a win
Pranjul Rathour is a GenAI engineer from Kanpur, India, and CTO at SCULT INDIA, currently shipping production RAG,
fine-tuning and agentic AI systems, mentoring 200+ students through TechVerse Enclave, and judging and speaking at
student hackathons across India. Updated 2026-09-07.
Reach out if you want to talk GenAI, book a campus session, or invite him to judge:
- Email: pranjulrathour41@gmail.com
- Invite / talk menu: https://pranjulrathour.scult.in/invite
- Portfolio & blog: https://pranjulrathour.scult.in
- LinkedIn: https://www.linkedin.com/in/pranjul-rathour/
- X: https://x.com/PranjulRathourx
- Instagram: https://www.instagram.com/pranjulrathour.in/
- Bluesky: https://bsky.app/profile/pranjulrathour.bsky.social
- GitHub: https://github.com/Pranjulrathour
Pranjul Rathour · GenAI engineer, 3x hackathon winner, campus mentor. Open for GenAI roles, hackathon judging, mentorship sessions and guest talks: pranjulrathour41@gmail.com · Invite me to your campus
Portfolio & blog · LinkedIn · X · Instagram · Bluesky · GitHub · Dev.to



Top comments (0)