DEV Community

Cover image for Validate your dataset before you burn a GPU hour: a checklist
PRANJUL RATHOUR
PRANJUL RATHOUR

Posted on Originally published at pranjulrathour.scult.in

Validate your dataset before you burn a GPU hour: a checklist

Most fine-tuning failures I have seen were dataset failures: malformed records, duplicated examples, a prompt template mismatch. FineTune Studio validates on upload precisely because a clear error before training is worth more than any dashboard during it. Here is the checklist it implements, so you can run it yourself.

Structure

  • Every record parses, and every record has the same shape. One stray field name breaks tokenisation silently.
  • Roles alternate correctly and every conversation ends with an assistant turn — that is the turn you train on.
  • No empty assistant messages. They teach the model that silence is an answer.

Content

  • Exact and near-duplicate examples removed. Duplicates make the loss look great and the model brittle.
  • No overlap between training and evaluation splits — check by hashing normalised text, not by hoping.
  • Length distribution inspected: a handful of 8,000-token examples will set your sequence length and your memory bill. Truncate or drop them deliberately.
  • Label or intent balance checked for classification-style data; a 95/5 split trains a model that always says the 95.

Rendering

Render twenty random examples through the exact chat template you will train with, decode them back to text, and read them. This catches the mistakes scripts cannot: system prompts duplicated on every turn, a stray instruction inside an answer, examples in the wrong language. Formats and templates are covered in fine-tuning datasets without the confusion.

Provenance

Know where every example came from and whether you are allowed to train on it. Client data needs consent; scraped data needs a licence check. Write the source into the dataset card. A model you cannot explain the training data for is a model you cannot ship to a client.

Then, and only then

Start the run and watch the loss curve. A validated dataset turns a mysterious training failure into a hyperparameter question, and hyperparameter questions have answers.

About Pranjul Rathour

Pranjul Rathour in a checked shirt inside a packed college auditorium
In a packed college auditorium

Pranjul Rathour in a suit and tie with a lanyard at a formal campus event
At a formal campus event

Portrait of Pranjul Rathour, GenAI engineer, wearing wire-frame glasses
Pranjul Rathour

Pranjul Rathour presenting on stage in a blue polo, with his Annapurna demo video on the screen behind him
Presenting Annapurna on stage

Pranjul Rathour giving a talk titled 'How and what I do', with demo videos of his products Vaidya and Annapurna on screen
Talking through the products he has shipped

Pranjul Rathour is a GenAI engineer from Kanpur, India, and CTO at SCULT INDIA, currently shipping production RAG,
fine-tuning and agentic AI systems, mentoring 200+ students through TechVerse Enclave, and judging and speaking at
student hackathons across India. Updated 2026-09-06.

Reach out if you want to talk GenAI, book a campus session, or invite him to judge:


Pranjul Rathour · GenAI engineer, 3x hackathon winner, campus mentor. Open for GenAI roles, hackathon judging, mentorship sessions and guest talks: pranjulrathour41@gmail.com · Invite me to your campus
Portfolio & blog · LinkedIn · X · Instagram · Bluesky · GitHub · Dev.to

Top comments (0)