π Technical Briefing: This tutorial is part of our deep-dive series on Agentic Workflows at Gate of AI. For the full technical breakdown, interactive code sandbox, and the native Arabic translation, visit the original article here.
Tutorial
Intermediate
Build an LLM Fine-Tuning Data Validation Pipeline
Use a three-stage fail-fast workflow to check fine-tuning data, run a small validation trial, detect runtime anomalies, and preserve a separate validation set for measuring model behavior.
Why Fine-Tuning Validation Must Come First
Fine-tuning an LLM is a pipeline rather than a single training command. Before meaningful computation begins, a team must establish that the configuration is syntactically valid, the data has the expected structure, and a short run can complete without obvious failures. The verified FT-Dojo research describes this as progressive validation: inexpensive checks run first, followed by increasingly costly checks. Configurations that fail a stage are rejected immediately instead of consuming resources in a full run.
This tutorial implements a small, provider-neutral version of that idea. It validates a JSON Lines dataset, checks the schema of every example, verifies paths and configuration values, performs a reduced mini-run over a sample, and reports runtime anomalies such as an empty dataset or non-finite loss values. It does not submit data to a commercial training API. That boundary is deliberate: provider-specific training formats and model eligibility are outside the verified context.
The workflow also separates training data from a validation set. During fine-tuning, the training loss is minimized on the training examples. Afterward, validation loss is computed on data that was not used to update the trainable parameters. This separation gives you evidence about generalization rather than merely showing that the model can fit the examples it saw.
For parameter-efficient fine-tuning, LoRA keeps most pretrained weights frozen and introduces a low-rank decomposition that is trained instead. The verified hyperparameter study highlights LoRA rank, scaling alpha, dropout, and learning rate as important variables. The correct engineering response is not to copy one value blindly, but to validate each candidate configuration and compare it on the same validation set.
Prerequisites
- Python 3.10 or newer.
- A JSONL dataset containing conversational examples.
- A separate JSONL validation set that does not overlap with the training examples.
- Basic knowledge of JSON, command-line execution, and Python virtual environments.
The validation utility below uses only Pythonβs standard library. This keeps the fail-fast stage easy to run before installing a training stack. A later training implementation can use the Hugging Face Transformers API, which is the API used for model handling, training, and validation in the verified hyperparameter study.
Step 1: Create the Project and Dataset Contract
Create a project directory and a virtual environment. The pipeline will treat each non-empty JSONL line as one example. Each example must contain a messages array with at least two objects. Every message needs a supported role and non-empty string content. The final message must be the target assistant response.
mkdir llm-finetuning-validation
cd llm-finetuning-validation
python -m venv .venv
# Linux or macOS
source .venv/bin/activate
# Windows PowerShell
# .venv\Scripts\Activate.ps1
mkdir -p data/reports src
touch src/__init__.py
Create data/train.jsonl with examples such as these:
{"messages":[{"role":"system","content":"Answer clearly."},{"role":"user","content":"What is a validation set?"},{"role":"assistant","content":"A validation set is held-out data used to measure a model during or after training."}]}
{"messages":[{"role":"system","content":"Answer clearly."},{"role":"user","content":"Why check the schema first?"},{"role":"assistant","content":"Schema checks catch malformed examples before an expensive training run begins."}]}
Create data/validation.jsonl separately. Do not copy the same records into both files. The verified research evaluates a fine-tuned model on a validation set, so the set must remain available for that purpose rather than being absorbed into training.
Step 2: Implement Static and Schema Validation
Static validation is the first and least expensive stage. It checks that the input path exists, every non-empty line is valid JSON, the record is an object, and the conversational structure satisfies the dataset contract. It also checks configuration values for LoRA rank, scaling alpha, dropout, learning rate, batch size, and mini-run length. These fields correspond to hyperparameters discussed in the verified context; the script validates their shape but does not claim that any particular value is optimal.
Create src/validate_pipeline.py:
from __future__ import annotations
import argparse
import json
import math
from pathlib import Path
from typing import Any
ROLES = {"system", "user", "assistant"}
def issue(stage: str, message: str, line: int | None = None) -> dict[str, Any]:
result = {"stage": stage, "message": message}
if line is not None:
result["line"] = line
return result
def validate_record(record: Any, line: int) -> list[dict[str, Any]]:
errors: list[dict[str, Any]] = []
if not isinstance(record, dict):
return [issue("schema", "record must be a JSON object", line)]
messages = record.get("messages")
if not isinstance(messages, list) or len(messages) < 2:
return [issue("schema", "messages must contain at least two items", line)]
roles: list[str] = []
for index, message in enumerate(messages):
if not isinstance(message, dict):
errors.append(issue("schema", f"message {index} must be an object", line))
continue
role = message.get("role")
content = message.get("content")
if role not in ROLES:
errors.append(issue("schema", f"unsupported role at message {index}", line))
else:
roles.append(role)
if not isinstance(content, str) or not content.strip():
errors.append(issue("schema", f"message {index} needs non-empty content", line))
if "user" not in roles:
errors.append(issue("schema", "record needs a user message", line))
if roles and roles[-1] != "assistant":
errors.append(issue("schema", "final message must be the assistant target", line))
return errors
def read_records(path: Path, limit: int | None = None) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
if not path.is_file():
return [], [issue("static", f"missing dataset path: {path}")]
records: list[dict[str, Any]] = []
errors: list[dict[str, Any]] = []
with path.open(encoding="utf-8") as handle:
for line_number, raw in enumerate(handle, 1):
if not raw.strip():
errors.append(issue("format", "blank lines are not allowed", line_number))
continue
try:
value = json.loads(raw)
except json.JSONDecodeError as error:
errors.append(issue("format", error.msg, line_number))
continue
record_errors = validate_record(value, line_number)
if record_errors:
errors.extend(record_errors)
elif limit is None or len(records) < limit:
records.append(value)
return records, errors
def validate_config(config: dict[str, Any]) -> list[dict[str, Any]]:
errors: list[dict[str, Any]] = []
positive = ("rank", "alpha", "learning_rate", "batch_size", "mini_run_records")
for name in positive:
value = config.get(name)
if not isinstance(value, (int, float)) or isinstance(value, bool) or value <= 0:
errors.append(issue("static", f"{name} must be a positive number"))
dropout = config.get("dropout")
if not isinstance(dropout, (int, float)) or not 0 <= dropout < 1:
errors.append(issue("static", "dropout must be at least 0 and below 1"))
return errors
def mini_run(records: list[dict[str, Any]]) -> list[dict[str, Any]]:
errors: list[dict[str, Any]] = []
if not records:
return [issue("runtime", "mini-run received an empty dataset")]
for index, record in enumerate(records, 1):
messages = record["messages"]
total_chars = sum(len(message["content"]) for message in messages)
if total_chars == 0:
errors.append(issue("runtime", "example produced zero content", index))
simulated_loss = 1.0 / (index + 1)
if not math.isfinite(simulated_loss):
errors.append(issue("runtime", "loss is not finite", index))
return errors
def main() -> None:
parser = argparse.ArgumentParser(description="Run fail-fast LLM fine-tuning validation.")
parser.add_argument("--train", required=True, type=Path)
parser.add_argument("--validation", required=True, type=Path)
parser.add_argument("--report", required=True, type=Path)
parser.add_argument("--rank", type=int, default=8)
parser.add_argument("--alpha", type=float, default=16)
parser.add_argument("--dropout", type=float, default=0.05)
parser.add_argument("--learning-rate", type=float, default=0.0001)
parser.add_argument("--batch-size", type=int, default=4)
parser.add_argument("--mini-run-records", type=int, default=8)
args = parser.parse_args()
config = {"rank": args.rank, "alpha": args.alpha, "dropout": args.dropout,
"learning_rate": args.learning_rate, "batch_size": args.batch_size,
"mini_run_records": args.mini_run_records}
errors = validate_config(config)
train, train_errors = read_records(args.train)
validation, validation_errors = read_records(args.validation)
errors.extend(train_errors)
errors.extend(validation_errors)
if not errors:
errors.extend(mini_run(train[:args.mini_run_records]))
errors.extend(mini_run(validation[:args.mini_run_records]))
report = {"approved_for_full_run": not errors, "config": config,
"train_records_checked": len(train),
"validation_records_checked": len(validation),
"errors": errors}
args.report.parent.mkdir(parents=True, exist_ok=True)
args.report.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
print(json.dumps(report, indent=2))
if errors:
raise SystemExit(1)
if __name__ == "__main__":
main()
Step 3: Run the Reduced Mini-Run
Run the validator before starting full fine-tuning:
python -m src.validate_pipeline \
--train data/train.jsonl \
--validation data/validation.jsonl \
--report data/reports/validation_report.json
The script stops before the mini-run if static, schema, or format checks fail. This is the fail-fast behavior described in the verified FT-Dojo work. A successful mini-run here is not evidence that the model will perform well. It only shows that the selected data and configuration pass inexpensive structural and short-run checks.
The default batch size of 4 is included because it was used in the verified hyperparameter study, not because it is universally correct. Hardware capacity, sequence length, model size, and training implementation can require different settings. Likewise, the example LoRA values are validation inputs, not recommended optima. Compare candidates using the same validation set and record the configuration for each run.
Step 4: Interpret Validation Loss and LoRA Experiments
Once static checks and the mini-run pass, a real training implementation can fine-tune the model and compute validation loss. The verified study uses the Hugging Face Transformers API for model handling, training, and validation, and adapts LoRA trainable parameters while keeping most pretrained weights frozen. The study investigates rank, scaling alpha, dropout, and learning rate because these settings affect downstream performance.
Keep the experiment table simple and reproducible. For every candidate, record the dataset revision, model identifier, LoRA rank, alpha, dropout, learning rate, batch size, number of training steps or epochs, training loss, and validation loss. Do not compare runs whose training data or validation data changed at the same time. A lower training loss alone is insufficient: the validation result is needed to assess whether the adapted model generalizes to held-out examples.
The verified research also describes black-box hyperparameter optimization with NOMAD. That is an experimental optimization approach, not a reason to skip validation. Every proposed configuration should still pass the cheap checks and the reduced run before expensive training is authorized.
Step 5: Add Runtime Sanity Checks
Runtime checks should look for failures that static validation cannot detect. The verified FT-Dojo description specifically identifies exploding loss, empty datasets, and invalid gradients as examples of runtime anomalies. The standard-library utility above checks for an empty mini-run and non-finite simulated values; a training integration should connect the same decision point to the actual loss and gradient values produced by the training framework.
Stop the run when a sanity check fails. Do not interpret a completed process as a successful experiment if the dataset was empty, the output structure was incompatible, the loss became non-finite, or gradients were invalid. The purpose of the staged design is to reject broken configurations early and return targeted diagnostics while little computation has been spent.
Step 6: Execute Full Fine-Tuning Only After Approval
A report with approved_for_full_run: true means only that the local gates passed. It does not guarantee quality, safety, or a useful final model. Before full training, inspect representative examples and confirm that the validation set is separate. Then run the chosen fine-tuning implementation and evaluate the resulting model on the unchanged validation set.
The verified experimental setup used four NVIDIA A100 GPUs with 80 GB memory, AdamW, and batch size 4. Treat those details as the conditions of that study. They are useful reference points for reproducing that experiment, but they are not a universal infrastructure requirement for every LLM or LoRA run.
Common mistake: changing the validation examples after every experiment makes scores impossible to compare. Freeze the validation set while comparing configurations, and change it only through a documented dataset revision.
Key Takeaways
- Use static and schema validation before any expensive computation.
- Check data format and run a reduced mini-run before full training.
- Detect empty data, non-finite loss, and invalid gradients during runtime.
- Keep a separate validation set and compute validation loss after fine-tuning.
- Treat LoRA rank, alpha, dropout, and learning rate as experiment variables.
- Use research-specific hardware and optimizer settings as documented conditions, not universal defaults.
Sources
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents.
- Hyperparameter Optimization for Large Language Model Instruction-Tuning.
- Hugging Face: Fine-tune a SmolLM on synthetic data.
Editorially reviewed by the Gate of AI Editorial & Engineering Teams, GateOfAI, LLC, Delaware, USA.
Top comments (0)