Disclaimer & Notice: This project is strictly an educational experiment conducted for self-learning purposes to understand the mechanics of language model fine-tuning and deployment. The resulting model weights must never be used as actual medical, wilderness, or life-safety advice.
Fine-tuning language models often feels wrapped in an intimidating vocabulary: weights, tokens, FP32 precision, loss curves, and quantization. Most guides throw these terms around without explaining what is actually happening under the hood.
I was curious about the actual engineering behind how models adapt, so I ran an end-to-end experiment using Google Colab. I took a compact open-source model, trained it on a specialized dataset to act as a concise emergency survival guide, and published the final artifact to the cloud.
Here is an honest, step-by-step breakdown of how it works—with every AI term explained right as it appears.
🎯 1. The Core Objective: What Are We Actually Doing?
To understand fine-tuning, it helps to distinguish the two main phases of building an AI model:
- Pre-training: Training a neural network from scratch on hundreds of terabytes of internet text so it learns grammar, facts, and reasoning. Think of it like someone spending 15 years in school reading an entire library.
- Fine-Tuning: Taking that already educated model and running additional training on a small, specialized dataset to teach it a specific role, tone, or format. It is like giving that graduate a focused 2-day workshop on emergency incident response.
- Supervised Fine-Tuning (SFT): The specific training style used here. "Supervised" simply means the training data contains clear pairs: a prompt paired with an ideal target answer. The model guesses an answer, compares its guess against the target, and adjusts itself.
The goal here was to train an emergency survival assistant—taking a model that naturally gives long, chatty responses and steering it toward direct, prioritized, life-saving steps.
🧠 2. The Base Model & Parameters
For this experiment, I selected Qwen/Qwen2.5-0.5B-Instruct.
- Base Model / Weights: The neural network file itself. It is essentially a massive collection of numbers (called weights) that determine how input text gets converted into output predictions.
- Parameters (0.5B): The adjustable numerical values inside the model. 0.5B stands for 0.5 billion (500 million) parameters. In a landscape where flagship models have 70B to 400B+ parameters, 0.5B is tiny. That small footprint is an advantage for experimentation: it loads quickly and fits entirely within the free GPU limits of Google Colab.
- Instruct Model: A model that has already been taught how to carry on a conversational back-and-forth, rather than just raw text autocomplete.
🔤 3. Preparing the Data: Tokens & Chat Templates
Computers cannot read plain text; they only do math with numbers. Before training, words must be translated into numerical sequences.
Tokens and Tokenizers
-
Token: A chunk of text—a word, syllable, or punctuation mark—represented as an integer. For example, the word
"freeze"might be mapped to ID4512. - Tokenizer: The software tool that slices raw text strings into arrays of token IDs before sending them to the model (and converts them back into human words when the model answers).
Chat Templates
When you have a multi-turn conversation, how does the model know who said what?
-
Special Tokens: Invisible control tags placed around text. Different model creators use different tags. Qwen uses
<|im_start|>and<|im_end|>to mark boundaries. - Chat Template: A formatting rule that wraps messages with those exact special tokens so the model recognizes system instructions, user prompts, and assistant replies without confusion:
messages = [
{"role": "system", "content": "You are an emergency survival assistant. Provide direct, step-by-step, life-saving instructions without fluff."},
{"role": "user", "content": inst},
{"role": "assistant", "content": resp}
]
# Applies Qwen's specific boundary tags:
formatted_text = tokenizer.apply_chat_template(messages, tokenize=False)
Labels
- Labels: The ground-truth answers in token form. During training, the framework compares the model's predicted token IDs against the labels to calculate where it made a mistake.
tokenized = tokenizer(formatted_texts, truncation=True, max_length=512, padding=False)
tokenized["labels"] = tokenized["input_ids"].copy()
⚙️ 4. The Training Loop & Hyperparameters
The training was orchestrated using Hugging Face's trl (Transformer Reinforcement Learning) library and its SFTTrainer.
This step often introduces several confusing configuration settings (known as hyperparameters—settings chosen by the engineer before training begins):
training_args = TrainingArguments(
output_dir="./survival_model_out",
per_device_train_batch_size=2,
gradient_accumulation_steps=2,
learning_rate=3e-5,
num_train_epochs=5,
logging_steps=10,
fp16=False, # FP32 precision
)
Here is what those parameters actually control:
- Precision & FP32 (32-bit Floating Point): In computing, decimal numbers are called "floating-point numbers." FP32 means each numerical weight uses 32 bits of computer memory. It gives maximum mathematical accuracy and prevents numerical instability during training. Because each number takes 4 bytes, 500 million parameters consume about 2 GB of memory in raw FP32.
- Epochs (5): One epoch equals one full pass through the entire dataset. Running 5 epochs means the model saw and practiced every survival scenario in the file 5 times.
- Batch Size (2): The number of training examples processed at the exact same moment on the GPU.
- Gradient Accumulation Steps (2): If your GPU runs low on memory, you cannot use a large batch size. Gradient accumulation lets the model calculate errors across smaller micro-batches (here, 2 batches of 2) and combine the updates, effectively behaving like a batch size of 4 without crashing the GPU.
-
Learning Rate (
3e-5/0.00003): The magnitude of adjustment applied to the model's weights after each error. If the learning rate is too high, the model overreacts and forgets what it already knew (catastrophic forgetting). If it is too low, the model learns too slowly to make progress.
💾 5. Saving & Storing the Model: Safetensors & Hugging Face
Once training finished, the model wrote its updated weights to disk inside a survival_model_final folder.
-
.safetensors: The file format storing the model's trained numbers. Historically, models used Python'spickleformat (.binor.pt), which could accidentally run malicious code when loaded..safetensorsis a modern, restricted format that is faster to load and safe against code execution exploits.
Why Not Push to GitHub?
The exported weights totaled around 1 GB.
- Standard Git / GitHub is designed for source code (small text files) and restricts any single file over 100 MB.
- The standard registry for storing large machine learning models is the Hugging Face Hub (essentially GitHub, but optimized for multi-gigabyte neural network files).
Using the huggingface_hub Python library, the weights and tokenizer were published directly from Colab:
👉 sahil2605/survival-qwen-0.5b
Anyone can now pull and test this checkpoint with standard Python scripts:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "sahil2605/survival-qwen-0.5b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
📊 6. What Was Achieved
- A Working End-to-End Pipeline: Successfully took raw conversational text records, tokenized them with native boundary delimiters, trained the network parameters, and hosted the final checkpoint on an open registry.
- Distinct Behavioral Shift: When tested with emergency questions (e.g., being stranded in freezing weather), the base model typically outputs polite introductory filler. The fine-tuned version immediately starts with prioritized, numbered triage steps (shelter, ground insulation, signaling).
- Hands-On Clarity: Running every phase demonstrated that fine-tuning is not mysterious intuition—it is an automated pipeline of text-to-number encoding, gradient optimization, and binary checkpoint storage.
🗺️ 7. What’s Next: Planned Experiments
This training run yielded a functional baseline checkpoint. The next iterations will test:
-
Quantization (4-Bit / GGUF): Converting the 32-bit floating-point weights down to 4-bit integers using tools like
llama.cpp. This compresses the file from ~1 GB down to ~300 MB, allowing it to run offline on budget laptops or mobile devices without needing a GPU. - Catastrophic Forgetting Evaluation: Running comparative benchmarks against the original base model to verify whether specializing in survival instructions harmed its general reasoning or grammar.
- Safety Boundaries & Guardrails: Adding refusal behaviors so the model safely declines clearly hazardous, nonsensical, or malicious requests instead of generating plausible-sounding hallucinations.
This write-up documents an ongoing personal learning path in machine learning engineering. If you have questions about running small LLM experiments or suggestions for evaluation, feel free to connect and share your thoughts!
Top comments (0)