Originally published on khadim.tech.
In short: gpt-oss-20b fine-tunes with QLoRA on a single consumer GPU (Unsloth puts it at about 14 GB of VRAM). Shipping it is where the chat template matters: gpt-oss uses OpenAI's Harmony format, so the GGUF needs a Harmony template with <|return|> as a stop token, or Ollama gives gibberish or never stops.
I fine-tuned gpt-oss-20b for STEM reasoning and published it three ways: a LoRA adapter, a merged 16-bit model and GGUF files. Below are the settings and versions behind it, the export code, and the Modelfile that makes the GGUF behave in Ollama.
QLoRA settings for gpt-oss-20b
The published run, as recorded on the model card:
| Setting | Value |
|---|---|
| Base model | openai/gpt-oss-20b (21B parameters, mixture of experts) |
| Method | QLoRA, 4-bit base weights |
| LoRA rank / alpha | 32 / 64 |
| Learning rate | 1e-4 |
| Batch size | 1, with gradient accumulation 16 |
| Optimizer | AdamW 8-bit |
| Epochs | 1 |
| Data | 4,260 training and 474 evaluation examples of STEM Q&A with chain-of-thought |
| Result | train loss 1.087, eval loss 0.837 |
| Framework | Unsloth + TRL |
The full training config, with target modules and scheduler, is in configs/gpt_oss_20b.yaml in kllm, the small toolkit I use for these runs.
How much VRAM does fine-tuning gpt-oss-20b need?
QLoRA keeps the base weights in 4-bit and trains only a small adapter, so a 20B model fits on one consumer card. Unsloth's gpt-oss guide lists about 14 GB of VRAM for gpt-oss-20b with QLoRA and recommends at least 16 GB for stable runs. LoRA on unquantized BF16 weights needs about 44 GB, which is why QLoRA is the practical choice on one card.
The environment kllm is verified on:
| Component | Version |
|---|---|
| GPU | RTX 5090 (32 GB) |
| OS | Ubuntu 24.04 |
| CUDA | 12.8 toolkit (13.0 driver) |
| Python | 3.12.3 |
| PyTorch | 2.9.1+cu128 |
| Unsloth | 2026.1.4 |
| transformers | 4.57.3 |
| vLLM | 0.15.0 |
Pin transformers. vLLM 0.15.0 needs exactly transformers 4.57.3, and other versions fail with import errors. Install in this order, and force the pin last.
uv pip install -U vllm --torch-backend=cu128 # cu121 for RTX 40 series
uv pip install unsloth unsloth_zoo bitsandbytes
uv pip install --force-reinstall "transformers==4.57.3"
Export to LoRA, merged and GGUF with Unsloth
After training, the same model object gives all three formats. These are the calls kllm makes:
# 1. LoRA adapter (~61 MB): for people who already run the base model
model.save_pretrained(output_dir)
tokenizer.save_pretrained(output_dir)
# 2. Merged 16-bit model (~41 GB): for Transformers or vLLM, no adapter handling
model.save_pretrained_merged(f"{output_dir}-merged", tokenizer, save_method="merged_16bit")
# 3. GGUF for llama.cpp, Ollama and LM Studio, one call per quantization
model.save_pretrained_gguf("gguf/gpt-oss-20b-stem", tokenizer, quantization_method="q4_k_m")
quantization_method takes f16, q8_0, q4_k_m and the other llama.cpp types listed in Unsloth's GGUF docs. The GGUF files for this model came out at:
| File | Size |
|---|---|
| f16 | 41.9 GB |
| Q8_0 | 22.3 GB |
| Q4_K_M | 15.8 GB |
If the GGUF step runs out of memory, Unsloth's docs suggest lowering maximum_memory_usage from its 0.75 default to around 0.5.
The Harmony chat template is what breaks
gpt-oss doesn't use ChatML. It was trained on OpenAI's Harmony format, where every message is wrapped in special tokens and the assistant writes on separate channels:
-
analysisfor its chain of thought, which end users shouldn't see -
finalfor the answer -
commentaryfor tool calls
A final answer ends with <|return|>, which OpenAI's docs call "a valid stop token indicating that you should stop inference." Messages kept in the conversation history end with <|end|> instead, while supervised training targets should end with <|return|>. OpenAI's model card is blunt about it: the models "should only be used with the harmony format as it will not work correctly otherwise."
So the template you ship has to match the one used in training. Unsloth's docs name a mismatch as the most common reason a model gives gibberish or endless output on other platforms. And <|return|> has to be a stop token, or generation runs past the answer.
This exact question, how to template a fine-tuned gpt-oss for Ollama, is still open on OpenAI's forum.
A working Ollama Modelfile for fine-tuned gpt-oss
This is the Modelfile that ships with the GGUF repo:
FROM ./gpt-oss-20b-finetuned-q8_0.gguf
TEMPLATE """<|start|>system<|message|>You are a helpful assistant trained by OpenAI.
Reasoning: medium
# Valid channels: analysis, commentary, final. Channel must be included for every message.<|end|><|start|>user<|message|>{{ .Prompt }}<|end|><|start|>assistant"""
PARAMETER stop "<|return|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 4096
PARAMETER num_predict 2048
Setting reasoning effort
Reasoning: medium in the system message sets how much the model thinks in its analysis channel before it answers. Harmony accepts low, medium and high, and medium is the default. To change it, edit that line in the TEMPLATE and run ollama create again. Ollama shows the analysis channel as "Thinking...", so low is the setting for short, fast answers.
ollama run hf.co or a Modelfile?
With the Modelfile, the template is the one you wrote:
ollama create gpt-oss-20b-stem -f Modelfile
ollama run gpt-oss-20b-stem
Ollama can also pull straight from Hugging Face:
ollama run hf.co/khadim-hussain/gpt-oss-20b-stem-reasoning-GGUF:Q4_K_M
That route doesn't read the Modelfile in the repo. Per Hugging Face's Ollama docs, it uses the chat template stored in the GGUF's tokenizer.chat_template metadata, or a file named template in the repo. The quantization tag is case-insensitive, and without a tag Ollama picks Q4_K_M. If you publish a Harmony model, put a template file in the repo or point people to the Modelfile.
Which format people actually download
All-time Hugging Face downloads as of October 2026:
| Format | gpt-oss-20b | Qwen3-14B |
|---|---|---|
| GGUF | 513 | 274 |
| Merged 16-bit | 70 | 42 |
| LoRA adapter | 46 | 51 |
Most people want a file they can run on their own machine. If you only publish one format, make it a GGUF with a working template. The adapter is worth adding for anyone who wants to keep training.
Troubleshooting
The failures documented in kllm's troubleshooting guide:
| Symptom | Fix |
|---|---|
| Training process dies with no error | Out of memory. Check `dmesg \ |
{% raw %}UnboundLocalError mentioning dropout |
Set lora_dropout: 0.0
|
ImportError: cannot import name ... from 'transformers' |
uv pip install --force-reinstall "transformers==4.57.3" |
Python.h: No such file or directory |
sudo apt-get install python3.12-dev |
The full project, with the Qwen3-14B run and the download numbers, is in the STEM reasoning LLMs case study.
Frequently asked questions
How much VRAM do you need to fine-tune gpt-oss-20b?
With QLoRA, Unsloth's gpt-oss guide puts gpt-oss-20b at about 14 GB of VRAM and recommends at least 16 GB for stable runs. LoRA on BF16 weights needs about 44 GB. kllm, the toolkit I trained with, is verified on an RTX 5090 with 32 GB.
Can I run a fine-tuned gpt-oss-20b in Ollama?
Yes. Export it to GGUF, then either run ollama create with a Modelfile that carries the Harmony template, or pull it with ollama run hf.co//:Q4_K_M. The second route uses the chat template embedded in the GGUF file, not the Modelfile in the repo.
Why does my fine-tuned gpt-oss model output gibberish or never stop?
Usually the chat template. The template at inference has to match the one used in training, and Harmony ends a final answer with <|return|>. If that token isn't a stop token, generation runs past the answer.
Which GGUF quantization should I use for gpt-oss-20b?
Q4_K_M (15.8 GB for this model) is the smallest file and the one Ollama picks by default. Use Q8_0 (22.3 GB) if you have the memory and want output closer to the 16-bit model.
How do I set reasoning effort for gpt-oss in Ollama?
Harmony reads it from the system message as Reasoning: low, medium or high, and medium is the default. In a Modelfile, change that line inside the TEMPLATE and run ollama create again.
Top comments (0)