DEV Community

RainyTech
RainyTech

Posted on AI-assisted

Fine-Tune an LLM on Your Mac with MLX: QLoRA in 16 Seconds (Real Run, No Cloud)

I fine-tuned a language model on a Mac. The training took about 16 seconds, used about 2.5 GB of memory, and needed no cloud and no NVIDIA card. Just Apple's MLX.

Before training, the model wrote a long, rambling answer. After training, it answered short and clear, exactly like my examples. And the version with the best score gave the worst answer. That twist is the most useful thing in this post.

Every number below comes from the recorded run, which I show start to finish in this video:

Test machine: MacBook Pro, M2 Pro, 16 GB unified memory · macOS 26.5 · MLX 0.32.3 · MLX LM 0.32.0. The commands are identical on newer Macs like the Mac mini M5 Pro; only speed and memory change.

What fine-tuning actually changes (LoRA in one minute)

A language model predicts the next token, over and over. Training compares its guess with the right answer, measures the error (the loss), and nudges the weights to make it smaller.

Updating all 3 billion weights of a 3B model would need far too much memory. LoRA freezes them and trains two tiny matrices next to each layer instead: the new weight is the old weight plus a small update, B × A.

In this run that meant 3.3 million trainable parameters out of 3.09 billion: 0.108% of the model. Because the base model is 4-bit, this is QLoRA.

Why MLX, and not PyTorch?

MLX is Apple's machine learning framework, built for Apple Silicon and its unified memory (the CPU and GPU share the same RAM, so the model isn't copied around). MLX LM adds ready-made commands to run and fine-tune language models.

PyTorch on a Mac uses a different backend (MPS). Don't take a CUDA notebook, swap cuda for mps, and expect it to train.

1. Install MLX LM

In a fresh Python 3.12 environment:

uv pip install "mlx-lm[train]"
Enter fullscreen mode Exit fullscreen mode

2. Convert the model to 4-bit

I used Qwen2.5-3B-Instruct. Hugging Face was blocked on my network, so I downloaded the official weights from ModelScope and checked their SHA-256 hashes. With normal access, you can use the ready 4-bit model from mlx-community instead.

mlx_lm.convert --hf-path ./models/qwen2.5-3b-instruct -q --q-bits 4 \
  --mlx-path ./models/Qwen2.5-3B-Instruct-4bit
Enter fullscreen mode Exit fullscreen mode

9 seconds. 6.2 GB became 1.6 GB (4.501 bits per weight). Quick check: 3B weights × ~4.5 bits ÷ 8 ≈ 1.7 GB. Remember that's only the weights; training needs more on top.

3. Get a baseline first

Before training anything, I asked the base model the same question I would use later, with the same system prompt and temperature 0 so it's repeatable:

Explain the difference between authentication and authorization.

The answer was correct but long: numbered points, bold text, 151 tokens, at 84 tokens/s and 1.86 GB peak memory. Save this output. Without a baseline you can't prove anything improved.

4. Create the training data (JSONL)

Each line of the file is one chat: a system message, a user question, and the ideal answer.

{"messages": [{"role": "system", "content": "You explain software concepts in one or two short, clear sentences."}, {"role": "user", "content": "What is the difference between authentication and authorization?"}, {"role": "assistant", "content": "Authentication verifies who you are. Authorization decides what you are allowed to do."}]}
Enter fullscreen mode Exit fullscreen mode

Two rules that save hours:

  • Use chat messages, not hand-written special tokens. MLX applies the model's own chat template.
  • Keep similar examples in the same split, or your test set leaks into training.

My demo had 20 examples: 16 train, 2 validation, 2 held back for testing (train.jsonl, valid.jsonl, test.jsonl in ./data). That's only enough to test the pipeline. A real project needs hundreds of reviewed examples.

5. Train the LoRA adapter

mlx_lm.lora --model ./models/Qwen2.5-3B-Instruct-4bit --train --data ./data \
  --adapter-path ./adapters/smoke --batch-size 1 --num-layers 8 \
  --max-seq-length 512 --learning-rate 1e-5 --iters 20 \
  --grad-checkpoint --mask-prompt
Enter fullscreen mode Exit fullscreen mode
  • --grad-checkpoint saves memory.
  • --mask-prompt means it only learns from the answers, not from your questions.
  • 20 iterations: this is a smoke test.

Results:

Start End
Training loss 4.95 1.70
Validation loss 7.03 2.81

About 16 seconds, about 2.5 GB peak memory, with checkpoints saved at step 10 and step 20.

6. Did it actually get better? (the twist)

Test loss on the 2 examples the model never saw:

Model Test loss
Base model 3.99
Step-10 adapter 2.24
Step-20 adapter 1.89
mlx_lm.lora --model ./models/Qwen2.5-3B-Instruct-4bit --test --data ./data --adapter-path ./adapters/smoke
Enter fullscreen mode Exit fullscreen mode

Step 20 looks like the clear winner. Then I asked it the question:

Step 20: "Authentication verifies who you are."

That's it. It forgot authorization. It learned to be short, but too short. The step-10 checkpoint, with the worse loss:

Step 10: "Authentication verifies who you are. Authorization decides what you can do."

Short, clear, complete: 47 tokens instead of 151.

Lower loss does not always mean a better assistant. Always read the answers and compare checkpoints before you keep one.

7. Fuse the adapter into a standalone model

To use the step-10 checkpoint, copy 0000010_adapters.safetensors into its own folder as adapters.safetensors, together with adapter_config.json (I called it adapters/smoke-it10). Then:

mlx_lm.fuse --model ./models/Qwen2.5-3B-Instruct-4bit \
  --adapter-path ./adapters/smoke-it10 --save-path ./models/rainy-qwen-3b
Enter fullscreen mode Exit fullscreen mode

Still 1.6 GB, and it runs at 81 tokens/s. One tip: test the fused model with the same system prompt you trained with. Without it, it falls back to full-length answers.

8. Serve it as a local API with FastAPI (and the one-word fix)

A small FastAPI app loads the model and adapter once and serves a chat route:

from fastapi import FastAPI
from pydantic import BaseModel
from mlx_lm import load, generate

SYSTEM = "You explain software concepts in one or two short, clear sentences."
model, tokenizer = load("./models/Qwen2.5-3B-Instruct-4bit", adapter_path="./adapters/smoke-it10")

app = FastAPI()

class Question(BaseModel):
    question: str

@app.post("/chat")
async def chat(q: Question):  # async: MLX streams belong to the thread that loaded the model
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": q.question}]
    prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
    return {"answer": generate(model, tokenizer, prompt=prompt, max_tokens=200)}
Enter fullscreen mode Exit fullscreen mode
uvicorn app:app --port 8000
curl -X POST localhost:8000/chat -H 'Content-Type: application/json' -d '{"question": "What is a reverse proxy?"}'
Enter fullscreen mode Exit fullscreen mode

With a normal def route, my first request crashed with:

There is no Stream(cpu, 0) in current thread
Enter fullscreen mode Exit fullscreen mode

MLX streams belong to one thread, and FastAPI runs normal def routes in a worker thread. Making the route async def keeps it on the thread that loaded the model. This is a single-user demo (no streaming, no auth, no queue), so keep it on your own machine.

What about Ollama and GGUF?

Ollama is a separate path: it won't use your MLX adapter automatically, and MLX LM's built-in GGUF export currently supports Llama, Mistral and Mixtral, not Qwen. (If you fine-tune a Llama model instead, you can fuse it and import the safetensors into Ollama; I did that in my Ollama course.)

Memory tips from the run

  • Chatting needs less memory than training: 1.9 GB to chat, 2.5 GB to train for this 3B model on short examples.
  • Longer sequences and bigger batches push memory up fast. If you run out: batch size 1, shorter --max-seq-length, fewer --num-layers.
  • The flag is --num-layers, not the old --lora-layers.
  • Answers repeating? Check your data for duplicates before blaming the model.
  • For a real run, more iterations with gradient accumulation (for example 400 iterations, accumulation 4 = 100 weight updates). Measure your own speed first: at 0.2 it/s, 400 iterations take about 33 minutes.

Fine-tuning vs RAG vs a better prompt

  • Answers about documents that change → use RAG.
  • Just a different tone → try a better system prompt first.
  • Fine-tune when you have real examples and a behavior you can measure.

A small fine-tune won't make a small model smarter. It makes it more consistent.

Keep a record of every experiment

The exact model, package versions, the command, the random seed, your data splits, every adapter (including early checkpoints), and the before/after answers on the same test prompts. That record is how you prove your model really improved.


So, can you fine-tune an LLM on a Mac? Yes. With MLX, a 3B model trains in seconds in under 3 GB of memory. Start with the 20-step trial, read the answers, then scale up.

What would you teach your own local model? Tell me in the comments.

Top comments (0)