DEV Community

Cover image for LoRA vs QLoRA vs full fine-tuning: cost, quality and when each makes sense
PRANJUL RATHOUR
PRANJUL RATHOUR

Posted on Originally published at pranjulrathour.scult.in

LoRA vs QLoRA vs full fine-tuning: cost, quality and when each makes sense

Every fine-tuning tutorial assumes hardware most students do not have. The choice between LoRA, QLoRA and full fine-tuning is mostly a hardware and data question, so let me lay out what each one actually does before you pick.

Full fine-tuning: every weight moves

You update all parameters of the model. It gives the most capacity to change behaviour and is what labs do for major model versions. It also needs memory for weights, gradients and optimiser states — several times the model size — which puts even a 7B model out of reach for a consumer GPU. For a student project it is almost never the right tool.

LoRA: train small adapters instead

LoRA freezes the base weights and trains low-rank matrices injected into attention (and often MLP) layers. The trainable parameters drop to a fraction of a percent, memory drops with them, and the result is a small adapter file you can swap in and out. Quality on narrow tasks is usually close to full fine-tuning.

QLoRA: LoRA on a quantised base

QLoRA loads the frozen base model in 4-bit precision and trains LoRA adapters on top. That is how FineTune Studio trains Qwen3-1.7B at around 3.2 GB of VRAM — hardware a student can borrow or rent for almost nothing. The adapters are trained in higher precision, so quality holds up well; the cost is slower steps because of dequantisation.

A decision guide

  • Under 8 GB VRAM, or a free cloud GPU → QLoRA. Nothing else fits.
  • 16–24 GB VRAM and a model up to 7–8B → LoRA in 16-bit for faster steps, QLoRA if you want a bigger base.
  • A narrow behaviour (format, tone, classification) → adapters are enough; full fine-tuning is waste.
  • Fundamentally new capabilities or a new language → full fine-tuning on serious hardware, and probably not your project this semester.

The part that matters more than the method

Whichever you choose, evaluate base against tuned on held-out examples before you believe anything — the workflow in how to evaluate a fine-tuned model honestly. Most disappointing fine-tunes were not the wrong method; they were the wrong dataset.

About Pranjul Rathour

Pranjul Rathour in a suit and tie with a lanyard at a formal campus event
At a formal campus event

Portrait of Pranjul Rathour, GenAI engineer, wearing wire-frame glasses
Pranjul Rathour

Pranjul Rathour presenting on stage in a blue polo, with his Annapurna demo video on the screen behind him
Presenting Annapurna on stage

Pranjul Rathour giving a talk titled 'How and what I do', with demo videos of his products Vaidya and Annapurna on screen
Talking through the products he has shipped

Pranjul Rathour in a black t-shirt holding a microphone in front of a chequered wall
On the mic

Pranjul Rathour is a GenAI engineer from Kanpur, India, and CTO at SCULT INDIA, currently shipping production RAG,
fine-tuning and agentic AI systems, mentoring 200+ students through TechVerse Enclave, and judging and speaking at
student hackathons across India. Updated 2026-09-06.

Reach out if you want to talk GenAI, book a campus session, or invite him to judge:


Pranjul Rathour · GenAI engineer, 3x hackathon winner, campus mentor. Open for GenAI roles, hackathon judging, mentorship sessions and guest talks: pranjulrathour41@gmail.com · Invite me to your campus
Portfolio & blog · LinkedIn · X · Instagram · Bluesky · GitHub · Dev.to

Top comments (0)