A small model is ready for traffic only when you can name the precision you will serve, whether an adapter passed the same evals as a fuller update, which machine will run it, and which live signal would make you pull it.
What does lower precision take away?
Quantization stores weights in fewer bits, often 8-bit or 4-bit instead of 16-bit. The model takes less memory and runs faster. That is how a capable model fits a laptop, a phone, or a cheaper GPU, and how serving cost drops when call volume is high. Networks often tolerate small numerical noise, so a careful cut can stay close to the full-precision model on the job you measured.
Quality is the loss. Push the bit width too low and accuracy drops, and the damage is uneven across tasks. I would treat precision as a dial on your evals: ship the lowest precision that still passes, and keep that file as the one you score. If training stays at full precision and only the server copy is quantized, the pass has to be on the quantized weights. Fergal Reid, chief AI officer at Intercom, said most of what his team does now is full supervised fine-tuning, unquantized, with reinforcement learning on some harder work. Match the number format in the eval to the number format in the deploy.
Is a LoRA adapter enough for this job?
LoRA freezes the pretrained model and trains small low-rank adapter matrices. Full fine-tuning updates every weight and produces another full copy per task. Adapters are tiny, so training is cheaper and faster, and you can store many task adapters instead of many full models. Adapters can be swapped or combined. Quantized LoRA (QLoRA) pushes that training cost lower still. I would start with an adapter when you want a specialized behavior and a full fine-tune is more than you need or can pay for.
Fine-tuning belongs on a repeated task, with examples that look like real use and an eval set that can show whether the change helps. A smaller specialized model can still lose once training and serving are both on the bill.
Reid said they started with LoRA and other parameter-efficient methods. Over time the work shifted toward full supervised fine-tuning, plus reinforcement learning for harder objectives. He also said the harder spend was infrastructure: an internal setup so a scientist can log in and run a distributed job on large GPUs on AWS. He described their standard instance at the time as one node of H200s. If the adapter passes the task evals you trust, stop. If it misses a policy you can already hit with a fuller update on those same evals, spend the full update on that task.
Should this run on a server or on the device?
Edge AI runs the model near the data. A camera that detects objects locally, or a phone that recognizes speech on device, is doing that. Processing close to the source can shorten response time and cut what you send over the network. It can also keep sensitive recordings local, depending on what the rest of the application stores and transmits. The weights still have to fit the device memory, the processing capacity, and the power budget. Smaller models and quantization help only if the task evals still pass. A hybrid split can keep urgent work on the device and send heavier requests to the cloud. The split depends on the application.
Maxime Labonne, head of post-training at Liquid AI, wanted an on-device model that is fast on a phone or a wearable. He argues you have to time the operators on the target hardware. Theoretical speed from the math can fail once the operator is on the device. His team measured on a Samsung phone for that reason, alongside a large set of pretraining evaluations (he cited over 100).
So on a Samsung phone, for example. So we would not be misguided into believing that our our operator is is really good. It's very fast. No. Like, we could measure it and make sure that it's actually working in practice.
Maxime Labonne, Head of Post-Training at Liquid AI, on Chain of Thought episode 43
The same small-model range also has a server job. Labonne said a model on the order of 350 million parameters can be deployed on GPU at scale for high-volume work, including turning unstructured text into a JSON object you specify, in areas such as ecommerce and finance. He also noted that API calls are not always possible, and that a car assistant that depends on online connectivity would not be very useful most of the time. If the call has to succeed offline, or the recording should stay on the device, put the model there and prove it fits. If you need many parallel calls and the data can leave the machine, host it yourself and remove the per-token dependency on an outside provider. Either way, time it on the box you will use.
Labonne is also clear about the ceiling: knowledge depends on parameter count, so distilling into a one-billion-parameter model will not make it as smart as a three-billion one, and small models miss frontier quality on complex workflows. Match the model to the task, then match the task to a machine you measured.
What should you watch after launch?
For each task, take the smallest model that still passes your evals. That is where cost, speed, and quality tend to balance in production. Offline passes are the entry ticket. Reid's team still back-tests on a suite built over time, including hallucinations customers reported in the wild, and they like to try a well-prompted but untrained model before they post-train, to see whether it is even close. Then they A/B a release candidate in production. Their reinforcement learning stays offline. They do not run an open-loop setup in production.
But then nothing beats testing in production. We test everything in production.
Fergal Reid, Chief AI Officer at Intercom, on Chain of Thought episode 49
On their customer-service agent, the ground truth is the hard resolution, where the user confirms the question was resolved. The soft resolution, where Fin answers, offers a human, and the user leaves without saying anything, is the proxy. On a model change he wants soft resolutions up and hard resolutions at least flat. If the proxy rises while confirmed success falls, the system is deflecting. He also argues that increased latency almost always leads to more resolutions, including hard ones, possibly because users read the wait as more work done. A back test that ignores latency can prefer a slower answer for the wrong reason.
You may not have his resolution labels. You do need one confirmed outcome and one proxy, logged with latency, on an A/B before the full ramp. The sketch below is only a gate I would write for myself. It is not a pipeline Intercom or Liquid AI runs.
# Sketch only. Illustrates a go/no-go. Not a vendor client.
def go_live(task_ok, quant_ok, adapter_ok, fits_device, ab_ok, latency_neutral):
if not task_ok:
return "hold: task eval failed"
if not quant_ok:
return "hold: scored precision is not the file you will serve"
if not adapter_ok:
return "hold: adapter missed; compare a fuller update"
if not fits_device:
return "hold: weights miss memory or power on the target"
if not ab_ok:
return "hold: confirmed outcome moved the wrong way"
if not latency_neutral:
return "hold: win is explained by delay"
return "ramp"
The longer conversations these checks come from sit on the AI engineering topic page, with transcripts and chapters.
Checklist
- Score the weights you will serve, at the bit width you will serve.
- Start with a LoRA adapter. Move to a full fine-tune only when the adapter misses the same evals.
- Time the model on the target phone, or on the server instance you will rent, and confirm it fits memory and power.
- A/B against a user-confirmed outcome. Keep the proxy from rising while confirmed success falls.
- Log latency next to quality so a slower answer cannot win by accident.
The longer explainer, with the episode clips, is on Chain of Thought. It draws on this episode.
Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.
Drafted with AI assistance from the episode transcripts.
Top comments (0)