Narrow steps that run on every request, on a model stronger than that step needs, are candidates for a cheaper model. Measured spend then shows where the savings lie. Prove the step on the strongest model, move that one job to a smaller or fine-tuned model, and keep the product metric where it was. Fergal Reid said Intercom's query canonicalization could have been like a quarter million dollars a month on GPT-4.1, and that a post-trained 14-billion-parameter Qwen model saved almost all of several hundred thousand a month in inference on that task alone.
Fergal Reid is Chief AI Officer at Intercom and spoke on Chain of Thought in February 2026. Loïc Houssier is CTO of Superhuman Mail and spoke in May 2026. Loïc starts on the best model and cuts cost only if the feature works. Fergal's canonicalization step was already on GPT-4.1, then a post-trained 14-billion-parameter model.
How do you choose which agent step can leave the frontier model?
Pick a step with a stable job, high call volume, and an output you can score without re-judging an open-ended answer. Fergal described Fin, Intercom's customer service agent, as a cluster of 10 or 15 machine learning problems. The hardest prompt, the one that answers the end user, stayed on a frontier Claude Sonnet model. Ancillary prompts were different. Those ran on the smallest, cheapest, lowest-latency model that still produced the end-user experience they wanted.
Query canonicalization was the step he priced. Before retrieval, an LLM summarizes the question and rewrites it into a canonical form. End users ask in colloquial ways that vary from person to person and can be inaccurate. Retrieval models are less powerful than a modern LLM, so the rewrite improved retrieval and, through that, the later answer. The job did not need GPT-4 or Claude Opus. GPT-4.1 did it, and the spend was large.
I don't remember exactly, but it could have been like a quarter million dollars a month on just that summarization task.
Fergal Reid, Chief AI Officer at Intercom, on Chain of Thought ep 49
A 14-billion-parameter Qwen model, post-trained for that summarization task, replaced the GPT-4.1 call. The swap saved almost all of several hundred thousand a month on inference and gave finer control over the step. Before post-training, the team runs a production test of a well-prompted but untrained model, to see whether it is close to the job.
Loïc splits models by use case the same way. There is no single model for the product. Auto labels began as classification on a strong model (he recalled a 4.0 model), run on every email a user has, which made it expensive. After the team knew which 10 or 12 labels were useful, including FYI, pitch, partner, and newsletter, the job looked like a typical BERT classifier. The move he describes from there is to fine-tune an open model a bit and put it on cheap, dedicated infrastructure for inference at scale. The agentic framework stayed on the latest Opus.
So depending on the use case, and usually we start by expensive, the best. And then if the feature is successful, how can we optimize? How can we drive the cost down?
Loïc Houssier, CTO at Superhuman Mail, on Chain of Thought ep 59
How do you hold quality when a cheaper model takes the step?
Keep an outcome a cheaper model cannot raise by doing less. Fergal's check is a production A/B test on two signals. A soft resolution is Fin giving an answer it treats as correct, offering a human, and the user leaving without a reply. A hard resolution is the user confirming the question was resolved, with something like thanks or yes. Whenever you give an answer, a hard resolution happens about 30 to 40 percent of the time, and that rate tends to be about 30 to 40 percent of the soft-resolution rate. Hard resolutions are the North Star and the ground truth.
And we check to see that the soft resolutions have gone up, but also that the hard resolutions don't go down, basically.
Fergal Reid, Chief AI Officer at Intercom, on Chain of Thought ep 49
The ratio of soft to hard resolutions should stay roughly constant for every model change, so the product does not become a deflection machine. The switch from GPT to Claude Sonnet was a multi-million end-user interaction A/B test. Fin was resolving more than a million conversations successfully a week, possibly a good bit over that, which is what makes those tests sensitive. They care about a tenth of a percentage point of resolution, and every backtest can deviate from what production does.
Resolution was about 30 percent at launch and was heading towards 70 percent. The vast majority of that improvement was in the surrounding systems, not the core frontier model: the retrieval model and a custom re-ranker built almost from scratch. An ancillary step can still be the expensive line and a real part of the metric. Build that step's back test from production cases, including customer-reported hallucinations, which they capture for that purpose. Set the pass bar before you look at the savings.
On a product without those two tiers, use the same shape of rule. The user-visible outcome stays flat or improves against the strong-model version. A proxy that moves because the model abstains more, or hands off more, does not pass.
How do you rank steps by monthly spend?
Sort monthly spend by step, and mark a row only when the job is narrow and the dollars are high. A rare frontier call can look pricey per request and still be the wrong thing to shrink. Rank the step that runs on every ticket or every email.
Illustrative sketch with hypothetical totals. Not a production report, and not either team's implementation.
steps = [
{"step": "canonicalize_query", "model": "large_general", "spend_usd": 4800, "narrow": True},
{"step": "answer_user", "model": "frontier", "spend_usd": 3100, "narrow": False},
{"step": "label_email", "model": "large_general", "spend_usd": 2600, "narrow": True},
{"step": "draft_reply", "model": "frontier", "spend_usd": 900, "narrow": False},
]
def rank_for_cheaper_model(steps, spend_floor=1000):
ranked = sorted(steps, key=lambda row: row["spend_usd"], reverse=True)
report = []
for row in ranked:
report.append({
"step": row["step"],
"model": row["model"],
"spend_usd": row["spend_usd"],
"try_cheaper_model": row["narrow"] and row["spend_usd"] >= spend_floor,
})
return report
narrow is a judgment you set after reading traces. A provider does not return it. If a step retries, add every attempt into that step's monthly total, or the ranking hides the failure. Then read a few full traces of the top flagged step. The steps these two teams moved were repeated jobs with an output they could grade: a canonical query, or a label from a set they had already settled. The open-ended answer stayed on the frontier model.
FAQ
Which Intercom step left GPT-4.1?
Query canonicalization, the summarization that rewrites the end user's question before retrieval. A post-trained 14-billion-parameter Qwen model replaced GPT-4.1 on that job. Claude Sonnet stayed on the hardest prompt.
What did Superhuman take off the strong model?
Auto-label classification. A strong model ran on every email until the team knew the useful set was 10 or 12 labels. Those labels then became a BERT-classifier-type job for a lightly fine-tuned open model on cheap, dedicated infrastructure. The agentic path stayed on the latest Opus.
What has to stay steady after the swap?
For Fin, hard resolutions should not go down, and the soft-to-hard ratio should stay roughly constant. On another product, hold the same user-visible outcome in a production comparison against the strong-model version. A lower bill with a worse outcome fails the change.
Takeaway
- Rank steps by monthly spend, with every attempt included, before you pick a model to replace.
- Prove the feature on the strongest model so you know the job is real.
- Move a narrow, high-volume step, such as canonicalization or a fixed label set.
- Run a well-prompted smaller model on the live task before you invest in post-training.
- Ship the cheaper model only when the production outcome holds. On Fin, hard resolutions do not fall.
The longer explainer, with the episode clips, is on Chain of Thought. It draws on this episode.
Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.
Drafted with AI assistance from the episode transcripts.
Top comments (0)