Wiring an LLM API into a product looks like a two-day job. The first version usually is. The part that takes longer is everything around the call itself, and that is where the money leaks.
This is the checklist I would hand someone who is about to put a model behind a real feature.
1. Count output tokens, not just input
Teams budget with the input price, then get an invoice that is several times higher. Output tokens usually cost multiples of input tokens, and chatty assistants produce far more output than input.
Two habits fix this: cap max_tokens per call, and log input and output counts separately from day one. A single combined number hides the line you actually need to watch.
2. Make retries explicit and cheap
A malformed response still costs tokens. If your client retries three times by default, a broken prompt is billed four times.
- Set an explicit retry count, not the library default.
- Only retry on transport errors and rate limits, not on every exception.
- Validate the shape of the response before you use it, so you retry less.
3. Turn on prompt caching if your prompt has a stable prefix
If every request carries the same system prompt, tool definitions and documentation, you may be paying for it repeatedly. Providers that support prefix caching charge less for the repeated part. This is usually the single largest line item you can cut without touching quality.
4. Decide what happens when the model is slow before it is slow
Streaming helps perceived speed but makes errors harder to handle: you have already sent half an answer when the connection drops. Pick a rule now - either you stream and buffer for the client, or you do not stream and hold a timeout budget.
Write the fallback down: smaller model, cached answer, or an honest "try again". "Nothing" is also a choice, just an expensive one.
5. Version your prompt like code
Prompts change behaviour in ways that are invisible in a diff. Keep them in files, review them, and record which version produced which request. When quality drops, the first question is always "what changed", and you want an answer that is not a guess.
6. Build a tiny eval set before you tune anything
Twenty real examples with expected properties beat a hundred vibes. You do not need a framework. A JSON file and a script that prints a pass rate is enough to stop you from shipping a regression because it read better in a demo.
7. Track cost per feature, not cost per account
An account that uses chat heavily and one that runs a nightly batch job are not comparable. Attribute spend to the feature that caused it. That is the number that tells you whether a feature is worth keeping.
8. Have an exit
Models change price, get deprecated, or get worse. Keep the call behind one interface, and keep the prompt portable. If swapping providers is a two-week project, that is a business risk, not an engineering preference.
The short version
Cap your output, retry on purpose, cache the prefix, budget the timeout, version the prompt, measure per feature, and keep the door open.
I keep pricing and capability notes for the models I use at haiai123 - 1,487 tools across 17 categories, with the data source and refresh date printed on every table.
Top comments (0)