“Cheapest GPT API” is an attractive search query—and a surprisingly hard claim to prove.
A price per million tokens is not enough to compare two APIs. One service may publish separate input, cached-input and output rates; another may publish a blended rate. The numbers describe different billing formulas until you test them with the same workload.
Disclosure: I work on Model.sale, a prepaid API service. This guide is about a comparison method, not a claim that one service is always cheapest.
Start with the units
For a split tariff, estimate a request as:
input_tokens × input_rate + cached_input_tokens × cached_rate + output_tokens × output_rate
Then add cache-write, tool, image, audio, regional or processing-mode charges when they apply. Reasoning tokens are generally included in reported output usage for APIs that expose them; check the model’s usage contract rather than counting visible text alone.
For a blended tariff, the estimate may instead be:
charged_tokens ÷ 1,000,000 × blended_rate
A blended rate can be simple to forecast, but it is not directly comparable to an input-only rate or to a split input/output price. The result depends on the workload’s input/output mix, cache hit rate, reasoning usage, minimum request charges and failed retries.
OpenAI’s current API pricing lists model-specific rates and processing modes. Its prompt-caching guide explains that cached input changes cost and that cache-write treatment varies. Always check the live documentation: model pricing and options change.
Build one representative workload
Before comparing providers, write down a small test plan:
- exact model ID and API endpoint;
- approximate input and output token counts;
- whether prompts reuse a stable prefix;
- reasoning effort, tools and other request options;
- whether the request is streamed;
- maximum acceptable time to first token;
- what counts as success.
Use at least three shapes: a short interactive prompt, a long-context coding task and an output-heavy task. If caching matters, run both a cold request and a repeated request with the same prefix. Do not assume a cache hit just because two prompts look similar.
Measure the final bill, not just the estimate
For each test, record:
- request ID, exact model and endpoint;
- input, cached-input, output and reasoning usage when reported;
- time to first token, total latency and whether the stream reached its terminal event;
- final ledger charge after settlement;
- any retry, timeout or incomplete response.
A streaming response can fail after producing useful text, or disconnect before final usage arrives. Decide how your application handles that case and verify the billing record before retrying. Otherwise, retries can distort both cost and reliability comparisons.
Use a small budget cap during experiments. Keep separate API keys for testing and production, and set spend or rate limits where available.
A minimal OpenAI-compatible smoke test
Discover the exact model IDs your account can call instead of copying a model name from an old post:
export OPENAI_BASE_URL="https://api.model.sale/v1"
read -rsp "API key: " OPENAI_API_KEY; echo
curl "$OPENAI_BASE_URL/models" \
-H "Authorization: Bearer $OPENAI_API_KEY"
Then select a model returned by that endpoint and make a tiny request with a low output limit. Keep the key in an environment variable or secret manager; never put it in a URL, public code sample, shell-history literal or screenshot.
For Model.sale, the public model catalog shows model status and the latest listed rate, while the pricing methodology describes its blended charged-token unit. A listed rate is not a promise that a temporarily unavailable model can accept requests, so check availability immediately before testing.
Compare cost and reliability together
A low nominal rate is not the whole production cost. Include successful output, retries, latency, stream completion and the engineering time needed to work around missing features. A useful report shows:
| Measure | Why it matters |
|---|---|
| Cost for the same token mix | Makes different billing formulas comparable |
| P50/P95 time to first token | Captures interactive experience |
| Successful terminal streams | Detects partial or broken responses |
| Usage reconciliation | Confirms the final charge matches reported usage |
| Rate limits and model availability | Shows whether the test can scale |
Publish your assumptions and test date. If the formulas cannot be normalized, say so instead of presenting a percentage “saving.” A workload-specific result is useful; a universal “cheapest API” ranking usually is not.
A practical decision rule
Choose the API that meets your model, feature, privacy and reliability requirements at a predictable cost for your workload. Re-run the same test after a price change or a major model update. For teams evaluating prepaid APIs, a calculator and a dated pricing page are more useful than a single headline number.
Further reading: Model.sale live models · pricing methodology · workload cost calculator.
Top comments (0)