A large language model (LLM) is the combination of a huge Transformer neural network (the design you will meet in detail later) and racks of high-performance GPUs. For a business leader, the big question today is less about how an LLM works inside and more about how to use it, and that means understanding the modern API landscape.
In plain terms: An API is a waiter between your app and a kitchen. Your app writes an order (the prompt), the waiter carries it to the kitchen (the model), and the finished dish (the answer) comes back. You never have to own or run the kitchen.
Diagram: An app or agent sends a prompt to the model's API. The model generates the answer one token at a time and streams it back. See the animated version.
What an LLM API call looks like
You send the model a message, and you get generated text back. The request is a small piece of structured data. This is the typical shape (field names vary a little between providers, so treat it as an illustration):
import json
# The shape of a typical LLM API request. Field names differ a little between providers.
request = {
"model": "your-chosen-model",
"messages": [
{"role": "system", "content": "You are a careful IT assistant. Answer in two sentences."},
{"role": "user", "content": "Summarize this incident: the VPN gateway restarted at 03:15."},
],
"max_tokens": 200, # a cap on the length of the answer, which also caps its cost
"temperature": 0.2, # lower means more predictable wording
"stream": True, # receive the answer token by token
}
print(json.dumps(request, indent=2))
The landscape in 2026
APIs used to be about one application talking to another. Increasingly they are being built so that AI models and autonomous agents can use them directly, and that changes what a company has to manage.
- Closed commercial APIs (such as OpenAI, Anthropic and Google): you manage no infrastructure and you get immediate access to state-of-the-art reasoning. Providers increasingly offer very large context windows, with some now reaching a million tokens or more, and prompt caching, which makes repeated long instructions cheaper.
- API gateways as AI control layers: a gateway sits in front of the models and manages who can call them, which APIs an AI system is allowed to discover and use, what policies apply, and what agents actually do, so that risks can be caught.
- The routing approach: many production apps send simple tasks to cheaper, faster models and reserve the expensive flagship model for hard, multi-step reasoning. This controls cost and reduces dependence on one vendor.
Diagram: Many production apps put a gateway in front of their models. It applies policies, logs usage, and routes simple tasks to cheaper models while keeping the flagship model for hard reasoning and a self-hosted model for private data. See the animated version.
Diagram: When many calls start with the same long system prompt, providers can cache it so repeated reads cost less. The bars are illustrative, since exact prices vary by provider. See the animated version.
Diagram: A modern API gateway acts as a control layer for AI: it checks who is calling, limits usage, applies policies, records what happens, routes requests and tracks cost, including actions started by agents. See the animated version.
Putting it together
Think of an AI app as three layers: your application or agent, a gateway that applies the rules, and one or more models. Open-weight models, which you can run yourself, belong in this picture too. They are useful when data must stay inside your own walls, and we will build exactly that in Week 3.
Questions worth asking about any LLM API:
- How much does it cost per token, and is output priced higher than input?
- How large is the context window, and is prompt caching available?
- Where does my data go, and is it used for training?
- What happens if the provider changes or retires the model?
Coming Up Next
Day 9: Advanced prompt engineering: few-shot examples, chain-of-thought and structured JSON.
LargeLanguageModels #GenerativeAI #CloudComputing #TechStrategy #APIs
Originally published at https://sureshpallapothu.in/blog/day-8-llm-api-landscape, where this post includes animated diagrams.
Top comments (0)