DEV Community

Cover image for Your AI Agent Has a Context Budget: Treat It Like a CPU Budget
Karthigayan Devan
Karthigayan Devan

Posted on

Your AI Agent Has a Context Budget: Treat It Like a CPU Budget

The 3 AM page

Picture this. You get paged at 3 AM for a production outage.

A teammate hands you one log file with the exact error in it. You find the problem in five minutes.

Now replay the same night. This time your teammate hands you that same log file, plus 500 unrelated log files, three months of Slack threads, 20 architecture docs, 100 dashboards, and every incident report from the last five years.

You have more information. You do not have more clarity. The answer is still in there somewhere, but now you have to dig for it, and you might miss it completely.

This is exactly what we are doing to our AI agents.

We keep making agents "smarter" by giving them more. More docs. More tools. More memory. More history. More MCP servers. More retrieved chunks. And then we act surprised when the agent gets slower, burns more tokens, picks the wrong tool, and ignores an instruction that was right there in the prompt.

I think the fix starts with one mindset change:

Context is a resource. It needs a budget, just like CPU and memory.

The context window is not your context budget

If a model supports a 200K token context window, that does not mean every request should fill 200K tokens.

We already know this lesson from infrastructure. If a node has 64 GB of RAM, we do not tell our service "great, load everything into memory." We set requests and limits. We decide what the workload actually needs, and we leave headroom.

The same thinking applies here:

Concept What it answers Example
Context window What the model can physically accept 200K tokens
Token budget How many tokens a call can consume or generate in total 50K tokens
Context budget How much information we intentionally put in front of the model for this task 20K tokens

The model's limit is capacity. The context budget is your operating point. Capacity and operating point are different things, and good engineers never confuse the two.

What actually eats your context

Every request is built from pieces, and every piece takes a bite:

Context
  ├── System instructions
  ├── User request
  ├── Conversation history
  ├── Memory
  ├── Tool definitions
  ├── Retrieved documents
  ├── Examples
  └── Tool results
Enter fullscreen mode Exit fullscreen mode

So instead of asking "how much context can my model handle?", I have started asking a different question:

"How much context should this task be allowed to consume?"

Take a simple request: "Find a vegetarian Indian restaurant near my hotel tonight."

The agent does not need the full restaurant database, every travel preference the user ever shared, every map tool, and the whole chat history. It needs five things: current location, cuisine, dietary preference, time, and a couple of search and routing tools.

That smaller context is not a limitation. It is an advantage.

Write the budget down

Here is the part that changed how I think about this. Once you write the budget down, context stops being an invisible blob inside an API call. It becomes something you can review, measure, and argue about in a design review.

For a customer support agent, it might look like this:

# Context budget for: customer-support-agent
total: 20000
allocations:
  system_instructions: 2000
  user_request: 1000
  recent_conversation: 3000
  relevant_memory: 2000
  tool_definitions: 3000
  retrieved_documents: 5000
  tool_results: 4000
Enter fullscreen mode Exit fullscreen mode

If you work with Kubernetes, this should feel familiar. It is basically a resource spec for the model's attention.

The numbers are not universal. The point is that you have numbers at all. Now you can ask much better questions:

  • Why did this request use 38K tokens when the budget is 20K?
  • Why are tool definitions eating 12K?
  • Why did retrieval return 15K tokens when the answer needed 2K?
  • Why is the customer's account ID in the prompt three times?

Not all tokens are worth the same

Say you have 10,000 tokens to spend.

Option one: a 6,000 token document that is loosely related to the question.

Option two: a 500 token snippet with the exact business rule that decides whether the request is allowed.

Option two is worth far more, at a twelfth of the cost. So I like to think in terms of context value density:

context value density = useful information / tokens consumed
Enter fullscreen mode Exit fullscreen mode

This is not a metric the model gives you. It is a habit of mind. Before anything goes into the context, ask one question:

"What decision will this help the agent make?"

If you cannot answer that, it probably does not belong in this request.

The six places your budget leaks

In my experience, context does not blow up in one big moment. It leaks, quietly, from a handful of places.

Leak 1: Tool definitions

Connect an agent to 20 MCP servers and you might expose 500 tools. If every request carries all 500 tool descriptions, you have spent a big chunk of your budget before the agent has even read the user's question.

The key idea here is:

Installed capability is not the same as visible capability.

An agent can have 500 tools installed and only see the 10 that matter right now:

                 User request
                      |
                      v
              +---------------+
              |  Tool router  |
              +-------+-------+
                      |
         +------------+------------+
         |            |            |
         v            v            v
       Maps       Database     Payments
         |            |            |
         +------------+------------+
                      |
                 3 to 10 tools
                      |
                      v
                    Agent
Enter fullscreen mode Exit fullscreen mode

This is also why on-demand skill loading (like Google Cloud's Skill Registry, where agents search for skills by intent and load only what fits) is such a good pattern. It is a context budget decision baked into the platform.

Leak 2: Retrieval

A typical RAG pipeline ends with "top 20 documents." Why 20? Why not 5? Why not 2?

Most of the time, "top 20" is a config default, not an engineering decision. If five chunks hold everything the agent needs, the other fifteen are not free. They dilute the signal.

A budget-aware pipeline looks more like this:

Question
   |
   v
Retrieve candidates
   |
   v
Rank -> Filter -> Dedupe -> Compress
   |
   v
Fit to retrieval budget (e.g. 5K tokens)
   |
   v
LLM
Enter fullscreen mode Exit fullscreen mode

The question shifts from "which documents are relevant?" to "what is the smallest set of information that gets this task done?"

Leak 3: Conversation history

A 40 minute conversation piles up fast. A naive agent replays all of it on every turn:

Message 1
Message 2
...
Message 100
Current request
Enter fullscreen mode Exit fullscreen mode

That is not memory. That is history replay.

A better approach keeps turning the conversation into state: important facts, decisions made, open tasks, and the current goal. The goal is not to remember every sentence. It is to keep what still matters.

Leak 4: Stale memory

Not every fact should live forever. Different facts have very different shelf lives:

Fact Shelf life
Preferred language Months or years
Home city Months or years
Current hotel Days
Current weather Hours
Login session state Minutes
Stock price Seconds

So give memory items metadata, just like a cache entry:

{
  "fact": "Restaurant is currently open",
  "source": "Places API",
  "timestamp": "2026-10-05T18:15:00Z",
  "ttl_seconds": 900
}
Enter fullscreen mode Exit fullscreen mode

Otherwise the agent makes a perfectly logical decision based on something that was true yesterday. The model was not wrong. Our context was stale.

Leak 5: Fat tool responses

This is my favorite one, because it hides so well.

We tune the prompt. We tune retrieval. We tune memory. Then a tool call comes back like this:

{
  "id": "...",
  "metadata": {},
  "debug": {},
  "internal_state": {},
  "records": ["... thousands of lines ..."]
}
Enter fullscreen mode Exit fullscreen mode

The tool worked fine. It just was not built for an AI consumer. A human UI can choose what to display. An agent swallows everything.

Agent-facing tools should support field selection, filtering, sorting, pagination, limits, and summaries. Instead of "here is everything I found," a good tool says "here are the five things that matter for your next step."

Leak 6: Paying for the same fact twice

Imagine the fact "the user's project is in us-central1" shows up in the system prompt, the chat history, memory, a tool response, and a retrieved doc.

That is one fact, paid for five times.

This is the same lesson as data normalization. Keep one canonical copy of each fact, and reference it, instead of scattering duplicates across the context.

One budget does not fit every task

A single global budget for the whole agent makes about as much sense as giving every pod the same CPU limit. Different work needs different resources:

User request
     |
     v
Task classifier
     |
     +---- simple lookup -------> ~5K
     |
     +---- support workflow ----> ~15K
     |
     +---- code migration ------> ~30K
     |
     +---- deep research -------> expandable, with checkpoints
Enter fullscreen mode Exit fullscreen mode

These numbers are only examples. Your evals should set the real ones.

Context is a runtime problem, not a prompt problem

A single LLM call is easy to budget. Agents are harder, because context grows on every loop:

User -> Agent -> Tool -> Result -> Agent -> Tool -> Result -> Agent -> Answer
Enter fullscreen mode Exit fullscreen mode

One tool returns 5K tokens. The next returns 8K. The agent adds 3K of reasoning notes. If you keep everything, the context keeps climbing until something breaks.

That is where a context manager comes in. It sits in the loop and makes a decision after every step: keep, prune, compress, or drop and re-fetch later.

            +------------------+
     +----> | Context manager  |
     |      +--------+---------+
     |               |
     |    +----------+----------+
     |    |          |          |
     |    v          v          v
     | Retrieve    Prune     Compress
     |    |          |          |
     |    +----------+----------+
     |               |
     |               v
     |              LLM
     |               |
     |               v
     |         Tool execution
     |               |
     +---------------+
          tool response
Enter fullscreen mode Exit fullscreen mode

What happens when you go over budget?

Every production agent needs an answer to this. Say the budget is 20K and the loop has piled up 27K. You have three choices, and they map nicely to things we already know from Kubernetes:

Strategy Kubernetes cousin Trade-off
Truncate the oldest content Evict the oldest pod Simple, but the oldest info is often the most important (like the original goal)
Reject the request OOMKilled Safe, but a bad user experience
Compress: dedupe, summarize, drop low value, keep critical facts Graceful degradation More work to build, but the agent keeps going

Whatever you pick, pick it on purpose. You do not want to discover your context strategy from a production incident.

A small code sketch

Here is a minimal Python version of a budget-aware context builder. It is not production code, but it shows the core ideas: priorities, per-task budgets, no duplicates, and a hard stop when a must-have item does not fit.

from dataclasses import dataclass
from collections import defaultdict

BUDGETS = {"lookup": 5_000, "support": 15_000, "migration": 30_000}

@dataclass
class ContextItem:
    source: str     # "system", "memory", "retrieval", "tool_result", ...
    text: str
    priority: int   # 1 = must keep, 2 = useful, 3 = nice to have
    tokens: int     # count with your model's tokenizer

def build_context(items: list[ContextItem], task_type: str):
    budget = BUDGETS[task_type]
    kept, seen, used = [], set(), 0

    for item in sorted(items, key=lambda i: i.priority):
        key = item.text.strip().lower()
        if key in seen:
            continue  # don't pay for the same fact twice

        if used + item.tokens > budget:
            if item.priority == 1:
                raise ValueError(
                    f"Must-keep item from '{item.source}' does not fit "
                    f"the {budget} token budget for '{task_type}'"
                )
            continue  # skip lower priority items that don't fit

        kept.append(item)
        seen.add(key)
        used += item.tokens

    # per-source breakdown, ready to emit as metrics
    breakdown = defaultdict(int)
    for item in kept:
        breakdown[item.source] += item.tokens

    return kept, used, budget, dict(breakdown)
Enter fullscreen mode Exit fullscreen mode

In real life you would swap the exact-match dedupe for something smarter, and replace "skip" with "summarize" for big low-priority items. But even this simple version gives you something most agents do not have today: a written rule for what gets in.

You can't manage what you don't measure

If context is a resource, it needs dashboards. The total token count is not enough. You need to know where the tokens came from:

context_budget        20,000
context_tokens        17,400
context_utilization   87%

by source:
  retrieval           31%
  tool_results        24%
  tools               17%
  conversation        12%
  system               8%
  memory               8%
Enter fullscreen mode Exit fullscreen mode

Now you have something actionable. Retrieval is the biggest spender? Tighten top-k or add a reranker. Tool results are huge? Fix the tool's response shape.

The most useful chart of all is context size against task success, from your own evals. A made-up example of what it might look like:

10K context  ->  87% success
20K context  ->  93% success
30K context  ->  92% success
50K context  ->  84% success
Enter fullscreen mode Exit fullscreen mode

If your real numbers look anything like this, then "just add more context" is actively making the agent worse past a certain point. Your sweet spot is where the curve peaks, not where the model's limit is.

This is not just a hunch. Researchers have started calling this effect context rot: model performance tends to drop as input length grows, even on simple tasks. The exact curve depends on the model and the workload, which is exactly why you have to measure your own.

The FinOps angle

Agents rarely make one model call. A single task might make ten.

30K tokens per step x 10 steps = 300K input tokens
12K tokens per step x 10 steps = 120K input tokens
Enter fullscreen mode Exit fullscreen mode

That is 60% fewer input tokens for the same task. And if the smaller context is also more focused, quality can go up at the same time. It is rare to get a cost win and a quality win from the same change, so this one is worth chasing.

This is why I do not see context budgeting as a prompt trick. It sits right where quality, latency, cost, reliability, and tool selection all meet. Those are production concerns, and they deserve production engineering.

Putting it all together

Here is how I would sketch a budget-aware agent:

                      User
                       |
                       v
               +----------------+
               | Task classifier|
               +-------+--------+
                       |
                       v
               +----------------+
               | Context budget |
               +-------+--------+
                       |
          +------------+------------+
          |            |            |
          v            v            v
       Memory      Retrieval      Tools
          |            |            |
          +------------+------------+
                       |
                       v
               +----------------+
               |    Context     |
               |   optimizer    |
               +-------+--------+
                       |
                       v
                      LLM
                       |
                       v
                Tool execution
                       |
                       v
               +----------------+
               | Context manager| ---> back to optimizer
               +----------------+
Enter fullscreen mode Exit fullscreen mode

The most important box in this picture is not the LLM. It is the layer that decides what the LLM gets to see.

A checklist you can steal

Context

  • What does this task actually need?
  • What gets loaded automatically, whether it is needed or not?
  • What is duplicated? What is stale?

Tools

  • How many tool definitions does the model see per request?
  • Can tools be discovered on demand instead?
  • Do tool responses include fields the agent never uses?

Retrieval

  • Why this many chunks? Who decided?
  • Do you rerank, filter, and dedupe before the model sees anything?
  • Are you measuring whether retrieved chunks were actually used?

Memory

  • What is short-term and what is long-term?
  • Which facts have a TTL?
  • What should the agent forget?

Observability

  • Do you track tokens per source, not just the total?
  • Do you know your context utilization?
  • Do you have an eval that shows where more context stops helping?

Runtime

  • Does every task get the same budget? Should it?
  • What happens when the budget is exceeded?
  • Can the agent re-fetch later instead of carrying everything around?

Limitations and honest caveats

A few things to keep in mind before you go and cut every prompt in half:

  • There is no universal number. The budgets in this post are examples. Your evals should set yours.
  • Compression is not free. Summarizing history costs an extra model call, and a bad summary can quietly drop the one fact that mattered.
  • Over-pruning is a real failure mode. An agent that cannot see the right tool or rule will fail just as badly as one buried in noise.
  • Token counts vary by model. The same text can be a different number of tokens on different models, so budget per model.
  • Context rot behaves differently across models. Some handle long inputs better than others. Measure on the model you actually ship.

Conclusion

We already know how to manage scarce resources. We have CPU scheduling, memory limits, connection pools, queues, and backpressure. Context is just the newest resource on that list, and agents need the same discipline.

The goal is not to give an agent the most information possible. The goal is to give it the smallest useful set of the right information, at the right moment.

So the next time your agent gives you a bad answer, resist the urge to ask "how do I give it more context?"

Ask instead: "What did I give it that it didn't need?"

That one question might lead you to a faster, cheaper, and more reliable system.

I am curious how others are handling this. Do you put a hard budget around your agent's context, or do you let it grow until the model's window becomes the limit? Tell me in the comments.

Further reading

Top comments (0)