We use AI agents to optimize cloud spending. Then we give them a very expensive bag of context.
MCP can make FinOps agents more useful by giving them access to real cloud data, but additional tools and context can also increase token usage and AI costs. The goal is not maximum context, but the smallest relevant context that produces a useful result.
Imagine asking an AI agent to investigate your AWS bill. You want it to look at your real costs, inspect your infrastructure, and explain what changed. A generic answer about oversized EC2 instances is no longer enough.
This is where MCP starts to matter. The Model Context Protocol gives AI applications a standard way to connect to external tools and data. Anthropic introduced MCP, and OpenAI also supports remote MCP servers in its Responses API.
More context makes the agent more useful. It can work with your environment instead of guessing from general knowledge. But that extra context also has a cost.
More context makes agents smarter
An AI model already knows a lot about AWS. It can list common reasons for a sudden increase in EC2 spending. What it does not know is what happened inside your account yesterday.
MCP can give the agent access to that missing information. It can work with cost data, infrastructure details, tags, and usage patterns. Now it has something concrete to investigate.
The difference is important. Instead of asking what usually causes EC2 costs to rise, you can ask why your costs rose. The agent can inspect the environment and connect the change to real resources.
This is one reason MCP is interesting for FinOps. The FinOps Foundation discusses use cases such as cost analysis, anomaly investigation, and connecting financial data with infrastructure decisions. Better context moves the agent from general advice toward answers about your actual environment.
But there is a tradeoff. The model now has more information to process before it can give you that better answer. And that context is not free.
Context is not free
Tools need descriptions so the model knows how to use them. Those descriptions may include parameters, schemas, instructions, and other metadata. Depending on the architecture, some of that information becomes part of the model context.
That means more tools can mean more tokens. More retrieved data can mean more tokens too. For token-priced models, this can directly increase cost.
This creates an unusual FinOps loop. We add AI to find cloud waste and improve infrastructure decisions. Then the AI system creates its own usage and its own bill.
Congratulations, you may have created a small FinOps problem inside your FinOps solution. That does not mean MCP is inefficient. It means the AI layer needs cost control too.
The new problem also looks strangely familiar. Cloud costs grow when resources are consumed without enough attention to value. AI costs can grow for the same reason when every available tool and piece of context is treated as necessary.
Not every tool needs to be in the room
Imagine an agent with access to AWS billing, GitHub, Slack, Jira, Kubernetes, Terraform, Datadog, databases, and many other systems. All of these connections may be useful at some point. They are unlikely to be useful for every question.
If you ask why EC2 spending increased yesterday, the agent probably does not need Slack administration tools. It probably does not need every GitHub capability either. It needs the smallest useful set of tools and data for that task.
This is why tool discovery and selective loading matter. A system can keep many capabilities available without putting all of them into active context at once. Modern agent tooling is increasingly moving in this direction.
This is not just about saving tokens. Too much context can also make the agent's job harder. More available information does not automatically mean a better decision.
The useful distinction is between available context and active context. An agent may have access to twenty systems but need only two for the current task. Good context management keeps that difference intentional.
The goal is selective context
The answer is not to make agents less capable. The answer is to load capabilities when they are useful. That lets the agent stay powerful without making every request unnecessarily large.
This matters even more when FinOps agents move beyond analysis. They may recommend changes or eventually execute controlled actions. At that point, cost sits alongside permissions, policies, auditability, and clear limits.
The FinOps Foundation makes a similar point in its work on Agentic FinOps. Early adoption is more practical with narrow use cases and clear policy boundaries. Giving an agent more authority makes control more important, not less.
Where this gets practical
Not every FinOps problem needs AI reasoning. Some sources of waste are already obvious and well understood, such as forgotten idle VMs or development and staging machines that keep running at night and on weekends. Once this pattern is known, there is little value in spending tokens to rediscover the same problem again and again.
This is where a specialized tool can be more efficient. Idlefy can audit cloud environments to identify idle machines and estimate the cost of that unused compute, then automate when Dev and Stage VMs should be running. The problem is detected, the waste is reduced, and no additional AI context is needed to solve the same predictable issue every day.
Idlefy also uses MCP to make relevant cloud optimization data available when needed, while keeping repeatable actions automated and predictable.
So is it worth it?
Creating another cost does not automatically make the solution bad. The question is whether the new cost creates more value than it consumes. An agent that costs $100 and finds $10,000 of avoidable spend may be an excellent trade.
The problem starts when nobody measures that trade. If the same agent repeatedly spends money analyzing something that could be handled with a simple rule, the economics change quickly. At that point, the AI itself has become a FinOps workload.
And maybe that is the real lesson here. MCP can help AI solve FinOps problems, but it also gives FinOps a new resource to manage.
The question is not whether more context costs money. The question is whether that context is worth what we pay for it.
Further reading
If you work with MCP, FinOps, or AI infrastructure, I would be curious to hear how you are thinking about context cost vs. context value.





Top comments (2)
One measurement that changes the math here: in our own MCP server's agent benchmark (24 tasks, three passes each, on gpt-5.4-mini), 48% to 90% of the prompt tokens were served from the provider's prompt cache depending on the run, because the tool definitions were byte-identical on every turn. Cached tokens are billed at a fraction of the price, so the first FinOps rule for MCP is to keep the tool list stable: no timestamps, session ids or per-user text in descriptions, or every turn pays full price. The second is to trim the schemas, not only the number of tools: zod writes JavaScript's safe-integer range into every integer field, and dropping that and two other keywords that constrain nothing took 357 tokens off our default list without changing a tool.
Really good point. Thank you for adding this. Prompt caching can clearly change the cost picture, and your example shows that MCP efficiency depends not only on the number of tools, but also on stable tool definitions and lean schemas. That fits well with the broader point of the article that context has a cost, but that cost is heavily influenced by implementation quality.