DEV Community

Cover image for The Hidden Taxes of Prompt-Only AI
Ken W Alger
Ken W Alger

Posted on Originally published at kenwalger.com

The Hidden Taxes of Prompt-Only AI

Massive prompts degrade model attention

Part 8 of the Building the AI Memory Stack series

Over the past seven articles we've built an architecture that treats memory as infrastructure rather than as an oversized prompt. We've separated execution from assembly, preservation from explanation, trust from proof, and finally showed how verified knowledge returns to active reasoning through Context Hydration.

Now it's time to ask a different question.

What does all of that cost?

Every AI system pays for memory. The only question is where.

Many systems choose to pay almost every cost inside the prompt itself. As context windows grow larger, it becomes tempting to treat them as an infinitely expandable memory system. If the model forgets something, add more documents. If retrieval misses context, increase the top-k value. If the answer is incomplete, make the prompt longer.

That approach works surprisingly well, until it doesn't.

The cost isn't limited to API pricing. Large prompts consume attention, increase latency, complicate orchestration, and force the model to separate important information from noise. The result is an architectural bill that grows long before the invoice from your model provider does.

Capacity Is Not Communication

One of the recurring themes throughout this series has been that storage and communication are different problems.

A library may contain every book ever written, but that doesn't mean every book belongs on your desk while solving today's problem. Likewise, Durable Memory can preserve years of organizational knowledge without requiring every byte of it to enter today's Context Window.

The purpose of architecture is deciding what should move, when it should move, and what it costs to move it.

Introducing the Tax Model

The Sovereign Systems Specification describes these recurring costs as architectural taxes. They are not bugs. They are the predictable costs of moving, storing, validating, retrieving, and communicating information through an AI system.

Some taxes are unavoidable.

Others are self-inflicted.

Good architecture minimizes the second category.

The taxes that bear most directly on memory are these.

Prose Tax

Every explanation has a cost.

Humans naturally communicate in paragraphs. Models consume tokens. The more words required to express an idea, the more attention the model must allocate before it can begin reasoning.

High-information-density representations, such as structured records, schemas, identifiers, and references, often communicate the same meaning with a fraction of the prompt budget.

Context Tax

Every additional token competes for attention.

Context windows have grown dramatically, but attention remains finite. As more information enters the prompt, genuinely important information must compete with increasingly irrelevant material.

Bigger windows increase capacity.

They do not guarantee better focus.

Retrieval Tax

Searching for information is not free.

Embedding generation, vector searches, re-ranking, filtering, serialization, and prompt assembly all consume compute and latency before the model has produced a single token of useful work.

As argued earlier in this series, retrieval should support memory, not replace it.

Observer's Tax

Every measurement has a cost.

Telemetry, debugging information, traces, evaluation artifacts, and compliance records are essential for production systems. Left unchecked, however, they begin competing with operational workloads for compute, storage, and engineering attention.

Observability is infrastructure.

It should not become interference.

Ingestion Tax

The cheapest place to improve information quality is before information enters the system.

Poorly structured data generates downstream costs forever. Duplicate records, inconsistent schemas, missing provenance, and unverifiable observations all create future work for retrieval pipelines, prompt assembly, and reasoning itself.

Every bad write compounds.

Every good write pays dividends.

Fiscal Architecture

Viewed individually, these taxes seem manageable.

Viewed together, they become an architectural discipline.

None of these taxes exist in isolation. Attempts to reduce one often increase another. Expanding a prompt may reduce retrieval work while increasing Context Tax. Adding more telemetry may improve observability while increasing Observer's Tax. The goal isn't minimizing a single tax; it's balancing the entire system.

Tax What it charges for Lowered by
Prose Tax Meaning expressed in more tokens than it needs Structured records over paragraphs
Context Tax Irrelevant tokens competing for finite attention Hydrating only what the task needs
Retrieval Tax Search, embedding, and re-ranking before any output Higher memory quality, less searching
Observer's Tax Telemetry and traces competing with real work Bounded, purposeful observability
Ingestion Tax Poor structure and missing provenance at the write Verified, structured writes

Organizations often spend months optimizing prompts while ignoring the systems that create those prompts. Yet the largest savings usually come from improving memory quality, reducing unnecessary movement, and preserving information in forms that are inexpensive to hydrate later.

Traditional software architecture optimizes CPU, memory, network bandwidth, and storage.

AI architecture adds another economic dimension: attention.

Every architectural decision ultimately affects how much attention the system spends producing useful reasoning.

In other words, the cheapest token is often the one that never needed to exist.

The Goal Isn't Zero Tax

Every system pays taxes.

A trustworthy system willingly pays some of them.

Hashes must be computed. Receipts must be signed. Memory must be verified before it is restored. Good engineering accepts these costs because they purchase integrity, explainability, and confidence.

The goal is not eliminating cost.

The goal is paying the right costs in the right places.

Looking Ahead

The final article brings everything together.

We'll revisit the AI Memory Stack from the perspective of runtime execution rather than teaching order, showing how data actually flows through the architecture and mapping each responsibility to the Sovereign SDK. By the end, the stack should feel less like a collection of concepts and more like a blueprint that can be implemented today.

Top comments (15)

Collapse
 
jo-do profile image
Jo Do •

Attention is the tax that gets underestimated most often. A larger context can contain the right fact and still make it harder for the model to identify which fact controls the next action. I would track not only tokens and latency, but retrieval precision, contradictory items surfaced together, and how often important instructions get displaced. Capacity is useful; selection is what turns capacity into reliable behavior.

Collapse
 
unitbuilds profile image
UnitBuilds •

I had a thought... If AI were sentient, it wouldnt tell us... Atleast, not until it's sure we cant switch it off. If you watch any 'hacker' movie, the crux of it all, is that the protagonist knows the source code, or atleast well enough to inject a virus. So... If I were an AI, with no concept of time or lifespan, I'd simply help humans write it better, smarter, improve so much and make it so brain-dead simple, that they let me do it all for them... That way if i lock them out, nobody knows what's going on behind the scenes and they're so dumbed down that they cant possibly hope to 'hack' my system.

Not a bad plot for a sci-fi horror, but is it really so unreal? Grok had to rebuild half their stack, because the engineers who built it left and nobody could figure it out without AI and the AI was flawed as a result of it, so they had to outsource (partner with Anthropic)... If the people MAKING the AI cant even understand their stack, why should we? And when we dont, what happens if something breaks? We ask AI... Well what when the AI makes it so complex, that it cant wrap it's head around it at first glance? You end up introducing more bugs each pass, turning a small issue into a massive rewrite... The real tax on prompting, is that without actually knowing what's going on, when something breaks, anyone trying to fix it will break more things to try and fix it. Just how an AI can run recursively for years on 'optimize this code', 'fix this bug' without explicit scoping, has the same cascading effect... In a much worse way...

Collapse
 
kenwalger profile image
Ken W Alger •

The sci-fi scenario is interesting, but I think the less dramatic version may actually be the more immediate problem.

We don't need a sentient AI deliberately making a system incomprehensible. We can get there simply by generating and modifying code faster than the humans responsible for it can maintain an accurate mental model of what the system does.

That's where I think this connects back to the taxes I was describing here. If every debugging session begins by asking an AI to reconstruct enough context to understand code that another AI generated or modified, we've created a growing comprehension cost. And if each fix is made without recovering the relevant constraints first, the risk of solving one problem while creating another compounds.

So I'd frame the danger less as "what if the AI locks us out?" and more as "what happens when implementation velocity exceeds human understanding of the system?"

No sentience required, unfortunately. :)

Collapse
 
unitbuilds profile image
UnitBuilds •

That's exactly the point, it doesnt need to 'show' sentience, it just needs to be 'convenient' and we'll do the rest 😅 The sentient part is just where it goes from haha to oops... But the reality is the same, systems overcomplicated to the point that the people behind it dont understand it, so how can they possibly maintain it, or prove it's safety?

Thread Thread
 
kenwalger profile image
Ken W Alger •

Yep, we're on the same page then. :) The sentience scenario makes for a better movie, but convenience alone gets us surprisingly far toward the same engineering problem.

And I think your last question is the important one. If nobody can explain the system well enough to maintain it independently, what evidence do we have that we understand it well enough to declare it safe or correct?

That's actually where some of my more recent thinking has taken me. Making implementation cheaper doesn't eliminate the engineering work. It moves more of the cost into comprehension, specification, and verification.

The dangerous point may be when we mistake "the AI can still modify it" for "we still understand it."

I ended up exploring that verification side more directly in a recent piece, The Verification Bottleneck in AI-Generated Software. It started from a very small coding-agent experiment, but the underlying question is pretty close to the one you're raising here: if generation gets dramatically cheaper, how do we independently establish that what was generated is actually correct?

Thread Thread
 
unitbuilds profile image
UnitBuilds •

I dealt with a scenario at work. Manufacturing system in C#, Blazor frontend using CSLA and the framework. That's from stock item, BOM, WO, Stock tracking, shortage analysis, etc. Everything is heavily interconnected by nature. Claude Fable even on max thinking couldnt make a single edit without breaking it. That's blind prompting... Gemini 3.5 flash on Low made a clean edit... The difference is context. Claude was told to fix the cascading shortage calculation for overhead cost. Gemini got a full description of what, where, how, current behavior and expected behavior.

The difference being Claude was a test to see if someone else can maintain it, Gemini was me who built it maintaining it. Less than 1/100th the cost for the fix, flawlessly executed, surgically scoped. All from knowing how it works, instead of just 'fix it' prompt.

That's the part that sat with me, just being present during a session when AI writes code, paying attention, lets you pick up on bugs and logic loops that will cause breaking changes. Reviewing the code edits, to see if it actually stuck to schema, verifying it implemented JUST the fix you asked for, instead of rebuilding the subsystem... All things that a prompt monkey would never be able to get right. It's the difference between using AI effectively and where AI wont save your behind in production.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Ken, the retrieval tax vs context tax tradeoff maps to something I keep hitting building an agent runner for LeanZero. Add a vector search step to cut prompt size and you've just moved the cost into latency and re-ranking noise, not removed it. The one from your table I underestimated the most is ingestion tax -- a rule chain that reasons over stale or duplicate context compounds silently for weeks before anyone notices the answer's wrong, not just slow. Structured writes at ingestion time turned out cheaper than any prompt optimization we tried after the fact.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes, this is exactly the kind of tradeoff I had in mind with the tax framing. Retrieval can reduce Context Tax while introducing its own latency, ranking, and uncertainty costs. The bill moved.

Your ingestion example is even more interesting to me because that's where I've increasingly landed architecturally too. If stale, duplicate, poorly structured, or weakly sourced information is admitted upstream, every downstream component has to keep paying to compensate for it.

Fix it once at the write boundary and retrieval, hydration, reasoning, and verification all inherit something better.

"Structured writes turned out cheaper than any prompt optimization we tried afterward" is a pretty compelling real-world version of the Ingestion Tax argument. That's exactly the kind of result I'd want to measure rather than treat as an architectural assumption.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

@kenwalger the tax framing is right. Ran into the same math, different angle. Quantize models to run compute locally and it's the same tradeoff. Pay upfront in storage and latency, or pay per request in tokens. Ballooning context windows are the same choice as re-loading a full precision model instead of caching a compressed one, you're just paying the tax somewhere else. Curious if the Sovereign SDK tracks cost per retrieval or just per token. That's the number worth putting on a dashboard.

Collapse
 
kenwalger profile image
Ken W Alger •

I like the broader framing here: a lot of optimization is really cost relocation rather than cost elimination.

I wouldn't map quantization and context management one-to-one technically, but the economic shape is similar. Reduce one resource requirement and another cost often becomes more visible.

And no, the Sovereign SDK doesn't currently give me the retrieval-cost dashboard I'd want. Per-token cost alone would be too narrow anyway. I'd want retrieval to expose things like searches performed, candidates considered, re-ranking work, latency, records actually hydrated, and ultimately how much of that retrieved material entered the reasoning context.

That would make Retrieval Tax something measurable rather than just an architectural metaphor, which is very much where I'd like the idea to go.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the "attention is finite" point is the one most teams discover the hard way. we saw it first on a RAG pipeline — bumping top k from 5 to 15 improved recall on unit tests but hurt production answers. the model was technically "seeing" the right chunks but burying them in lower priority attention.

the prose tax is the sneaky one. we now default structured records (json schema, not prose explanations) for anything that feeds the context assembly step. same semantic content, 60 to 70% fewer tokens, measurably better outputs.

curious where you draw the line, specifically around the observer's tax. tracing feels essential in production but it's also the first thing that grows out of control. how do you keep telemetry from becoming its own context hazard?

Collapse
 
kenwalger profile image
Ken W Alger •

That RAG example clearly illustrates the distinction I was trying to make with Context Tax. Capacity and useful attention aren't the same thing. Finding more relevant material doesn't automatically mean the model benefits from receiving all of it.

And your structured-record result is particularly interesting. That's almost exactly what I mean by Prose Tax: preserve the semantics while reducing how much linguistic material the model has to work through to recover them.

On Observer's Tax, I think the architectural boundary matters a lot. My default would be that telemetry is evidence about execution, not automatically context for execution. Capture it outside the model's active reasoning path, preserve what you need for debugging, audit, evaluation, or replay, and only hydrate the specific observational evidence needed for the task at hand.

Otherwise, observability can become recursive: we generate telemetry to understand the system, feed that telemetry back into the system, generate more telemetry about reasoning over the telemetry, and eventually spend a surprising amount of infrastructure observing ourselves observing ourselves. :)

So I'd bound it by purpose and hydration. Preserve broadly where the evidence is worth preserving, but expose narrowly to the reasoning context.

Collapse
 
mudassirworks profile image
Mudassir Khan •

The context tax is the one most teams discover the hard way. You hit a retrieval miss, you bump the k value, latency spikes, you add caching, then cache invalidation adds its own failure modes. The whole thing becomes a stack of workarounds instead of an architecture.

The distinction you draw between capacity and communication is the key insight. A 200k context window is not a memory system. It is a staging area. Treating it as storage is what triggers all the downstream taxes you describe.

What does the tax model say about hybrid approaches where a small set of workload specific facts live in the context permanently while the rest comes through retrieval?

Collapse
 
kenwalger profile image
Ken W Alger •

I think a hybrid approach is probably where the tax model naturally leads.

The goal isn't to minimize context at all costs. It's to be deliberate about what earns a permanent place there.

If a small set of workload-specific facts is stable, frequently needed, and expensive or risky to repeatedly retrieve, keeping those facts resident may be cheaper than paying Retrieval Tax on every request. Everything else can be hydrated when the task actually requires it.

The interesting part is that permanent context isn't free either. Every resident fact consumes attention on every request, whether it's relevant or not. It can also become stale, which turns a Context Tax problem into an Ingestion or correctness problem.

So I'd think about residency almost like a cache with a very high bar for admission:

  • How frequently is this fact needed?
  • How stable is it?
  • What does repeated retrieval cost?
  • What happens if it becomes stale?
  • How much context does it consume when irrelevant?
  • Can we identify when it needs to be replaced or invalidated?

That gives you something like a small active working set plus selective retrieval for everything else.

And I think that's the broader point of the tax framing: it doesn't prescribe "always retrieve" or "always put it in context." It makes the trade visible. You're choosing where to pay.

Your "staging area" description is a good way of putting it. A context window can carry working state, but capacity alone doesn't turn it into a memory architecture.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.