DEV Community

Cover image for Your Reader Has a Context Window Too
Remus Lazar
Remus Lazar

Posted on Originally published at Medium

Your Reader Has a Context Window Too

AI made writing cheap. It did not make reading cheap. That gap is where most of my collaboration problems came from this year.

The message that proved his point

A while ago a colleague sent me a message I have been thinking about ever since.

He had counted the characters. His document was about 13,100 characters. My feedback on it was 17,100. So my feedback was 30% longer than the thing it was feedback on, and he found that absurd.

I did what I suspect a lot of people would do. I disagreed, and I explained why at length. I pointed out that resolving ambiguity is genuinely more expensive than creating it. I mentioned that the working session behind those 17,100 characters had produced roughly 126,800 characters of dialogue that I had read and processed, about ten times his original text, and that a raw character comparison therefore said nothing about the actual work.

Every word of that was true. It was also the single worst message I sent all year.

He said the text was too long. I answered with something much longer, about how much text I had produced. I proved his point better than he could have.

The asymmetry nobody planned for

Here is what actually happened to my workflow over the last three years, stated plainly. Producing text went from expensive to nearly free. Consuming text did not change at all.

That is the whole thing. Everything else is a consequence.

I can now take a messy pile of thoughts, structure them, check them for contradictions, get pushback on them, rewrite them three times, and end up with a dense, well-organised document in the time it used to take me to write a rough first draft. My throughput went up by something like an order of magnitude. It is the single biggest productivity change of my working life and I am not giving it back.

I should be honest about one thing, because it would be easy to read this as a story about AI. The long documents were not new. Digging through my Slack history, I found the same argument with a colleague in 2021, two years before any of these tools existed: too much context, too much reasoning, too long to read. I have always written this way. What used to hold it in check was not judgment. It was time. There were only so many hours in which to produce 17,000 characters of feedback, and that limit was quietly doing the reader's flow control for me. AI did not create the problem. It took the brake off a car I had been driving too fast for years.

But the person receiving that document is exactly as fast as they were in 2019. Same eyes, same working memory, same number of hours. Their side of the pipe did not widen by one bit.

We built a 10x transmitter and attached it to an unchanged receiver. In any other engineering context we would immediately recognise what happens next.

Receiver throttling

A wireframe human head in profile, labelled HUMAN RECEIVER. Streams of input tokens funnel into a gate labelled THROTTLING, then through a narrow LIMITED BANDWIDTH pipe into a CONTEXT WINDOW inside the head marked

Both diagrams made with an image model, labels mine.

If you have ever debugged a system where a fast producer feeds a slow consumer, you know the failure modes. The queue grows. Latency climbs. Eventually you either drop messages or the consumer stalls out entirely.

Human collaboration has the same failure modes, and they look like this.

Dropped messages. They skim. They read the first paragraph and the last one. They answer the easiest of your five questions and silently discard the other four.

Growing latency. Your document sits unread for three days, not because it is unimportant but because they cannot find a block of time large enough to process it.

Consumer stall. They stop engaging with your output entirely and start responding in single words. If you have ever had a long, careful message answered with ok, you have seen a stalled consumer.

The mistake I made for a long time was reading these signals as a lack of interest, or as sloppiness. They are not. They are flow control. A slow consumer that starts dropping packets is not being rude. It is protecting itself from a producer that is not respecting the link speed.

The reader's context window

The framing that finally made this click for me was borrowing the language of the models themselves.

A model has a context window: a hard limit on how much it can hold at once. Push past it and things fall out, usually the material in the middle, quietly, without an error message.

People have exactly the same constraint, and we have always known it. What is new is that we now have a vocabulary precise enough to reason about it. But the human version has two properties the model version does not.

First, tokens cost the reader something. For a model, filling the context is billed to whoever runs the inference. For a person, every token is paid in attention, and attention comes out of a fixed daily budget that also has to cover their own work. A colleague once told me that reading my output and catching up on the threads around it consumed roughly half the time he had available for the project that day. I had read that sentence before, but I had not really heard it. Half his day. For input.

Second, and this is the part I got wrong for years:

The window is not the bottleneck. The retrieval is.

The experiment that changed my mind

After the argument about the 30 percent, I rewrote the document. I put a compact decision table at the top, one row per open point. I tagged every point by type. I collapsed all the reasoning behind expandable sections.

I did not delete anything. I added the tagging and the cross-references, which means the new version was longer than the one he had complained about.

His reaction to it was one word: great.

Same content. More characters. Completely different reception.

So it was never the volume. It was the time to first useful token, how long he had to read before he found the thing he actually had to act on. In the first version that was several minutes. In the second it was about ten seconds. The reasoning was still all there, but it had moved from mandatory to available, and that turns out to be the entire difference.

Length is a bad proxy for cost. It is just the only one that is easy to measure, which is why we all reach for it. The real cost is how long until the reader can do something with this.

The compression tax

Which leads to the uncomfortable conclusion, and the reason I now think most AI writing advice is aimed at the wrong problem.

The advice is usually to write less. That is not right, and it is not even achievable, because the volume is a symptom. The right move is to spend part of the time AI saved you on compression, not on more output.

That is a real tax and it is the least fun part of the workflow. Generating a dense, thorough, well-argued document is now the easy half. Deciding what the reader has to do, putting that first, and demoting everything else to optional is the work, and it is work the tools do not do for you, because they do not know your reader.

Five rules, all learned by breaking them

  1. The ask goes in line one. What you need, by when, in what form. Everything else is context, and context is optional by definition.
  2. A reply to criticism must be shorter than the criticism. No exceptions. If I need more than five lines to respond to too long, I am not clarifying, I am defending.
  3. Never cite your own effort. How much work went into it is invisible to the reader and irrelevant to their decision. Mentioning it only invites them to price it.
  4. One message, one open question. A message with five asks gets answered on the easiest one.
  5. The ten second test. Open your own message and look for the action you need. If you cannot find it in ten seconds, restructure, do not just shorten.

The part that is actually hard

There is one more thing, and I do not have a clean rule for it.

The same colleague told me that reading my output felt strange in a way he struggled to name. Not wrong, the content was fine. But the voice was gone. He put it as: text has a personal note the way a voice does, and this one was somebody else's.

He was right, and it was not about the AI producing the sentences. It was that in optimising for structure and density, I had squeezed out every trace of what I actually thought about any of it. The document was correct and it was empty.

The cheap fix, which works better than it has any right to: two or three lines in your own voice on the front. What you think. What you are unsure about. What worries you. Ninety seconds of typing, no tooling, and it turns a report back into a message from a person.

The productivity gain is real. I am ten times faster at producing than I was. But that number is measured at the wrong end of the pipe, and I spent about six months celebrating it before I noticed that nobody at the other end had gotten any faster at receiving.

They still have not. They are not going to.


If you have solved this in your own team, I would genuinely like to hear how. Preferably briefly.

Top comments (0)