Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.
An AI agent can remember too much.
That sounds like a strange problem. More context should mean more information, so more context should mean a better agent.
In practice, the opposite can happen.
A coding agent may start with a system prompt, repository conventions, an issue, a few files, tool outputs, test failures, and then accumulate dozens of turns of its own reasoning. By turn 30, the agent may be carrying 30,000 tokens of history to answer a question that depends on 2,000 of them.
The engineering problem is therefore:
How do we preserve the useful state of a long context while throwing away most of the tokens?
One answer is gisting.
Gisting treats a long context as something that can be compressed into a much smaller representation and then reused.
The idea has an interesting history. In 2023, Jesse Mu, Xiang Lisa Li, and Noah Goodman published Learning to Compress Prompts with Gist Tokens at NeurIPS. Their motivating example was almost mundane: a fixed instruction prompt gets encoded again and again, even though its meaning has not changed. Their solution was to train the model itself to turn the prompt into a few internal "gist" tokens.
For agents, this idea becomes much more interesting.
1. The real problem is context, not memory
Consider a coding agent working on this task:
Fix the authentication bug in the payments service.
The agent might accumulate:
System instructions 1,500 tokens
Repository instructions 2,000
Issue description 800
Relevant source files 4,500
Tool output 3,000
Turn 1 1,200
Turn 2 1,800
Turn 3 2,100
Turn 4 2,400
...
Turn 15 2,000
Eventually:
Total context ~= 25,000 tokens
Yet much of that history is redundant.
The agent does not need to remember:
"I opened auth.py."
"I then searched for validate_token."
"Then I inspected line 84."
"Then I ran pytest."
What it really needs is something closer to:
Relevant state:
- Authentication uses JWT tokens.
- Token validation happens in auth.py:84.
- The failure occurs when expired tokens are refreshed.
- The refresh path bypasses validate_token().
- tests/test_auth.py reproduces the bug.
- A patch was attempted but failed because refresh_token() expects
a normalized subject claim.
The first representation is history.
The second is state.
Gisting is the attempt to move from the first representation toward the second.
This is a useful way to think about agent context:
Raw history
|
v
Compression
|
v
Compact representation
|
v
Future reasoning
The goal is not to preserve every word.
The goal is to preserve the information required for future computation.
2. What gisting actually is
There are two ideas that are often mixed together.
Text summarization
A normal agent might ask an LLM:
Summarize the previous 20 turns in 800 tokens.
The result is ordinary natural language:
The user is debugging JWT refresh behavior...
This can be passed to almost any LLM.
Learned gist tokens
The original gisting paper does something different.
Suppose the original prompt is:
t = [t1, t2, ..., tn]
The model inserts a small number of special tokens:
t1 t2 ... tn <G1> <G2> ... <Gk>
The Transformer is trained with an attention mask that forces later tokens to access the earlier prompt only through the gist tokens.
Conceptually:
Original prompt
|
v
[G1] [G2] [G3] [G4]
|
v
Future input
The gist tokens therefore become a bottleneck.
The model has to learn:
"What information from this prompt must survive so that I can perform the task later?"
The important detail is that these are internal model representations, not miniature English summaries.
The paper trained LLaMA-7B and FLAN-T5-XXL models and reported prompt compression of up to 26x, with reductions in FLOPs and storage while retaining similar output quality on its evaluations.
This is closer to learned memory than to summarization.
3. Why the attention mask matters
This is the clever part of the original method.
Without any restriction, the model can cheat.
Imagine:
Prompt:
"Translate the following text into French."
Gist tokens:
<G1> <G2>
Input:
"The cat is sleeping."
If the later input can still directly attend to the original prompt, the model has no reason to encode anything into <G1> and <G2>.
So gisting changes the attention pattern.
Normally, later tokens can attend to previous tokens:
Prompt ----------------------> Future tokens
\ /
\ /
Gist tokens ------>
With gist masking:
Prompt ---> Gist tokens ---> Future tokens
^
|
compressed information
The future part cannot directly look back at the original prompt.
That creates a training pressure:
If the future needs information,
the information must pass through the gist.
This is essentially a learned information bottleneck.
The authors point out that the masking change is very small at the implementation level: the model architecture does not need to become a new neural architecture. The attention pattern is changed so the model learns the compression behavior during instruction tuning.
That distinction matters.
Gisting is not:
summary = model(prompt)
It is:
gist = learned_compressor(prompt)
output = model(gist, new_input)
where gist is represented inside the Transformer's activation space.
4. The economics are simple
The reason engineers should care about this is mathematical.
For a standard Transformer, self-attention over n input tokens has roughly:
O(n^2)
attention complexity during the prefill phase.
During generation, cached keys and values change the picture. Each newly generated token still needs to attend across the existing context, so the attention work is roughly proportional to:
O(n)
per generated token.
That means an agent with a huge context pays repeatedly for the same historical information.
Suppose:
Original context = 12,000 tokens
Compressed context = 1,000 tokens
Future generation = 300 tokens
Number of turns = 20
Ignoring constants and other model costs, the context-attention work associated with generation is approximately:
Full history:
12,000 * 300 * 20
= 72,000,000
Compressed:
1,000 * 300 * 20
= 6,000,000
So the difference is:
66,000,000
context-token attention interactions.
That is a 12x reduction in the context length seen by every generated token.
There is a catch: compression itself costs compute.
So gisting only makes economic sense when the compressed representation is reused enough times.
Let:
Ccompress = cost of creating the gist
Cfull = cost of processing the full context per future turn
Cgist = cost of processing the compressed context per future turn
R = number of future reuses
Gisting pays off when:
Ccompress + R * Cgist < R * Cfull
or:
R > Ccompress / (Cfull - Cgist)
This gives an important systems principle:
Compression is an amortization problem.
A 20-token system prompt used once does not need elaborate compression.
A 20,000-token agent history reused for another 40 turns is a different engineering problem.
The same logic applies to token billing.
Suppose compression reduces each future request by 11,000 input tokens and the agent makes 20 requests:
11,000 * 20
= 220,000 input tokens saved
At an input price of P dollars per million tokens:
Savings = 0.22 * P dollars
The exact dollar value depends on the model and pricing, but the scaling argument does not.
5. Where this becomes useful for agents
An agent has several different kinds of context.
They should not all be compressed in the same way.
Stable instructions
Examples:
Coding conventions
Security policies
Tool usage rules
Output format
Repository architecture rules
These are excellent candidates for caching or learned compression because they are reused repeatedly.
Episodic history
Examples:
Previous tool calls
Previous failed approaches
Old intermediate reasoning
Past conversation turns
This is where context gisting becomes useful.
The agent can periodically replace:
Turn 1
Turn 2
Turn 3
...
Turn 18
with:
Agent state as of turn 18
Exact evidence
This category is different:
Database rows
Source-code lines
Exception messages
API responses
Financial values
Configuration values
Compressing these aggressively can be dangerous.
An agent may need the exact string:
"expected audience=payments-api"
A summary such as:
"The JWT audience configuration was incorrect."
is semantically close but operationally weaker.
This suggests a useful architecture:
Agent Context
|
+----------+----------+
| | |
Stable Episodic Exact
context state evidence
| | |
Cache Gist/ Retrieve
summary on demand
In other words, gisting should usually reduce the working context, not become the only copy of the information.
For a coding agent, a sensible system might look like:
+----------------+
| Agent |
+-------+--------+
|
+--------------+--------------+
| | |
v v v
System rules Agent gist Retrieved code
| | |
+--------------+--------------+
|
v
LLM context
The gist remembers the trajectory.
Retrieval restores exact facts.
The current user request provides the immediate objective.
6. The compression bottleneck can destroy information
There is a fundamental problem with any lossy context compression:
The compressor does not know what the future question will be.
Suppose an agent summarizes a debugging session:
We fixed the OAuth callback bug.
That sounds useful.
Three turns later the user asks:
Which exact redirect URI caused the failure?
The answer may already have disappeared from the compressed representation.
This is why context compression behaves differently from ordinary text summarization.
A summary can preserve semantic meaning while destroying future utility.
Recent evaluations of gist-token compression found several failure patterns, including information being lost at segment boundaries, information that is locally surprising being dropped, and information being degraded across multiple compression stages. The 2025 ACL study also found that gist-based compression can perform well on some retrieval and long-document tasks while struggling on tasks requiring precise recall.
This leads to an important engineering rule:
Do not ask:
"What is important?"
Ask:
"What information might future computation require?"
Those are different questions.
For an agent, one practical response is to preserve certain classes of information explicitly:
Goals
Decisions
Constraints
Open questions
Identifiers
Errors
Exact values
Files changed
Commands executed
Artifacts produced
Then allow everything else to enter the compressed state.
A useful pattern is:
Gist:
high-level state and reasoning
Retrieval store:
exact historical evidence
Current context:
current task and retrieved evidence
That is much safer than:
100,000 tokens
|
v
800-token summary
|
v
delete everything
7. Gisting changes the architecture of an agent
Once context is treated as a compressible state, an agent starts to look less like a chatbot and more like a state machine.
Instead of:
history = history + new_turn
you can think in terms of:
state_(t+1) = compress(state_t, observation_t)
and later:
action_t = model(state_t, current_observation_t)
This is a much more scalable abstraction.
For example:
┌──────────────┐
│ Current task │
└──────┬───────┘
|
v
┌──────────────┐
│ Agent model │
└──────┬───────┘
|
v
┌──────────────┐
│ Tool actions │
└──────┬───────┘
|
v
┌──────────────┐
│ Observations │
└──────┬───────┘
|
v
┌──────────────┐
│ Gisting │
└──────┬───────┘
|
v
┌──────────────┐
│ Compact state│
└──────────────┘
|
└──────> next turn
This is the deeper idea behind gisting.
The agent does not need to carry its entire past forward.
It needs to carry forward a representation of its past that remains useful for future actions.
That is a very different mental model from simply "increase the context window."
There is also an important practical limitation: the original learned gist-token approach produces model-specific internal representations. An application using a closed LLM API generally cannot take those hidden activations and feed them back later as arbitrary input. For API-based systems, natural-language compression methods such as prompt compression are therefore often easier to deploy, while true gist-token methods are more natural when you control the model and inference stack. Work on long-context compression has also shown that gist-style methods can degrade as the context becomes longer, so compression ratio alone is not a sufficient metric.
The interesting engineering question is therefore not:
"How large should the context window be?"
It is:
"What is the smallest representation of the agent's past that preserves the information needed for future computation?"
That is the problem gisting puts on the table.
And once agents become long-running systems, that problem starts to look less like prompt engineering and more like designing a memory hierarchy.
Conclusion
LLM agents are often built around an implicit assumption:
More history = more intelligence
A better model may be:
More relevant state = more intelligence
Gisting is one attempt to make that transition.
It treats context as a representation that can be compressed, cached, and reused rather than as an ever-growing transcript.
The next generation of agents may therefore have something resembling:
working memory
episodic memory
compressed state
retrievable evidence
rather than one giant conversation buffer.
That raises an interesting question for anyone building agents:
What information should an agent be allowed to forget?
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production reliable and secure without slowing you down.
I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.
Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.
Spend code review effort where business risk is highest — not spread evenly across every diff.
⭐ Star it on GitHub:
HexmosTech
/
LiveReview
Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.
blast-radius-demo.mp4
LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
Here's the goal:
- A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
- A 300-line UI change in one file, fully covered by…
Click below to try LiveReview with your codebase:





Top comments (0)