DEV Community

Shrijith Venkatramana
Shrijith Venkatramana

Posted on AI-assisted

Gisting: Compressing LLM Agent Context

Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.


An AI agent can remember too much.

That sounds like a strange problem. More context should mean more information, so more context should mean a better agent.

In practice, the opposite can happen.

A coding agent may start with a system prompt, repository conventions, an issue, a few files, tool outputs, test failures, and then accumulate dozens of turns of its own reasoning. By turn 30, the agent may be carrying 30,000 tokens of history to answer a question that depends on 2,000 of them.

The engineering problem is therefore:

How do we preserve the useful state of a long context while throwing away most of the tokens?

One answer is gisting.

Gisting treats a long context as something that can be compressed into a much smaller representation and then reused.

The idea has an interesting history. In 2023, Jesse Mu, Xiang Lisa Li, and Noah Goodman published Learning to Compress Prompts with Gist Tokens at NeurIPS. Their motivating example was almost mundane: a fixed instruction prompt gets encoded again and again, even though its meaning has not changed. Their solution was to train the model itself to turn the prompt into a few internal "gist" tokens.

For agents, this idea becomes much more interesting.

1. The real problem is context, not memory

Consider a coding agent working on this task:

Fix the authentication bug in the payments service.
Enter fullscreen mode Exit fullscreen mode

The agent might accumulate:

System instructions                         1,500 tokens
Repository instructions                     2,000
Issue description                             800
Relevant source files                       4,500
Tool output                                 3,000
Turn 1                                       1,200
Turn 2                                       1,800
Turn 3                                       2,100
Turn 4                                       2,400
...
Turn 15                                      2,000
Enter fullscreen mode Exit fullscreen mode

Eventually:

Total context ~= 25,000 tokens
Enter fullscreen mode Exit fullscreen mode

Yet much of that history is redundant.

The agent does not need to remember:

"I opened auth.py."

"I then searched for validate_token."

"Then I inspected line 84."

"Then I ran pytest."

Enter fullscreen mode Exit fullscreen mode

What it really needs is something closer to:

Relevant state:
- Authentication uses JWT tokens.
- Token validation happens in auth.py:84.
- The failure occurs when expired tokens are refreshed.
- The refresh path bypasses validate_token().
- tests/test_auth.py reproduces the bug.
- A patch was attempted but failed because refresh_token() expects
  a normalized subject claim.
Enter fullscreen mode Exit fullscreen mode

The first representation is history.

The second is state.

Gisting is the attempt to move from the first representation toward the second.

This is a useful way to think about agent context:

Raw history
    |
    v
Compression
    |
    v
Compact representation
    |
    v
Future reasoning
Enter fullscreen mode Exit fullscreen mode

The goal is not to preserve every word.

The goal is to preserve the information required for future computation.

2. What gisting actually is

There are two ideas that are often mixed together.

Text summarization

A normal agent might ask an LLM:

Summarize the previous 20 turns in 800 tokens.
Enter fullscreen mode Exit fullscreen mode

The result is ordinary natural language:

The user is debugging JWT refresh behavior...
Enter fullscreen mode Exit fullscreen mode

This can be passed to almost any LLM.

Learned gist tokens

The original gisting paper does something different.

Suppose the original prompt is:

t = [t1, t2, ..., tn]
Enter fullscreen mode Exit fullscreen mode

The model inserts a small number of special tokens:

t1 t2 ... tn <G1> <G2> ... <Gk>
Enter fullscreen mode Exit fullscreen mode

The Transformer is trained with an attention mask that forces later tokens to access the earlier prompt only through the gist tokens.

Conceptually:

Original prompt
      |
      v
[G1] [G2] [G3] [G4]
      |
      v
Future input
Enter fullscreen mode Exit fullscreen mode

The gist tokens therefore become a bottleneck.

The model has to learn:

"What information from this prompt must survive so that I can perform the task later?"

The important detail is that these are internal model representations, not miniature English summaries.

The paper trained LLaMA-7B and FLAN-T5-XXL models and reported prompt compression of up to 26x, with reductions in FLOPs and storage while retaining similar output quality on its evaluations.

This is closer to learned memory than to summarization.

3. Why the attention mask matters

This is the clever part of the original method.

Without any restriction, the model can cheat.

Imagine:

Prompt:
"Translate the following text into French."

Gist tokens:
<G1> <G2>

Input:
"The cat is sleeping."
Enter fullscreen mode Exit fullscreen mode

If the later input can still directly attend to the original prompt, the model has no reason to encode anything into <G1> and <G2>.

So gisting changes the attention pattern.

Normally, later tokens can attend to previous tokens:

Prompt ----------------------> Future tokens
       \                     /
        \                   /
         Gist tokens ------>
Enter fullscreen mode Exit fullscreen mode

With gist masking:

Prompt ---> Gist tokens ---> Future tokens
                 ^
                 |
        compressed information
Enter fullscreen mode Exit fullscreen mode

The future part cannot directly look back at the original prompt.

That creates a training pressure:

If the future needs information,
the information must pass through the gist.
Enter fullscreen mode Exit fullscreen mode

This is essentially a learned information bottleneck.

The authors point out that the masking change is very small at the implementation level: the model architecture does not need to become a new neural architecture. The attention pattern is changed so the model learns the compression behavior during instruction tuning.

That distinction matters.

Gisting is not:

summary = model(prompt)
Enter fullscreen mode Exit fullscreen mode

It is:

gist = learned_compressor(prompt)
output = model(gist, new_input)
Enter fullscreen mode Exit fullscreen mode

where gist is represented inside the Transformer's activation space.

4. The economics are simple

The reason engineers should care about this is mathematical.

For a standard Transformer, self-attention over n input tokens has roughly:

O(n^2)
Enter fullscreen mode Exit fullscreen mode

attention complexity during the prefill phase.

During generation, cached keys and values change the picture. Each newly generated token still needs to attend across the existing context, so the attention work is roughly proportional to:

O(n)
Enter fullscreen mode Exit fullscreen mode

per generated token.

That means an agent with a huge context pays repeatedly for the same historical information.

Suppose:

Original context = 12,000 tokens
Compressed context = 1,000 tokens
Future generation = 300 tokens
Number of turns = 20
Enter fullscreen mode Exit fullscreen mode

Ignoring constants and other model costs, the context-attention work associated with generation is approximately:

Full history:

12,000 * 300 * 20
= 72,000,000

Compressed:

1,000 * 300 * 20
= 6,000,000
Enter fullscreen mode Exit fullscreen mode

So the difference is:

66,000,000
Enter fullscreen mode Exit fullscreen mode

context-token attention interactions.

That is a 12x reduction in the context length seen by every generated token.

There is a catch: compression itself costs compute.

So gisting only makes economic sense when the compressed representation is reused enough times.

Let:

Ccompress = cost of creating the gist

Cfull = cost of processing the full context per future turn

Cgist = cost of processing the compressed context per future turn

R = number of future reuses
Enter fullscreen mode Exit fullscreen mode

Gisting pays off when:

Ccompress + R * Cgist < R * Cfull
Enter fullscreen mode Exit fullscreen mode

or:

R > Ccompress / (Cfull - Cgist)
Enter fullscreen mode Exit fullscreen mode

This gives an important systems principle:

Compression is an amortization problem.

A 20-token system prompt used once does not need elaborate compression.

A 20,000-token agent history reused for another 40 turns is a different engineering problem.

The same logic applies to token billing.

Suppose compression reduces each future request by 11,000 input tokens and the agent makes 20 requests:

11,000 * 20
= 220,000 input tokens saved
Enter fullscreen mode Exit fullscreen mode

At an input price of P dollars per million tokens:

Savings = 0.22 * P dollars
Enter fullscreen mode Exit fullscreen mode

The exact dollar value depends on the model and pricing, but the scaling argument does not.

5. Where this becomes useful for agents

An agent has several different kinds of context.

They should not all be compressed in the same way.

Stable instructions

Examples:

Coding conventions
Security policies
Tool usage rules
Output format
Repository architecture rules
Enter fullscreen mode Exit fullscreen mode

These are excellent candidates for caching or learned compression because they are reused repeatedly.

Episodic history

Examples:

Previous tool calls
Previous failed approaches
Old intermediate reasoning
Past conversation turns
Enter fullscreen mode Exit fullscreen mode

This is where context gisting becomes useful.

The agent can periodically replace:

Turn 1
Turn 2
Turn 3
...
Turn 18
Enter fullscreen mode Exit fullscreen mode

with:

Agent state as of turn 18
Enter fullscreen mode Exit fullscreen mode

Exact evidence

This category is different:

Database rows
Source-code lines
Exception messages
API responses
Financial values
Configuration values
Enter fullscreen mode Exit fullscreen mode

Compressing these aggressively can be dangerous.

An agent may need the exact string:

"expected audience=payments-api"
Enter fullscreen mode Exit fullscreen mode

A summary such as:

"The JWT audience configuration was incorrect."
Enter fullscreen mode Exit fullscreen mode

is semantically close but operationally weaker.

This suggests a useful architecture:

             Agent Context
                  |
       +----------+----------+
       |          |          |
    Stable      Episodic    Exact
    context      state      evidence
       |          |          |
    Cache       Gist/       Retrieve
                 summary     on demand
Enter fullscreen mode Exit fullscreen mode

In other words, gisting should usually reduce the working context, not become the only copy of the information.

For a coding agent, a sensible system might look like:

                    +----------------+
                    | Agent          |
                    +-------+--------+
                            |
             +--------------+--------------+
             |              |              |
             v              v              v
        System rules    Agent gist     Retrieved code
             |              |              |
             +--------------+--------------+
                            |
                            v
                       LLM context
Enter fullscreen mode Exit fullscreen mode

The gist remembers the trajectory.

Retrieval restores exact facts.

The current user request provides the immediate objective.

6. The compression bottleneck can destroy information

There is a fundamental problem with any lossy context compression:

The compressor does not know what the future question will be.

Suppose an agent summarizes a debugging session:

We fixed the OAuth callback bug.
Enter fullscreen mode Exit fullscreen mode

That sounds useful.

Three turns later the user asks:

Which exact redirect URI caused the failure?
Enter fullscreen mode Exit fullscreen mode

The answer may already have disappeared from the compressed representation.

This is why context compression behaves differently from ordinary text summarization.

A summary can preserve semantic meaning while destroying future utility.

Recent evaluations of gist-token compression found several failure patterns, including information being lost at segment boundaries, information that is locally surprising being dropped, and information being degraded across multiple compression stages. The 2025 ACL study also found that gist-based compression can perform well on some retrieval and long-document tasks while struggling on tasks requiring precise recall.

This leads to an important engineering rule:

Do not ask:

"What is important?"

Ask:

"What information might future computation require?"
Enter fullscreen mode Exit fullscreen mode

Those are different questions.

For an agent, one practical response is to preserve certain classes of information explicitly:

Goals
Decisions
Constraints
Open questions
Identifiers
Errors
Exact values
Files changed
Commands executed
Artifacts produced
Enter fullscreen mode Exit fullscreen mode

Then allow everything else to enter the compressed state.

A useful pattern is:

Gist:
    high-level state and reasoning

Retrieval store:
    exact historical evidence

Current context:
    current task and retrieved evidence
Enter fullscreen mode Exit fullscreen mode

That is much safer than:

100,000 tokens
       |
       v
800-token summary
       |
       v
delete everything
Enter fullscreen mode Exit fullscreen mode

7. Gisting changes the architecture of an agent

Once context is treated as a compressible state, an agent starts to look less like a chatbot and more like a state machine.

Instead of:

history = history + new_turn
Enter fullscreen mode Exit fullscreen mode

you can think in terms of:

state_(t+1) = compress(state_t, observation_t)
Enter fullscreen mode Exit fullscreen mode

and later:

action_t = model(state_t, current_observation_t)
Enter fullscreen mode Exit fullscreen mode

This is a much more scalable abstraction.

For example:

             ┌──────────────┐
             │ Current task │
             └──────┬───────┘
                    |
                    v
             ┌──────────────┐
             │ Agent model  │
             └──────┬───────┘
                    |
                    v
             ┌──────────────┐
             │ Tool actions │
             └──────┬───────┘
                    |
                    v
             ┌──────────────┐
             │ Observations │
             └──────┬───────┘
                    |
                    v
             ┌──────────────┐
             │   Gisting    │
             └──────┬───────┘
                    |
                    v
             ┌──────────────┐
             │ Compact state│
             └──────────────┘
                    |
                    └──────> next turn
Enter fullscreen mode Exit fullscreen mode

This is the deeper idea behind gisting.

The agent does not need to carry its entire past forward.

It needs to carry forward a representation of its past that remains useful for future actions.

That is a very different mental model from simply "increase the context window."

There is also an important practical limitation: the original learned gist-token approach produces model-specific internal representations. An application using a closed LLM API generally cannot take those hidden activations and feed them back later as arbitrary input. For API-based systems, natural-language compression methods such as prompt compression are therefore often easier to deploy, while true gist-token methods are more natural when you control the model and inference stack. Work on long-context compression has also shown that gist-style methods can degrade as the context becomes longer, so compression ratio alone is not a sufficient metric.

The interesting engineering question is therefore not:

"How large should the context window be?"

It is:

"What is the smallest representation of the agent's past that preserves the information needed for future computation?"

That is the problem gisting puts on the table.

And once agents become long-running systems, that problem starts to look less like prompt engineering and more like designing a memory hierarchy.

Conclusion

LLM agents are often built around an implicit assumption:

More history = more intelligence
Enter fullscreen mode Exit fullscreen mode

A better model may be:

More relevant state = more intelligence
Enter fullscreen mode Exit fullscreen mode

Gisting is one attempt to make that transition.

It treats context as a representation that can be compressed, cached, and reused rather than as an ever-growing transcript.

The next generation of agents may therefore have something resembling:

working memory
episodic memory
compressed state
retrievable evidence
Enter fullscreen mode Exit fullscreen mode

rather than one giant conversation buffer.

That raises an interesting question for anyone building agents:

What information should an agent be allowed to forget?



Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production reliable and secure without slowing you down.

I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.

Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.

Spend code review effort where business risk is highest — not spread evenly across every diff.

⭐ Star it on GitHub:

GitHub logo HexmosTech / LiveReview

Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview

gitleaks.yml osv-scanner.yml govulncheck.yml semgrep.yml dependabot-enabled mcp-testcases.yml

LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.

blast-radius-demo.mp4

LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
















The exact math, not a black box Visualize blast radius at a glance Every factor that feeds the score

How does Blast Radius scoring work? (a more technical explanation)

Here's the goal:

  • A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
  • A 300-line UI change in one file, fully covered by…




Click below to try LiveReview with your codebase:

LiveReview Banner

Top comments (0)