Short answer: Context window management techniques come down to five controls before every model call:
compress the input, split anything too long into segments, truncate pipeline context at a fixed
cap, dedup overlapping passages, and enforce a monthly quota. SmartGate exposes all five as MCP
tools — the default compression ratio is 0.5, and the Free plan meters 2M tokens a month at 120
MCP requests per minute per key.
Key takeaways
-
Compression is a ratio, and the caller wins.
smart_context_gatedefaults to 0.5, but an explicit ratio in the request overrides the team default; the precedence lives inresolve_compress_ratio. -
Split before you compress.
splitForPlaygroundCompresscaps total input, hard-splits oversized blocks, and reports partial coverage instead of silently dropping a tail. -
Truncate only where a cap exists.
_truncate_pipeline_textreturns text untouched when no per-template cap is set, and otherwise cuts at exactly the cap. -
Dedup is a threshold, not a delete.
smart_dedupships at a 0.9 threshold, andfind_representativereranks survivors for diversity. -
Bound the candidate set or you pay for it.
compute_candidate_limitderives the rerank budget — 10% of the corpus, floor 100, ceiling 1,000. -
Memory and quotas are both team-scoped.
smart_memorymoves history out of the prompt with add/search/get/delete, andcheckQuotareturns oneexceededdecision per month. - Do this before your first loop: set the compression ratio and the monthly cap, then connect one MCP client and start free.
The short version for whoever signs the invoice
A context window is a budget, and an agent spends it every turn whether or not it learns
anything. The expensive habits are boring ones: re-sending the same fetched page, pasting a whole
issue thread to answer one question, keeping full history because trimming it feels unsafe. That
is plumbing, not a model problem, and it is fixable before the model is called.
SmartGate is an MCP-native algorithm gateway for token control, traffic shaping, and agent audit
— not a model host. You keep your agent and its host model, and connect the gateway as one more
MCP server at the stateless endpoint (https://smartgate.network/api/mcp, Streamable HTTP POST).
What changes is that the tool side of the loop — fetch, search, compress, dedup, budget, memory,
pipelines — is measured.
Free gives 2M tokens/month, all seven tools (smart_fetch, smart_search,
smart_context_gate, smart_dedup, smart_budget_guard, smart_memory, smart_pipe), 120
MCP requests/min per key, and 7-day logs, no card. Pro starts at $18/month ($5 first
month), 300 req/min/key, 30-day logs, capped near $36/month; Teams starts at $55/month
with 600 req/min/key and 90-day logs; Enterprise is contract-based at 1,200 req/min/key. The
billing line is "Pay for the platform. Share only when you save." — the share begins only after
$15 of measured savings (pricing).
What the research says about long context
There is a decent pile of published work on long context, and it points the same way these twelve
code paths do. Anthropic's context-engineering write-up frames the job as curating the smallest
set of high-signal tokens rather than filling the window
(Anthropic).
Chroma's Context Rot study reports reliability degrading as input length grows, even for an
unchanged task (Chroma Research). Lost in the
Middle showed that information buried mid-input is used less reliably than the same information
at the edges (arXiv:2307.03172), LLMLingua showed prompt
compression preserving task performance while cutting tokens
(arXiv:2310.05736), and MemGPT is the canonical argument for
moving history into an external memory tier (arXiv:2310.08560).
Two conclusions follow, both operational. Extra tokens dilute attention and raise cost at the same
time, so more context is not more understanding; and the fix is mechanical — compress, segment,
truncate, dedup, offload — which is why it belongs in a gateway every tool call passes through
rather than in each agent's prompt template. MCP defines tools and resources a client discovers
per session (MCP specification), so a
gateway can sit in that path without any agent being rewritten. How that gateway fits a wider
platform, and which controls belong in it rather than in each client, is covered in
Enterprise AI Gateway Architecture Best Practices.
smart_context_gate: compress the context window before the model call
Compression is the cheapest lever in the stack. The handler takes the long text, a target ratio,
and an optional purpose that pre-filters paragraphs toward the current goal, then dispatches
through the gateway's audit wrapper:
# backend/smartgate/api/mcp.py — source lines 150–174 (smart_context_gate)
@server.tool(
name="smart_context_gate",
description=TOOL_DESCRIPTIONS["smart_context_gate"],
annotations=tool_annotations("smart_context_gate"),
)
async def smart_context_gate(
text: str = Field(description="Long text to compress before the host LLM call."),
ratio: float = Field(
default=0.5,
description="Target compression ratio (e.g. 0.3–0.7).",
),
purpose: str | None = Field(
default=None,
description="Optional goal to pre-filter paragraphs (step intent, user query).",
),
) -> str:
_, registry = _app_state()
module = registry.get("context_gate")
ctx = _tool_ctx()
return await _run_with_audit(
"compress",
ctx,
module.process(ctx, text=text, ratio=ratio, purpose=purpose),
{"ratio": ratio, "purpose": purpose},
)
Three details matter. The ratio defaults to 0.5 — a target, not a promise, with 0.3–0.7 as
the range where compression stops being free. The purpose argument makes compression
goal-aware, keeping paragraphs that serve the current question instead of the first half of the
document. And the call is audited as compress with the ratio recorded, so a thin answer can be
traced back to a number.
resolve_compress_ratio: who decides the effective compression ratio
One setting, three sources, and an explicit order of precedence:
# backend/smartgate/core/effective_settings.py — source lines 69–74 (resolve_compress_ratio)
def resolve_compress_ratio(request_state: Any, body_ratio: float | None) -> float:
if body_ratio is not None:
return body_ratio
es = getattr(request_state, "effective_settings", None) or {}
defaults = es.get("defaults") or {}
return float(defaults.get("compress_ratio", 0.5))
A ratio supplied in the request body wins, because the caller knows what this specific call is
for. Otherwise the team's effective settings supply the default (compress_ratio, again 0.5),
so a team tunes compression once rather than in every prompt. What the function avoids is the
pattern that makes tuning impossible: a hardcoded ratio buried in a call chain. If compression
output looks wrong, this is the function that says which of the three sources to change.
splitForPlaygroundCompress: split long input before compression
You cannot compress a large page as one blob, and the failure mode of trying is a silent tail
truncation. So this function splits first: it caps the input at a maximum total size and sets a
partialCoverage flag when anything was dropped, splits the capped text into blocks, and
hard-splits any block still larger than the segment limit:
# lib/smartgate/playground-content.ts — source lines 113–166 (splitForPlaygroundCompress)
function splitForPlaygroundCompress(text: string): PlaygroundSplitResult {
const totalChars = text.length;
const capped = text.slice(0, PLAYGROUND_MAX_TOTAL_CHARS);
const partialCoverage = totalChars > PLAYGROUND_MAX_TOTAL_CHARS;
let blocks = splitIntoBlocks(capped);
let hardSplit = false;
const expanded: string[] = [];
for (const block of blocks) {
if (block.length > PLAYGROUND_SEGMENT_CHARS) {
hardSplit = true;
expanded.push(...hardSplitBlock(block, PLAYGROUND_SEGMENT_CHARS));
} else {
expanded.push(block);
}
}
blocks = expanded;
const segments: string[] = [];
let buf = "";
const flushSegment = () => {
if (buf.trim()) {
segments.push(buf.trimEnd());
buf = "";
}
};
for (const block of blocks) {
const candidate = buf ? `${buf}\n\n${block}` : block;
if (candidate.length <= PLAYGROUND_SEGMENT_CHARS) {
buf = candidate;
} else {
flushSegment();
buf = block;
}
}
flushSegment();
if (segments.length === 0 && capped.trim()) {
segments.push(capped.trim());
}
const processedChars = segments.reduce((sum, s) => sum + s.length, 0);
return {
segments,
processedChars,
totalChars,
hardSplit,
partialCoverage,
};
}
The packing loop accumulates blocks in a buffer until the next block would overflow the segment,
then flushes and starts again. The return value is the caller's receipt — segments to compress,
processedChars against totalChars, plus hardSplit and partialCoverage as two honesty flags.
flushSegment: when the compressor flushes a segment
Every packing loop needs one boundary rule, and this is it:
# lib/smartgate/playground-content.ts — source lines 135–140 (flushSegment)
const flushSegment = () => {
if (buf.trim()) {
segments.push(buf.trimEnd());
buf = "";
}
};
A segment is emitted only when the buffer holds non-whitespace, trailing whitespace is trimmed
before the push, and the buffer resets so the next segment starts clean. Two properties matter
downstream: nothing empty is ever emitted, so no compressor call is spent on whitespace, and
trimming happens at the boundary rather than per block, so a segment ending mid-sentence keeps its
internal newlines. The helper closes over the surrounding buffer, which is why it lives inside the
function it serves.
_truncate_pipeline_text: truncate pipeline text when the budget is tight
Truncation is the most primitive control, and this is the honest version of it: per-template caps,
no cap means no cut, and a cut means a cut rather than a summary.
# backend/smartgate/core/pipeline.py — source lines 176–180 (_truncate_pipeline_text)
def _truncate_pipeline_text(text: str, pipeline_template: str) -> str:
cap = PIPELINE_CONTEXT_GATE_MAX_CHARS.get(pipeline_template or "")
if not cap or len(text) <= cap:
return text
return text[:cap]
The lookup is keyed by pipeline template, so research and read runs can carry different ceilings,
and an unconfigured template passes text through instead of inventing a default. When a cap does
apply, the result is text[:cap] — the beginning of the text and nothing else. That bluntness is
the point: deterministic, free, and the last line of defence before a pipeline stage hands an
oversized blob to a model.
smart_dedup: semantic dedup for overlapping context passages
Agents rarely overflow a window with one huge document. They overflow it by assembling the same
document repeatedly: the same changelog from three searches, the same paragraph from two fetches,
a summary and its source.
# backend/smartgate/api/mcp.py — source lines 176–196 (smart_dedup)
@server.tool(
name="smart_dedup",
description=TOOL_DESCRIPTIONS["smart_dedup"],
annotations=tool_annotations("smart_dedup"),
)
async def smart_dedup(
texts: list[str] = Field(description="List of text passages to deduplicate."),
threshold: float = Field(
default=0.9,
description="Similarity threshold (0.0–1.0); higher keeps fewer duplicates.",
),
) -> str:
_, registry = _app_state()
module = registry.get("dedup")
ctx = _tool_ctx()
return await _run_with_audit(
"dedup",
ctx,
module.process(ctx, texts=texts, threshold=threshold),
{"threshold": threshold},
)
The tool takes a list of passages and a similarity threshold that defaults to 0.9, and its own
description states the trade-off: higher keeps fewer duplicates. That phrasing matters, because
0.9 means near-identical rather than merely related — lowering it starts discarding passages that
are only topically adjacent. The call is audited as dedup with the threshold recorded, which is
the field you want when two similar paragraphs both survived.
Dedup at this layer is what keeps an advanced RAG architecture
affordable: a four-stage retrieval stack pays for every passage it carries forward, and the
near-duplicates are the cheapest ones to drop.
find_representative: the representative passage in a duplicate cluster
Removing exact duplicates is the easy half. The hard half is what to keep when ten passages cover
the same ground, because centrality and redundancy look alike:
# backend/smartgate/modules/dedup/algorithm.py — source lines 326–352 (find_representative)
def find_representative(
self,
records: Sequence[Record],
selection_size: int = 10,
candidate_limit: int | Literal["auto"] = "auto",
diversity: float = 0.5,
strategy: Strategy | str = Strategy.MMR,
) -> FilterResult:
"""
Find representative samples from a given set of records against the fitted index.
First, the records are ranked using average similarity.
Then, the top candidates are re-ranked using Maximal Marginal Relevance (MMR)
to select a diverse set of representatives.
:param records: The records to rank and select representatives from.
:param selection_size: Number of representatives to select.
:param candidate_limit: Number of top candidates to consider for diversity reranking.
Defaults to "auto", which calculates the limit based on the total number of records.
:param diversity: Trade-off between diversity (1.0) and relevance (0.0). Default is 0.5.
:param strategy: Diversification strategy (MMR, MSD, DPP, COVER, SSD). Default is MMR.
:return: A FilterResult with the diversified candidates.
"""
ranking = self._rank_by_average_similarity(records)
if candidate_limit == "auto":
candidate_limit = compute_candidate_limit(total=len(ranking.selected), selection_size=selection_size)
return self._diversify(ranking, candidate_limit, selection_size, diversity, strategy)
Records are first ranked by average similarity to the fitted index, producing a candidate pool;
the pool is then re-ranked with Maximal Marginal Relevance to select passages that are relevant
and mutually diverse. diversity sits at 0.5 by default — the midpoint between relevance and
diversity — with selection_size at 10 representatives. The strategy is swappable (MMR, MSD, DPP,
COVER, SSD), and candidate_limit: auto defers to the next function rather than guessing a pool
size inline.
compute_candidate_limit: how many dedup candidates to consider
Diversity reranking is expensive work, so the candidate pool is a cost decision:
# backend/smartgate/modules/dedup/utils.py — source lines 88–113 (compute_candidate_limit)
def compute_candidate_limit(
total: int,
selection_size: int,
fraction: float = 0.1,
min_candidates: int = 100,
max_candidates: int = 1000,
) -> int:
"""
Compute the 'auto' candidate limit based on the total number of records.
:param total: Total number of records.
:param selection_size: Number of representatives to select.
:param fraction: Fraction of total records to consider as candidates.
:param min_candidates: Minimum number of candidates.
:param max_candidates: Maximum number of candidates.
:return: Computed candidate limit.
"""
# 1) fraction of total
limit = int(total * fraction)
# 2) ensure enough to pick selection_size
limit = max(limit, selection_size)
# 3) enforce lower bound
limit = max(limit, min_candidates)
# 4) enforce upper bound (and never exceed the dataset)
limit = min(limit, max_candidates, total)
return limit
Four rules apply in a fixed order: 10% of the corpus by default, raised so there are always at
least as many candidates as representatives requested, raised again to a floor of 100, then
clamped to a ceiling of 1,000 and to the dataset size. The order matters — the floor applies
after the fraction, so small corpora still get a real pool, and the ceiling applies last, so a
large corpus cannot turn one dedup call into an unbounded scan.
get_pure_token: count pure tokens inside the compressor
Token accounting inside a compressor has a nasty failure mode: subword markers counted as content.
Word-piece tokenizers prefix continuation pieces and SentencePiece marks word starts, so a naive
counter treats markers as text and ratios get measured against padded numbers. This function
strips the marker for the two tokenizers the compressor supports, and refuses the rest:
# backend/smartgate/modules/context_gate/utils.py — source lines 105–111 (get_pure_token)
def get_pure_token(token, model_name):
if "bert-base-multilingual-cased" in model_name:
return token.lstrip("##")
elif "xlm-roberta-large" in model_name:
return token.lstrip("▁")
else:
raise NotImplementedError()
The raise NotImplementedError() branch is the design choice worth copying. An unsupported
tokenizer fails loudly instead of returning a count that is mostly marker characters, which is what
makes the ratio numbers elsewhere in this article trustworthy: either the count is right, or the
call tells you it cannot count.
smart_memory: move agent history into team memory
The most effective way to shrink a context window is to stop putting things in it. Memory makes
that possible: instead of carrying a conversation's history in every prompt, the agent writes facts
into a team-scoped store and reads back only what the current step needs:
# backend/smartgate/api/mcp.py — source lines 280–300 (smart_memory)
) -> str:
_, registry = _app_state()
module = registry.get("memory")
ctx = _tool_ctx()
params = _non_empty(
action=action,
query=query or None,
text=text or None,
messages=messages,
user_id=user_id or None,
memory_id=memory_id or None,
top_k=top_k,
threshold=threshold,
metadata=metadata,
)
return await _run_with_audit(
f"memory_{action}",
ctx,
module.process(ctx, **params),
{"action": action},
)
action is one of add, search, get or delete; search takes a query, add takes text or a
messages payload, and top_k caps results at 10 by default over a similarity threshold of
0.1. The _non_empty wrapper turns empty strings into absent arguments, so a caller that
always passes user_id="" does not create a scoped-in-theory record, and each operation is audited
as memory_<action>. Memory belongs to the team and an optional user_id, so two agents on one
key share a store — split keys when that is not what you want.
checkQuota: the monthly token quota check per team
Every control above saves tokens; this one bounds the month:
# lib/usage/index.ts — source lines 101–129 (checkQuota)
async function checkQuota(
teamId: string,
monthlyTokenLimit?: number | bigint | null,
plan?: string | null
): Promise<UsageQuota> {
const normalizedPlan = (plan ?? "FREE").toUpperCase();
let limit =
monthlyTokenLimit != null ? Number(monthlyTokenLimit) : undefined;
if (limit == null || !Number.isFinite(limit) || limit <= 0) {
limit = PLAN_DEFAULT_LIMITS[normalizedPlan] ?? PLAN_DEFAULT_LIMITS.FREE!;
}
const { monthly } = await getCurrentUsage(teamId);
if (limit >= Number.MAX_SAFE_INTEGER / 2) {
return {
limit: Infinity,
used: monthly,
remaining: Infinity,
exceeded: false,
};
}
return {
limit,
used: monthly,
remaining: Math.max(0, limit - monthly),
exceeded: monthly >= limit,
};
}
The precedence inside checkQuota is the whole story. The plan name is upper-cased with FREE as
the default; an explicit monthly limit is used only when it is present, finite and positive;
otherwise the plan's default limit applies, falling back to FREE. Current usage comes from the
team's monthly counter. One branch treats any limit at or above half of the platform's maximum
safe integer as infinity — an explicit no-cap rather than a number large enough to look like one,
with exceeded hard-coded false. Otherwise the caller gets limit, used, remaining and a
single exceeded boolean, which is what separates a quota report from a quota. What that cap is
worth in money, and how the savings get measured, is the other half of token control, worked
through in
Token Optimization Techniques for AI Apps.
smart_pipe: a multi step pipeline so context is not refetched
The last technique removes repetition instead of shrinking it. Fetching, compressing and
remembering as three separate tool calls means three round trips and usually three copies of the
same text in the transcript:
# backend/smartgate/api/mcp.py — source lines 343–387 (smart_pipe)
payload = await engine.run(
ctx,
template=template,
steps=steps,
inputs=run_inputs,
)
data = payload if isinstance(payload, dict) else {"result": payload}
merged_data = {
**data,
"results": payload.get("results") or [],
}
steps_meta = payload.get("pipeline_steps") or []
pipeline_ok = bool(steps_meta) and all(s.get("success") for s in steps_meta)
failed_step = next((s for s in steps_meta if not s.get("success")), None)
failed_result = next(
(r for r in (payload.get("results") or []) if not r.get("success")),
None,
)
step_error = None
if failed_result:
step_error = failed_result.get("error")
elif failed_step:
step_error = failed_step.get("error")
if not pipeline_ok:
merged_data["failed_step"] = failed_step.get("step") if failed_step else None
merged_data["failed_tool"] = failed_step.get("tool") if failed_step else None
merged_data["step_error"] = step_error
tool_result = ToolResult(
success=pipeline_ok,
data=merged_data,
error=None if pipeline_ok else (step_error or "pipeline step failed"),
meta={},
)
audit_extra = {
**params,
"trace_kind": "pipeline",
"pipeline_template": template or None,
"pipeline_steps": payload.get("pipeline_steps") or [],
}
return await _run_with_audit(
"pipe",
ctx,
_identity_result(tool_result),
audit_extra,
)
The excerpt is the return path, and it is where the design pays off. The engine runs the template
or the explicit step list and reports per-step metadata; pipeline_ok is true only when there is
at least one step and every step succeeded. On failure the response carries failed_step,
failed_tool and step_error separately, so "the fetch broke" and "the query returned nothing"
are distinguishable — the difference between retrying and rephrasing. The run is also one audit
trace with trace_kind: pipeline, so a research workflow is one entry, not four calls.
How SmartGate compares
The category splits by where the trimming happens and what it can enforce. SmartGate's scope is
deliberately narrow: it governs tool traffic, and it charges a share only after it has saved
you money.
| Where the trimming happens | What it can enforce | What you pay | |
|---|---|---|---|
| SmartGate | Gateway, in the tool path: compression, segmenting, truncation, dedup, team memory, quotas | Hard monthly token cap per team, per-key MCP rate limits, per-call audit rows | Free: 2M tokens/mo, all seven tools, 120 req/min/key. Pro from $18/mo (first month $5), ~$36/mo cap, share only after $15 saved (pricing) |
| App-side trimming in your own code | Inside your agent: prompt templates, summary buffers, manual chunking | Whatever you write; nothing outside your process, nothing per team | Engineering time; no rate limit, cap, or audit trail shared across clients |
| Local stdio compression servers | On the developer's machine, in-process | Compression and filtering for one host; nothing cross-team | Free/open source; you own the runtime, the versions, and the blast radius |
| LLM routers and observability proxies | Model calls, prompt routing and token accounting | Model-side spend and telemetry | Usage-based on model traffic — a different line item from tool traffic |
Two honest readings. If your problem is one developer wanting shorter prompts in one editor, a
local compression server or a summary buffer in your own code is a reasonable answer: free,
private, enough. The moment the question becomes which agent spent the team's tokens and what stops
it next month, trimming in application code stops being sufficient — that is a quota and an audit
trail, and those live in something every client shares. The savings share is what makes the
incentive legible: a team that saves nothing pays only the platform fee.
How to get started
-
Connect one client. Mint a key and add the hosted, stateless MCP endpoint
(
https://smartgate.network/api/mcp, POST, Streamable HTTP) as a server in your MCP host, so the seven tools appear in the agent's tool list. -
Turn on compression and dedup before autonomy. Run
smart_fetchplussmart_context_gateover material you already pull by hand, addsmart_dedupwhere your searches overlap, and watch the token count on the usage report. -
Pick the ratio deliberately. Start from the 0.5 default, then set the team's
compress_ratioto whatever survives your own question-answering checks. -
Set the cap, then move history out of the prompt. Give the team a monthly token limit so
checkQuotahas something to enforce, push history intosmart_memory, and fold multi-step research intosmart_pipeso the same page is fetched once.
Start with the Free plan — 2M tokens/month, all seven tools, no card:
start free, then compare caps, rate limits
and log retention windows against your workload on the
pricing page.
FAQ
Which technique should I do first?
Compression, because it applies to material you already fetch and needs no new habit. Run a page
you would normally paste through smart_context_gate at the default ratio of 0.5 and compare the
answer; if quality survives, keep the default, and if not, raise the ratio for that use case.
Does compression lose information?
Yes, by design. The ratio is a target for how much text to keep, not a guarantee about which
sentences matter, so anything that must be quoted exactly should stay outside the compressed path.
The purpose argument improves the odds by naming the current question, but it does not make
compression lossless.
What is the difference between dedup and compression?
Compression shrinks each passage; dedup removes passages. smart_dedup compares candidates at a
threshold that defaults to 0.9, so it targets near-identical overlaps rather than merely related
text, and find_representative then picks a diverse survivor set from each duplicate cluster.
When should I truncate instead of compress?
When the tail is known to be worthless, or when you need a deterministic ceiling that costs nothing.
Truncation cuts at a configured cap and never summarises, which makes it the cheapest and least
intelligent option — a backstop behind compression, not the primary control.
How is a context window different from a token quota?
The window is what the model can read in one call; the quota is what your team may spend in a month.
Compression, segmenting, truncation and dedup act on the first. The monthly cap enforced through the
gateway acts on the second, and it is the one that stops an agent loop from billing all night.
Where does team memory fit into a context window?
It is the alternative to keeping history in the prompt. Instead of re-sending a long conversation
every turn, the agent writes facts and messages into team-scoped memory and reads back only the top
matches for the current step, with a default result cap of 10 and a similarity threshold applied.
Limitations and what this does not do
- Compression is lossy and ratio-dependent. A 0.5 target is a starting point, not a validated setting for your corpus; test ratios against questions whose answers you already know.
-
Truncation is blind.
_truncate_pipeline_textreturns the head of the text at a configured cap — it cannot tell a boilerplate footer from a crucial paragraph, and the tail is gone. -
Dedup and diversity are trade-offs you own. A 0.9 threshold lets lightly reworded duplicates
through, and raising
diversitycan drop the single best passage from the representative set. - Quotas are per team, not per agent. Agents sharing a key share the cap and the rate limit; split keys for independent budgets, and note that log retention (7/30/90/180 days) is a plan entitlement, not an archive.
-
A gateway is not a sandbox or a summarizer. It cannot judge whether a fetched page should have
been fetched, and it does not replace prompt design — compaction and clear instructions remain
your responsibility.Sources
SmartGate — product site and pricing: https://smartgate.network · https://smartgate.network/pricing
SmartGate — hosted MCP endpoint and tool reference: https://smartgate.network/docs
Anthropic — Effective context engineering for AI agents: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
Anthropic — Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
Chroma Research — Context Rot (reliability degrades as input grows): https://research.trychroma.com/context-rot
Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172): https://arxiv.org/abs/2307.03172
LLMLingua: prompt compression with preserved task performance (arXiv:2310.05736): https://arxiv.org/abs/2310.05736
MemGPT: memory tiers for long conversations and documents (arXiv:2310.08560): https://arxiv.org/abs/2310.08560
LangChain — Context engineering for agents: https://blog.langchain.com/context-engineering-for-agents/
Model Context Protocol — specification (tools, resources, transports): https://modelcontextprotocol.io/specification/2026-07-28
Method note
The code in this article is not transcribed. Each block was cut directly out of the slice body
returned by the SmartGate slice API and then re-asserted byte-for-byte as a substring of that body
before publication; the first line inside every fence records the file and the exact source lines.
Symbols were pinned with whole-name containment (rule A level 2) and confirmed by a server-side
slot-proof call before any of them entered the text. Where a slice carried a long pydantic field
list, the quoted window was narrowed to the logic rather than reproducing the whole handler;
nothing was rewritten, and repository-relative paths are shown as they are.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | smart_context_gate compress the context window before the model call | smart_context_gate |
backend/smartgate/api/mcp.py |
150–174 | rule A L2 → slot-proof | 39c700a39edd |
| 2 | resolve_compress_ratio resolve the effective compression ratio | resolve_compress_ratio |
backend/smartgate/core/effective_settings.py |
69–74 | rule A L2 → slot-proof | fa11ceea7701 |
| 3 | splitForPlaygroundCompress split long input before compression | splitForPlaygroundCompress |
lib/smartgate/playground-content.ts |
113–166 | rule A L2 → slot-proof | 270a1ccfefed |
| 4 | flushSegment flush a compression segment | flushSegment |
lib/smartgate/playground-content.ts |
135–140 | rule A L2 → slot-proof | d2656cf27e56 |
| 5 | _truncate_pipeline_text truncate pipeline text when the budget is tight | _truncate_pipeline_text |
backend/smartgate/core/pipeline.py |
176–180 | rule A L2 → slot-proof | 68e07702bc1b |
| 6 | smart_dedup semantic dedup of overlapping context passages | smart_dedup |
backend/smartgate/api/mcp.py |
176–196 | rule A L2 → slot-proof | 22b01981f934 |
| 7 | find_representative find the representative passage in a duplicate cluster | find_representative |
backend/smartgate/modules/dedup/algorithm.py |
326–352 | rule A L2 → slot-proof | 11f6bd2056df |
| 8 | compute_candidate_limit compute the dedup candidate limit | compute_candidate_limit |
backend/smartgate/modules/dedup/utils.py |
88–113 | rule A L2 → slot-proof | 9800d86e921c |
| 9 | get_pure_token count pure tokens inside the compressor | get_pure_token |
backend/smartgate/modules/context_gate/utils.py |
105–111 | rule A L2 → slot-proof | d0f34a4f3bd0 |
| 10 | smart_memory move agent history into team memory | smart_memory |
backend/smartgate/api/mcp.py |
280–300 | rule A L2 → slot-proof | f9c4d58b5501 |
| 11 | checkQuota monthly token quota check per team | checkQuota |
lib/usage/index.ts |
101–129 | rule A L2 → slot-proof | cf8c0cdcf082 |
| 12 | smart_pipe multi step pipeline so context is not refetched | smart_pipe |
backend/smartgate/api/mcp.py |
343–387 | rule A L2 → slot-proof | 91ee6d5bb70b |
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before
publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.
Top comments (0)