Introduction
Modern AI Agents are frequently constrained by context window limits, token cost and the degradation of reasoning quality as dialogue history grows. Simple prompt stuffing leads to expanded context, rising inference latency and gradual information dilution. This article systematically breaks down three core engineering components for building stable long-context agents: layered Skill on-demand loading, Agent state bar tracking, and multi-strategy context compression. It also introduces context isolation as an alternative paradigm for complex task workflows. All experimental data and architectural conclusions in this article are derived from practical Agent harness development.
The core design philosophy of modern Skill systems can be summarized as: keep lightweight metadata resident in the prompt, while loading complete Skill content only at runtime when the task requires it. This split design balances low baseline token overhead and flexible task capability expansion.
Layered Skill Architecture and On-Demand Loading Mechanism
Skill is a standardized set of task specifications, operation workflows and judgment rules for AI agents. Rather than embedding all task instructions into the system prompt permanently, the layered structure separates routing metadata, core workflow and supplementary reference documents into three tiers.
Tier 1: Routing Metadata for Decision Making
The description field inside Skill metadata serves as the core signal for routing decisions. It must remain concise to control resident token consumption, while written strictly as routing conditions: defining when this Skill should be triggered and what capabilities it provides. Overly broad descriptions such as “help with backend tasks” cause false positive matching and misrouting. A well-written description clearly specifies task scenarios, input types and boundary conditions.
This metadata stays in the agent catalog persistently. It occupies a small fixed token footprint, and the full Skill body will not be loaded unless the routing logic determines a match.
Tier 2: Core Workflow, Loaded On Demand
When the agent judges that a task requires a particular Skill, the full SKILL.md content is injected into the context at runtime. There are two common trigger modes. The first mode is explicit invocation: users input command prefixes such as /pptx, and the client intercepts the request and loads the Skill. The second mode is implicit matching: the agent automatically selects and loads the Skill according to task intent, without extra ReAct turns.
Both approaches converge at the same execution logic. Tools like Claude Code insert the full Skill text as a user message at the point of tool invocation. The model executes tool calls according to the workflow defined inside the Skill, and does not permanently retain the full Skill content after the tool result is returned.
Take the PPTX Skill as an example. Its core workflow defines how to extract text from PowerPoint files via markdown conversion, decompress PPTX packages to read raw XML structures, and enforce constraints on file path operations.
Tier 3: Supplementary Reference Documents
The main Skill file can reference additional markdown documents such as html2pptx.md and reference.md. These sub-files contain detailed workflow templates and low-level technical specifications. The agent selectively reads these sub-documents only when task execution hits relevant steps, avoiding loading all reference materials from the start.
How to Write a Production-Grade Skill
A usable Skill converts team experience into executable instructions for the model. A newly onboarded engineer should be able to understand applicable scenarios, execution sequence, confirmation checkpoints and completion criteria by reading this document.
According to the structured Skill framework, a complete Skill includes four core sections:
- Role and scope: Define target tasks and output quality standards for this Skill.
- Core judgment rules: List the most critical decision points, with positive and negative examples.
- Forbidden operations and boundary constraints: Clearly specify prohibited actions and safety boundaries.
- Step-by-step execution workflow: Break down the whole task into ordered operations, with checkpoints for human confirmation.
Different Agent harness frameworks implement Skill catalogs in different ways. OpenAI Codex rebuilds the Skill catalog at every context construction step, marking selected Skill content with dedicated tags. Other harnesses fetch Skill content through separate tool read operations. Despite implementation differences, nearly all mature systems follow the principle: keep lightweight directory metadata resident, load full text on demand. This design is the key to combining dynamic task capability with manageable context overhead.
Agent State Bar: Track Runtime Status Without Model Fine-Tuning
The Agent state bar is a lightweight context compression and state tracking technique. It injects machine-readable runtime metadata into the dialogue context, allowing the agent to record progress, retry history and environmental information. The state bar is non-intrusive to underlying LLMs; it works on almost all model variants without parameter fine-tuning.
The state bar contains multiple independent modules working together to create emergent effects. Individually each component brings limited gains, but combined they drastically improve task completion rates.
Tool Call Counter
A global dictionary records the invocation count of each tool. Every tool response annotates the sequence number, for example Tool call #3 for 'read_file'. This counter helps the model recognize repeated failure patterns. After the first failed directory lookup, the second attempt lists directory contents, and on the third failure the agent abandons the path and searches for alternative solutions. This counter implicitly builds cost awareness; the agent can detect excessive retries for a single operation.
TODO List Management
Inspired by human attention management practices, the TODO module provides two dedicated operations: rewrite_todo_list and update_todo_status. Each TODO item carries an identifier, task content, status flag (pending/in_progress/completed/cancelled) and timestamp.
The TODO list acts as external memory for agents, just like task checklists for human engineers. Experimental data shows agents equipped with TODO lists finish tasks in an average of 15 iterations. Agents without TODO management require 21 iterations for the same work.
Detailed Error Information
This module packages failure logs into four parts: error type and description, complete JSON parameters, invocation stack trace, and targeted remediation advice. For example, upon FileNotFoundError, the advice suggests verifying paths, checking working directories and switching to absolute paths.
After enabling this module, agent success rates for troubleshooting tasks rose from 60% to 95%. The agent evolves from blind repeated retries to targeted fault diagnosis and resolution.
System Status Sensing
The state bar injects real-time environmental metadata: current timestamp, working directory, operating system, shell interpreter and Python version. Working directory tracking is especially critical. When the agent executes cd, the state bar automatically updates this field, ensuring subsequent file operations run under the correct path. OS metadata enables platform-aware decisions, for instance selecting apt on Linux and brew on macOS.
Two Critical Rules for Maintaining State Bars
- State counters must be maintained by code. If LLM is used to aggregate counts, the model will trust its own generated numbers unconditionally and cannot recalculate values. LLMs easily introduce counting errors, which creates hidden risks for task logic.
- Handle raw context with caution. The state bar is a lossy summary of history. If all required information can be represented inside state fields, raw dialogue records may be pruned to save tokens. However, once a task dimension falls outside tracked state fields, deleting raw records will cause sharp accuracy drops.
Context Compression Strategies: Reduce Context Without Destroying Reasoning
Skill and state bars control what information enters context. Context compression solves the opposite problem: removing redundant content from accumulated dialogue history. Compression is not only for fitting within window limits; it improves reasoning quality and mitigates context anxiety.
Three Core Motivations for Context Compression
- Meet length and cost constraints. Context windows have hard limits such as 128k tokens. Uncompressed tool outputs quickly consume available space, while longer context also increases API cost and inference latency.
- Improve reasoning quality. Summarized knowledge is easier for models to use than raw verbose records. Long uncompressed history disperses model attention; key facts become buried in redundant text. Compression converts conclusions derived through reasoning into directly retrievable knowledge.
- Alleviate context anxiety. When an LLM detects that the context window is nearly exhausted, it may terminate unfinished tasks prematurely. Proactive compression before hitting window limits stabilizes decision quality.
The Internal Nature of Context Compression: Retrieval Rather Than Reasoning
LLM attention is optimized for searching existing content rather than automatic aggregation across long sequences. Compression transforms raw verbose records into structured summaries stored inside context. The difference between state bars and compression: state bars maintain metadata continuously, while compression calls LLM once to condense large text blocks.
A simple example illustrates this principle. Suppose context contains inspection records for 100 animal cages, listing black and white cats repeatedly. When asked to count total cats, an uncompressed model needs to scan every entry sequentially, consuming massive reasoning effort. If the summary precomputes the result “90 black cats, 10 white cats”, the model directly retrieves this number. Compression shifts computational burden from runtime reasoning to pre-summary processing.
Unmanaged long context also triggers context rot, a separate issue from window overflow. Overflow means content cannot fit into the window. Context rot means content fits, but the model cannot locate critical information. As token volume rises, attention is spread thinner, and each token carries lower effective weight. Irrelevant records gradually dominate the context, and agent decision quality declines.
Comparative Experiment of Six Compression Strategies
A research task was designed to identify and track career milestones of OpenAI co-founders. The task required multi-step information aggregation, with raw material ranging from tens of thousands to hundreds of thousands of tokens. All strategies ran within a 128k context limit.
| Strategy | Description | Compression Rate | Observations |
|---|---|---|---|
| 1. No compression | Preserve raw tool outputs completely | 0% | After repeated retrieval, raw content exceeded 128k token limit at the 5th iteration. |
| 2. Non-task-aware compression | Summarize each tool result independently | 10.9% | Independent summaries lose cross-document correlation; events described across multiple pages become fragmented. |
| 3. Merged summary compression | Combine all results into one single summary | 43.9% | Global summaries save token space, but risk losing fine-grained details and cannot distinguish information relevance. |
| 4. Context-aware compression | Compress with reference to query intent and accumulated information | 3.0% | Summaries retain key facts targeted to ongoing task queries. |
| 5. Context-aware compression + reference links | Preserve compressed summary and add source URL references | - | Lossy summary paired with lossless source links for traceback. |
| 6. Adaptive windowing | Trigger compression only when token usage exceeds 80% of window capacity | 76% | Retain full raw records in early iterations; activate batch compression near threshold, maximizing original information integrity. |
Adaptive windowing has three core mechanisms: threshold trigger, batch compression, and marker tagging. The system continuously monitors prompt token usage. When usage exceeds the 80% threshold, it compresses all untagged tool outputs in one batch, marking compressed content so it will not be reprocessed repeatedly.
Production-Grade Hierarchical Compression Pipeline
Production Agent systems rarely rely on a single compression algorithm. Taking Claude Code architecture as reference, mature context management includes five stacked layers:
- Tool result preview control: Store large raw outputs on disk, only inject preview snippets into prompt. Once a decision is made, the result snapshot is frozen.
- Deduplication and low-value pruning: Directly delete redundant or rarely referenced content, without extra summarization.
- API-side compression: Offload compression work to server-side endpoints. This approach reduces local token consumption and avoids repeated reconstruction cost.
- Structured rolling digest: Generate sequential structured summaries, preserving independent records for each turn instead of merging everything into one block.
- Full global compression: The final fallback layer. Global compression also has staged controls and circuit breakers. If repeated compression attempts keep failing, the circuit breaker stops compression to avoid infinite loops. Production data shows many conversations get stuck in failed compression cycles; circuit breakers prevent this failure mode.
Core Design Principles for Compression
- Non-uniform information value: Key decision points have higher priority than supporting evidence.
- Semantic integrity: Preserve entity relationships and timelines; avoid distorting facts during summarization.
- Task relevance: Compression scope is bounded by the active task. Cross-task content should be isolated.
- Agent awareness: The compression module must understand the agent’s current subtask and reasoning goals.
- Compression transparency: The summary should be readable and auditable, so developers can inspect how history is condensed.
Compression still carries overhead. Every compression step invokes an LLM call, but experiments show it reduces total token consumption by over 75% and improves task success rates. The investment for compression yields positive returns for long-running multi-turn agent workflows.
Context Isolation: Replace Compression With Sub-Agents
Compression is inherently lossy and requires post-repair via extra LLM calls. Context isolation avoids this problem from the beginning by spawning independent sub-agents. Each sub-agent works inside its own private context window and returns only final conclusions to the parent agent. Raw intermediate history stays within the sub-agent and is not passed upward.
This paradigm trades extra model invocation for clean context separation. The parent agent delegates subtasks, sub-agents handle heavy retrieval and tool operations, and only deliver condensed results. Claude Code’s Task tool and Deep Research module both implement this isolation pattern.
In multi-agent production systems, developers need unified management for model endpoints, authentication and request routing. 4sapi, an API gateway, simplifies multi-model access orchestration for distributed agent workflows.
Conclusion
Building reliable long-context AI agents requires systematic optimization across the whole context lifecycle. Skill on-demand loading controls what instructions enter the prompt. Agent state bars track runtime progress with lightweight metadata, without modifying underlying models. Context compression reduces token volume and mitigates context rot, while context isolation provides a lossless alternative by delegating work to independent sub-agents.
The experimental data and layered architecture in this article provide reusable patterns for Agent harness development. Teams can combine these technologies according to task complexity, window size and cost budget. The key insight is to avoid simply stuffing everything into the prompt; design layered loading, state tracking and history reduction to maintain stable reasoning over extended multi-turn tasks.
International access: https://4sapi.com
Domestic access: https://4sapi.cn
Top comments (0)