DEV Community

Cover image for Hardening Claude Code at Scale: Surviving Context Bloat, Rate-Limit Cascades, and Prompt Cache Eviction
yan_cheng
yan_cheng

Posted on

Hardening Claude Code at Scale: Surviving Context Bloat, Rate-Limit Cascades, and Prompt Cache Eviction

At 3:17 AM during an automated monorepo migration, our background agent cluster incinerated $940 of API credit in twenty-two minutes before triggering a hard upstream 429 lock. The culprit was not an infinite loop—it was a cascading context blowout inside anthropics/claude-code. When a terminal coding agent recursively indexes multi-file ASTs, inspects compiler traces, and spawns sub-shells, prompt context expands monotonically. Without an intermediate mediation layer, every minor retry resends a 140k-token payload over cold connections, evicting upstream prompt caches and stalling the entire development pipeline.

As systems engineers integrating anthropics/claude-code into enterprise development fleets, we cannot treat terminal agents as single-user toys. In an engineering org with dozens of engineers running autonomous background workflows, unmanaged CLI instances cause socket churn, concurrent rate-limit collisions, and crippling token spend.

Here is our independent architectural teardown, production proxy hardening, and benchmark analysis for running anthropics/claude-code reliably under enterprise workloads.


The Anatomy of Context Degradation in Terminal Agents

Under the hood, anthropics/claude-code interfaces with developers via an interactive Node.js REPL, translating agentic tool calls (ReadMultipleFiles, Grep, Bash) into structured messages dispatched to upstream model endpoints. In solitary interactive sessions, human-in-the-loop pauses mask latency spikes. But when deployed for automated repo sweeps or multi-ticket refactoring, three distinct systemic failure modes emerge:

  1. Monotonic Context Bloat: Every bash execution, file diff, and git log appends directly to the active conversation history. By turn twelve of an active debugging session, an agent routinely transmits >80,000 tokens per round-trip.
  2. Cache Eviction Thrashing: Provider prompt caching depends on byte-for-byte prefix stability. If an agent prepends dynamic timestamps or fluctuating workspace paths to system instructions, it breaks cache affinity, dropping cache hit ratios from >85% to near zero.
  3. Synchronized 429 Retry Storms: When multiple developer terminals encounter upstream concurrency throttles simultaneously, naive client-side exponential backoff algorithms synchronize, hammering upstream rate-limit counters into an unrecoverable saturation trap.

To establish deterministic operational boundaries, we decoupled client CLI instances from upstream model providers by routing all CLI traffic through an edge mediation topology:

+-----------------------------------------------------------------------+
|                    Developer Workstation / CI Runner                  |
|                                                                       |
|  +-----------------------------------------------------------------+  |
|  |                 anthropics/claude-code CLI Node.js              |
|  |             (ANTHROPIC_BASE_URL=http://127.0.0.1:8080)          |  |
|  +--------------------------------+--------------------------------+  |
+-----------------------------------|-----------------------------------+
                                    | HTTP/1.1 Keep-Alive
                                    v
+-----------------------------------------------------------------------+
|                    Local Sidecar Mediation Proxy                      |
|                                                                       |
|  +-----------------------------------------------------------------+  |
|  |  Envoy Proxy v1.31 (Sliding-Window Rate Limiter & Health Pool)  |  |
|  +--------------------------------+--------------------------------+  |
|                                   | Upstream Retries (Decorrelated)   |
|                                   v                                   |
|  +-----------------------------------------------------------------+  |
|  |  LiteLLM Proxy Router (Prompt Prefix Pinning & Token Governance)|  |
|  +--------------------------------+--------------------------------+  |
+-----------------------------------|-----------------------------------+
                                    | TLS / HTTP/2 Pooled Stream
                                    v
    +----------------------------------------------------------------+
    |           B-Lost Unified High-Throughput Model Gateway         |
    |                   (https://api.b-lost.com/v1)                  |
    +----------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Hardening the Edge: Envoy and Routing Mediation

To prevent runaway CLI processes from destabilizing API quotas, we deployed an Envoy sidecar proxy coupled with an upstream mediation router. Envoy manages connection reuse, TCP keep-alives, and decorrelated jittered backoffs, while the router guarantees prefix stability for prompt cache persistence.

Below is the production Envoy listener and cluster definition enforcing strict circuit breaking and connection timeouts:

# /etc/claude-code-relay/envoy.yaml
static_resources:
  listeners:
    - name: local_claude_ingress
      address:
        socket_address:
          address: 127.0.0.1
          port_value: 8080
      filter_chains:
        - filters:
            - name: envoy.filters.network.http_connection_manager
              typed_config:
                "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
                stat_prefix: ingress_claude
                route_config:
                  name: local_route
                  virtual_hosts:
                    - name: claude_backend
                      domains: ["*"]
                      routes:
                        - match:
                            prefix: "/"
                          route:
                            cluster: upstream_mediator
                            timeout: 180s
                            retry_policy:
                              retry_on: "5xx,gateway-error,connect-failure,reset"
                              num_retries: 3
                              retry_back_off:
                                base_interval: 1.5s
                                max_interval: 15s
                http_filters:
                  - name: envoy.filters.http.router
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router
  clusters:
    - name: upstream_mediator
      connect_timeout: 5s
      type: LOGICAL_DNS
      dns_lookup_family: V4_ONLY
      lb_policy: ROUND_ROBIN
      circuit_breakers:
        thresholds:
          - priority: DEFAULT
            max_connections: 64
            max_pending_requests: 128
            max_requests: 256
      load_assignment:
        cluster_name: upstream_mediator
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address:
                      address: 127.0.0.1
                      port_value: 4000
Enter fullscreen mode Exit fullscreen mode

Behind Envoy, the mediator routes requests to cached enterprise gateway endpoints, enforcing strict token budget clamps and system prompt pinning:

# /etc/claude-code-relay/litellm_config.yaml
model_list:
  - model_name: claude-3-7-sonnet-20250219
    litellm_params:
      model: anthropic/claude-3-7-sonnet-20250219
      api_base: https://api.b-lost.com/v1
      api_key: os.environ/BLOST_API_KEY
      rpm: 120
      tpm: 160000
      max_retries: 3
      timeout: 180

router_settings:
  routing_strategy: usage-based-routing-v2
  enable_pre_call_checks: true
  retry_after_policy:
    min_backoff_seconds: 2
    max_backoff_seconds: 30
    jitter: true

general_settings:
  master_key: os.environ/PROXY_MASTER_KEY
  pass_through_headers:
    - "anthropic-beta"
    - "anthropic-version"
Enter fullscreen mode Exit fullscreen mode

Configure your developer shell to route claude-code through the hardened proxy:

# Point Claude Code CLI to the hardened local edge relay
export ANTHROPIC_BASE_URL="http://127.0.0.1:8080"
export ANTHROPIC_API_KEY="sk-proxy-internal-dev-key"
export CLAUDE_AUTO_COMPACT_WINDOW="45000"

# Launch Claude Code within the protected boundary
claude
Enter fullscreen mode Exit fullscreen mode

Empirical Stress Benchmark: Direct CLI vs. Mediated Relay

To quantify the operational impact of our proxy topology, we benchmarked anthropics/claude-code across a standardized full-stack refactoring workload: migrating a 42,000-line TypeScript repository from CommonJS to ESM, including dependency graph resolution, AST transformation, and automated Jest test execution.

We ran five parallel agent workers under three discrete routing configurations:

Architecture Setup Total Transmitted Tokens Prompt Cache Hit Rate Aggregate Cost Wall-Clock Completion Upstream 429 Errors
Direct CLI (Unmanaged Claude Code) 6.84M 18.2% $29.14 54m 30s 23 retries / 3 failures
Direct CLI (With Static Memory Pruning) 5.12M 39.5% $20.80 43m 12s 9 retries / 0 failures
Mediated Envoy Relay + Cache Pinning 4.98M 89.6% $4.72 24m 45s 0 retries / 0 failures

Key architectural insights from our benchmark runs:

  1. Prompt Cache Retention Determines Unit Economics: In multi-turn coding sessions, re-reading repo summaries and linter configurations constitutes over 80% of aggregate input volume. Caching prefixes at the gateway reduced token expenditures by 83.8%.
  2. Circuit Breaking Prevents Cascading Outages: Unmitigated 429 errors trigger client-side spin loops that burn thread pools. Envoy's connection pools absorbed transient upstream backpressure without dropping client sessions.
  3. P95 Latency Drops Exponentially: Because cached prompt tokens evaluate significantly faster upstream, overall task completion time dropped by 54.6%, transforming batch agent refactors into viable CI verification gates.

The Operational Dilemma

Adopting anthropics/claude-code at engineering scale creates a structural dilemma for platform teams: Should we treat autonomous agent runtimes as unprivileged CLI tools governed by local workstation cgroups and sidecars, or should all agent state and file mutations be pushed entirely into centralized remote container clusters?

Local execution provides zero-friction filesystem access, but centralized proxies are mandatory to prevent quota exhaustion and runaway bills. How does your engineering organization manage agent rate limits, prompt caching, and token governance under concurrent developer load? Drop your architecture or battle scars in the comments below.


Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by b-lost.com — an enterprise AI gateway offering 0.58x-0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.

Top comments (0)