DEV Community

Tamiz Uddin
Tamiz Uddin

Posted on Originally published at tamiz.pro

Building Agent-Native Apps Without Burning Your Budget: Practical Patterns for Sandboxed AI Agent Architectures

Originally published on tamiz.pro.

AI agents are the most powerful and most expensive primitive in modern software. A single agentic loop that explores, plans, calls tools, retries, and synthesizes can burn through a month's API budget in hours — and your monitoring won't catch it until the invoice arrives. The core engineering challenge isn't capability; it's containment. How do you give agents autonomy to accomplish complex tasks while ensuring a runaway loop doesn't liquidate your runway?

This deep dive covers five production-proven patterns for building sandboxed, budget-aware AI agent architectures, with complete implementation code, architectural diagrams, and the operational metrics you need to keep costs predictable while preserving agent effectiveness.

Table of Contents

1. The Agent Cost Problem: Where Budgets Actually Burn

Before diving into solutions, you need a precise mental model of where agent costs accumulate. Unlike a simple request-response API call, an agent performs a loop: plan → act → observe → replan. Each iteration multiplies cost through several channels simultaneously.

The Cost Multiplication Taxonomy

Cost Channel Typical Amplifier Example Scenario
LLM inference Iteration count × context length 15-step reasoning loop with growing conversation history
Tool execution Retry storms, unbounded loops Agent retries a failed API call 47 times
Context accumulation O(n²) token growth Each loop appends prior outputs to the prompt
Parallel exploration Branching factor × depth Agent spawns 5 sub-agents, each spawning 3
Silent failures Error recovery loops Agent keeps retrying a misconfigured tool

The worst-case scenario is a feedback loop: a tool returns a confusing error, the agent re-plans using a larger context, the new plan calls more tools, which produce more errors, and the context grows quadratically. A single user request can cascade into thousands of tokens and multiple tool calls per second.

The solution isn't to reduce agent capability — it's to build containment layers that let agents work autonomously within hard financial boundaries.

2. Pattern One: Tiered Compute Isolation

Concept

Not every agent task deserves the same compute resources. A simple summarization task and a multi-step research task have fundamentally different cost profiles. Tiered compute isolation assigns each agent invocation to a resource tier based on task complexity, with hard ceilings per tier.

                    ┌─────────────────────────┐
                    │   Task Complexity       │
                    │   Classifier            │
                    └────────┬────────────────┘
                             │
              ┌──────────────┼──────────────┐
              ▼              ▼              ▼
     ┌────────────┐ ┌────────────┐ ┌────────────┐
     │  Micro     │ │  Standard  │ │  Extended  │
     │  (1 model  │ │  (1 model  │ │  (multi-   │
     │   call,    │ │   + tools, │ │   model,   │
     │   ≤2K tok) │ │   ≤15K tok)│ │   ≤50K tok)│
     └────────────┘ └────────────┘ └────────────┘
Enter fullscreen mode Exit fullscreen mode

Implementation

The classifier is itself a lightweight model call — a single prompt asking a small/fast model to categorize the task. This costs a fraction of what the actual task will cost, so it's economically sound.

from dataclasses import dataclass
from enum import Enum
import time


class ComputeTier(Enum):
    MICRO = "micro"        # Single LLM call, no tools
    STANDARD = "standard"  # LLM + tools, bounded iterations
    EXTENDED = "extended"  # Multi-model, multi-step reasoning


@dataclass
class TierConfig:
    tier: ComputeTier
    max_iterations: int
    max_tokens_total: int
    max_tool_calls: int
    allowed_models: list[str]
    timeout_seconds: int
    max_concurrent_subagents: int


TIER_REGISTRY: dict[ComputeTier, TierConfig] = {
    ComputeTier.MICRO: TierConfig(
        tier=ComputeTier.MICRO,
        max_iterations=1,
        max_tokens_total=2_000,
        max_tool_calls=0,
        allowed_models=["gpt-4o-mini", "claude-haiku"],
        timeout_seconds=30,
        max_concurrent_subagents=0,
    ),
    ComputeTier.STANDARD: TierConfig(
        tier=ComputeTier.STANDARD,
        max_iterations=8,
        max_tokens_total=15_000,
        max_tool_calls=12,
        allowed_models=["gpt-4o", "claude-sonnet"],
        timeout_seconds=120,
        max_concurrent_subagents=2,
    ),
    ComputeTier.EXTENDED: TierConfig(
        tier=ComputeTier.EXTENDED,
        max_iterations=25,
        max_tokens_total=50_000,
        max_tool_calls=40,
        allowed_models=["gpt-4o", "claude-sonnet", "claude-opus"],
        timeout_seconds=300,
        max_concurrent_subagents=5,
    ),
}


class TierClassifier:
    """Classifies incoming agent tasks into compute tiers."""

    def __init__(self, model_client):
        self.model_client = model_client

    def classify(self, user_request: str, context: dict) -> ComputeTier:
        classification_prompt = f"""Classify this task into one of three tiers:
- MICRO: Simple single-step tasks (summarize, translate, extract)
- STANDARD: Multi-step tasks requiring tool use (search, research, code review)
- EXTENDED: Complex multi-phase tasks (system design, deep research, multi-file refactoring)

Request: "{user_request[:500]}"
Context summary: {context.get('summary', 'none')}

Return ONLY the tier name."""

        response = self.model_client.complete(
            model="gpt-4o-mini",
            prompt=classification_prompt,
            max_tokens=20,
        )

        tier_str = response.strip().upper()
        if "EXTENDED" in tier_str:
            return ComputeTier.EXTENDED
        elif "STANDARD" in tier_str:
            return ComputeTier.STANDARD
        return ComputeTier.MICRO
Enter fullscreen mode Exit fullscreen mode

The key insight: the classification call itself costs roughly 50-100 tokens, which is negligible compared to the 2,000-50,000 token budgets it governs. You spend pennies to avoid burning dollars.

Tier-Based Model Routing

Each tier restricts which models the agent can use. Micro tasks go to the cheapest capable model. Extended tasks can access premium models but only when the agent's reasoning actually benefits from it — enforced by a quality gate that compares model outputs.

3. Pattern Two: Token Budgeting with Circuit Breakers

Concept

A circuit breaker is a reliability pattern from distributed systems that prevents cascading failures. Applied to agent token budgets, it monitors consumption velocity and trips when spending exceeds safe thresholds, halting the agent loop before the budget is exhausted.

The circuit breaker has three states:

  1. CLOSED — Normal operation. Tokens are tracked but no intervention occurs.
  2. HALF-OPEN — Budget threshold crossed (e.g., 70%). The agent is allowed to continue but with reduced capabilities (fewer tools, smaller context windows).
  3. OPEN — Hard budget exceeded. The agent loop terminates immediately. The partial result is returned with a status flag.

Implementation

import time
from dataclasses import dataclass, field
from enum import Enum
from typing import Callable


class CircuitState(Enum):
    CLOSED = "closed"
    HALF_OPEN = "half_open"
    OPEN = "open"


@dataclass
class TokenBudget:
    """Token budget with circuit breaker enforcement."""
    total_budget: int
    consumed: int = 0
    state: CircuitState = CircuitState.CLOSED
    warning_threshold: float = 0.70  # 70% triggers HALF_OPEN
    hard_threshold: float = 1.00     # 100% triggers OPEN
    velocity_window_seconds: int = 60
    max_velocity_tokens_per_second: int = 500
    _token_log: list[tuple[float, int]] = field(default_factory=list)

    def consume(self, amount: int) -> bool:
        """Attempt to consume tokens. Returns False if budget exceeded."""
        if self.state == CircuitState.OPEN:
            return False

        self.consumed += amount
        now = time.time()
        self._token_log.append((now, amount))

        # Clean old log entries
        cutoff = now - self.velocity_window_seconds
        self._token_log = [(t, a) for t, a in self._token_log if t >= cutoff]

        ratio = self.consumed / self.total_budget

        if ratio >= self.hard_threshold:
            self.state = CircuitState.OPEN
            return False
        elif ratio >= self.warning_threshold:
            self.state = CircuitState.HALF_OPEN

        return True

    def get_velocity(self) -> float:
        """Calculate current token consumption rate."""
        if len(self._token_log) < 2:
            return 0.0
        elapsed = self._token_log[-1][0] - self._token_log[0][0]
        if elapsed == 0:
            return 0.0
        total_in_window = sum(a for _, a in self._token_log)
        return total_in_window / elapsed

    def is_velocity_exceeded(self) -> bool:
        return self.get_velocity() > self.max_velocity_tokens_per_second


class BudgetEnforcedAgent:
    """Agent loop with token budget circuit breaker."""

    def __init__(self, model_client, budget: TokenBudget, tools: list):
        self.model_client = model_client
        self.budget = budget
        self.tools = tools
        self.iteration = 0

    def run(self, task: str, max_iterations: int = 10) -> dict:
        messages = [{"role": "user", "content": task}]
        result = {"status": "completed", "output": None, "iterations": 0}

        for self.iteration in range(1, max_iterations + 1):
            # Check circuit breaker state
            if self.budget.state == CircuitState.OPEN:
                result["status"] = "budget_exhausted"
                result["partial_output"] = messages[-1].get("content", "")
                break

            # In HALF_OPEN state, reduce context by trimming older messages
            effective_messages = messages
            if self.budget.state == CircuitState.HALF_OPEN:
                effective_messages = self._trim_context(messages, max_messages=4)

            # Check velocity
            if self.budget.is_velocity_exceeded():
                result["status"] = "velocity_throttled"
                result["partial_output"] = messages[-1].get("content", "")
                break

            # Make LLM call
            response = self.model_client.chat(
                messages=effective_messages,
                tools=self.tools if self.budget.state != CircuitState.HALF_OPEN else self.tools[:2],
            )

            # Track token consumption
            tokens_used = response.get("usage", {}).get("total_tokens", 0)
            if not self.budget.consume(tokens_used):
                result["status"] = "budget_exhausted"
                result["partial_output"] = response.get("content", "")
                break

            messages.append({"role": "assistant", "content": response.get("content")})

            # Check if agent is done
            if not response.get("tool_calls"):
                result["output"] = response.get("content")
                result["iterations"] = self.iteration
                break

        return result

    def _trim_context(self, messages: list, max_messages: int) -> list:
        """Reduce context in HALF_OPEN state to slow consumption."""
        if len(messages) <= max_messages:
            return messages
        # Keep system message + last N messages
        if messages[0].get("role") == "system":
            return [messages[0]] + messages[-(max_messages - 1):]
        return messages[-max_messages:]
Enter fullscreen mode Exit fullscreen mode

Velocity-Based Throttling

The velocity check adds a second axis of protection. Even if the total budget hasn't been reached, a sudden spike in token consumption (indicating a runaway loop) triggers throttling. This catches the scenario where an agent enters a pathological retry loop that would exhaust a large budget in seconds.

4. Pattern Three: Sandboxed Tool Execution

Concept

Agent tools are the highest-risk component for cost overruns. A tool that returns large payloads (e.g., a web search returning 50KB of HTML), a tool that triggers external side effects (e.g., sending emails), or a tool that loops internally (e.g., a recursive file processor) can each independently blow past budget limits.

Sandboxed tool execution wraps every tool call with resource constraints: output size limits, execution time limits, retry limits, and side-effect gates.

Implementation

import asyncio
import hashlib
import json
from dataclasses import dataclass
from typing import Any, Callable


@dataclass
class ToolBudget:
    max_calls: int
    calls_made: int = 0
    max_output_tokens: int = 2_000
    max_execution_seconds: float = 10.0
    max_retries: int = 2


class SandboxedTool:
    """Wraps a tool function with resource constraints."""

    def __init__(self, name: str, func: Callable, budget: ToolBudget):
        self.name = name
        self.func = func
        self.budget = budget
        self._call_log: list[dict] = []

    async def execute(self, arguments: dict) -> dict:
        # Enforce call limit
        if self.budget.calls_made >= self.budget.max_calls:
            return {
                "error": "tool_call_limit_exceeded",
                "tool": self.name,
                "max_calls": self.budget.max_calls,
            }

        # Enforce execution time limit with asyncio timeout
        for attempt in range(self.budget.max_retries + 1):
            try:
                self.budget.calls_made += 1
                start = asyncio.get_event_loop().time()

                result = await asyncio.wait_for(
                    self.func(**arguments),
                    timeout=self.budget.max_execution_seconds,
                )

                elapsed = asyncio.get_event_loop().time() - start

                # Enforce output size limit
                result_str = json.dumps(result) if not isinstance(result, str) else result
                if len(result_str) > self.budget.max_output_tokens * 4:  # ~4 chars per token
                    result = self._truncate_output(result, self.budget.max_output_tokens)

                self._call_log.append({
                    "attempt": attempt,
                    "elapsed": elapsed,
                    "output_size": len(result_str),
                    "success": True,
                })

                return {"result": result, "metadata": {"elapsed": elapsed}}

            except asyncio.TimeoutError:
                self._call_log.append({"attempt": attempt, "error": "timeout"})
                if attempt == self.budget.max_retries:
                    return {"error": "tool_timeout", "tool": self.name}

            except Exception as e:
                self._call_log.append({"attempt": attempt, "error": str(e)})
                if attempt == self.budget.max_retries:
                    return {"error": "tool_failed", "tool": self.name, "detail": str(e)}

        return {"error": "unexpected_exit", "tool": self.name}

    def _truncate_output(self, result: Any, max_tokens: int) -> Any:
        """Truncate tool output to prevent context bloat."""
        max_chars = max_tokens * 4
        if isinstance(result, str):
            return result[:max_chars] + "\n[TRUNCATED]"
        result_str = json.dumps(result)
        if len(result_str) <= max_chars:
            return result
        # For structured data, keep the first N items
        if isinstance(result, list):
            truncated = result[:10]
            return {"items": truncated, "truncated": True, "total_count": len(result)}
        return json.loads(result_str[:max_chars])

    @property
    def stats(self) -> dict:
        return {
            "tool": self.name,
            "calls_made": self.budget.calls_made,
            "max_calls": self.budget.max_calls,
            "avg_elapsed": sum(c["elapsed"] for c in self._call_log if "elapsed" in c) / max(len([c for c in self._call_log if "elapsed" in c]), 1),
        }


class ToolRegistry:
    """Central registry of sandboxed tools with per-tool budgets."""

    def __init__(self):
        self._tools: dict[str, SandboxedTool] = {}

    def register(self, name: str, func: Callable, budget: ToolBudget) -> None:
        self._tools[name] = SandboxedTool(name, func, budget)

    async def execute(self, name: str, arguments: dict) -> dict:
        if name not in self._tools:
            return {"error": "tool_not_found", "tool": name}
        return await self._tools[name].execute(arguments)

    def get_all_stats(self) -> dict:
        return {name: tool.stats for name, tool in self._tools.items()}
Enter fullscreen mode Exit fullscreen mode

Side-Effect Gating

For tools that produce irreversible side effects (sending emails, making payments, modifying databases), add a confirmation gate. The agent must produce a structured confirmation that matches the expected action before the tool executes:

class SideEffectGate:
    """Requires explicit confirmation before executing side-effect tools."""

    def __init__(self, allowed_actions: set[str]):
        self.allowed_actions = allowed_actions

    async def execute(self, action: str, payload: dict, confirm: dict) -> dict:
        if action not in self.allowed_actions:
            return {"error": "action_not_allowed", "action": action}

        # Verify the agent's confirmation matches the actual payload
        if not self._confirmations_match(payload, confirm):
            return {
                "error": "confirmation_mismatch",
                "expected": confirm,
                "actual": payload,
            }

        # Now actually execute
        return await self._do_execute(action, payload)

    def _confirmations_match(self, payload: dict, confirm: dict) -> bool:
        """Check that agent's stated intent matches actual payload."""
        for key, expected in confirm.items():
            actual = payload.get(key)
            if actual != expected:
                return False
        return True
Enter fullscreen mode Exit fullscreen mode

5. Pattern Four: Semantic Caching and Deduplication

Concept

Agents frequently make redundant LLM calls with semantically similar prompts — especially in multi-step workflows where the agent re-asks similar questions after gathering new information. Semantic caching intercepts these calls and returns cached results when the incoming prompt is sufficiently similar to a previously answered one.

Unlike exact-match caching (which fails when prompts differ by a single word), semantic caching uses embedding similarity to match semantically equivalent queries.

Implementation

import hashlib
import time
from collections import OrderedDict
from dataclasses import dataclass
from typing import Optional

import numpy as np


@dataclass
class CacheEntry:
    embedding: np.ndarray
    response: str
    model: str
    timestamp: float
    hit_count: int = 0
    tokens_saved: int = 0


class SemanticCache:
    """Embedding-based cache for LLM responses."""

    def __init__(
        self,
        embedding_model: str = "text-embedding-3-small",
        similarity_threshold: float = 0.92,
        max_entries: int = 1000,
        ttl_seconds: int = 3600,
    ):
        self.embedding_model = embedding_model
        self.similarity_threshold = similarity_threshold
        self.max_entries = max_entries
        self.ttl_seconds = ttl_seconds
        self._cache: OrderedDict[str, CacheEntry] = OrderedDict()
        self._hits = 0
        self._misses = 0

    async def get_or_compute(
        self,
        prompt: str,
        model: str,
        compute_fn: Callable[[], str],
        token_count: int = 0,
    ) -> str:
        """Check cache first, compute if miss."""
        embedding = await self._embed(prompt)

        # Find best match
        best_match: Optional[CacheEntry] = None
        best_score = 0.0

        for entry in self._cache.values():
            if time.time() - entry.timestamp > self.ttl_seconds:
                continue
            score = np.dot(embedding, entry.embedding) / (
                np.linalg.norm(embedding) * np.linalg.norm(entry.embedding)
            )
            if score > best_score:
                best_score = score
                best_match = entry

        if best_match and best_score >= self.similarity_threshold:
            best_match.hit_count += 1
            best_match.tokens_saved += token_count
            self._hits += 1
            self._cache.move_to_end(best_match.key)
            return best_match.response

        # Cache miss — compute and store
        self._misses += 1
        response = await compute_fn()

        cache_key = hashlib.md5(prompt.encode()).hexdigest()
        self._cache[cache_key] = CacheEntry(
            embedding=embedding,
            response=response,
            model=model,
            timestamp=time.time(),
        )

        # Evict oldest if over capacity
        while len(self._cache) > self.max_entries:
            self._cache.popitem(last=False)

        return response

    async def _embed(self, text: str) -> np.ndarray:
        """Generate embedding for text (delegated to embedding API)."""
        # In production, this calls your embedding provider
        # Placeholder for illustration
        return np.random.randn(1536)  # Replace with actual embedding call

    @property
    def hit_rate(self) -> float:
        total = self._hits + self._misses
        return self._hits / total if total > 0 else 0.0

    @property
    def stats(self) -> dict:
        return {
            "hits": self._hits,
            "misses": self._misses,
            "hit_rate": self.hit_rate,
            "entries": len(self._cache),
            "total_tokens_saved": sum(e.tokens_saved for e in self._cache.values()),
        }
Enter fullscreen mode Exit fullscreen mode

Deduplication Across Sub-Agents

When an orchestrator spawns multiple sub-agents working on parallel tasks, they often query the same external APIs or ask similar questions. A shared semantic cache at the orchestrator level prevents redundant work:

class SharedContextCache:
    """Cross-agent cache for shared knowledge."""

    def __init__(self, semantic_cache: SemanticCache):
        self.semantic_cache = semantic_cache
        self._agent_registry: dict[str, dict] = {}

    async def agent_query(
        self, agent_id: str, question: str, compute_fn: Callable[[], str]
    ) -> str:
        """Query with cross-agent deduplication."""
        result = await self.semantic_cache.get_or_compute(
            prompt=f"[agent:{agent_id}] {question}",
            model="shared",
            compute_fn=compute_fn,
        )

        # Track what each agent has queried
        if agent_id not in self._agent_registry:
            self._agent_registry[agent_id] = {"queries": [], "cache_hits": 0}
        self._agent_registry[agent_id]["queries"].append(question)

        return result

    def get_agent_stats(self, agent_id: str) -> dict:
        return self._agent_registry.get(agent_id, {})
Enter fullscreen mode Exit fullscreen mode

6. Pattern Five: Progressive Escalation

Concept

Progressive escalation is the agent equivalent of starting with a small investment and scaling up only when justified. Rather than committing to a full multi-step, multi-tool, premium-model execution from the start, the agent begins with the cheapest approach and escalates only when the initial result is insufficient.

The escalation ladder:

  1. Direct answer — Single LLM call with no tools
  2. Augmented — Single LLM call with tool retrieval (RAG)
  3. Reasoned — Multi-step reasoning with tool use
  4. Ensemble — Multiple model calls with synthesis

Implementation

class EscalationLevel(Enum):
    DIRECT = 0       # One LLM call
    AUGMENTED = 1    # RAG + one LLM call
    REASONED = 2     # Multi-step with tools
    ENSEMBLE = 3     # Multi-model synthesis


class ProgressiveEscalationAgent:
    """Escalates through complexity levels until result quality is sufficient."""

    def __init__(self, model_client, retriever, quality_evaluator, budget: TokenBudget):
        self.model_client = model_client
        self.retriever = retriever
        self.quality_evaluator = quality_evaluator
        self.budget = budget

    async def execute(self, task: str) -> dict:
        last_result = None

        for level in EscalationLevel:
            if not self.budget.consume(0):  # Check budget still available
                return {"status": "budget_exhausted", "output": last_result}

            # Execute at current escalation level
            result = await self._execute_at_level(task, level)
            last_result = result.get("output")

            # Evaluate quality
            quality = await self._evaluate_quality(task, result)

            if quality >= self._threshold_for_level(level):
                return {
                    "status": "completed",
                    "output": result.get("output"),
                    "escalation_level": level.name,
                    "quality_score": quality,
                }

        # Exhausted all levels — return best result
        return {
            "status": "max_escalation_reached",
            "output": last_result,
            "escalation_level": EscalationLevel.ENSEMBLE.name,
        }

    async def _execute_at_level(self, task: str, level: EscalationLevel) -> dict:
        if level == EscalationLevel.DIRECT:
            return await self._direct_answer(task)
        elif level == EscalationLevel.AUGMENTED:
            return await self._augmented_answer(task)
        elif level == EscalationLevel.REASONED:
            return await self._reasoned_answer(task)
        elif level == EscalationLevel.ENSEMBLE:
            return await self._ensemble_answer(task)

    async def _direct_answer(self, task: str) -> dict:
        response = await self.model_client.chat(
            messages=[{"role": "user", "content": task}],
            model="gpt-4o-mini",
            max_tokens=500,
        )
        return {"output": response["content"], "model": "gpt-4o-mini"}

    async def _augmented_answer(self, task: str) -> dict:
        context = await self.retriever.search(task, top_k=5)
        prompt = f"Use this context to answer:\n\nContext:\n{context}\n\nQuestion: {task}"
        response = await self.model_client.chat(
            messages=[{"role": "user", "content": prompt}],
            model="gpt-4o",
            max_tokens=1000,
        )
        return {"output": response["content"], "model": "gpt-4o", "context_used": True}

    async def _reasoned_answer(self, task: str) -> dict:
        # Multi-step with tool use
        agent = BudgetEnforcedAgent(
            self.model_client,
            TokenBudget(total_budget=self.budget.total_budget // 3),
            tools=self.model_client.get_tools(),
        )
        result = await agent.run(task, max_iterations=6)
        return {"output": result.get("output"), "iterations": result.get("iterations")}

    async def _ensemble_answer(self, task: str) -> dict:
        # Multiple models, synthesized
        responses = await asyncio.gather(
            self._direct_answer(task),
            self._augmented_answer(task),
        )

        synthesis_prompt = f"""Two responses to the same question:

Response 1: {responses[0].get('output')}
Response 2: {responses[1].get('output')}

Question: {task}

Synthesize the best answer, noting where responses agree and disagree."""

        response = await self.model_client.chat(
            messages=[{"role": "user", "content": synthesis_prompt}],
            model="claude-sonnet",
            max_tokens=1500,
        )
        return {"output": response["content"], "model": "claude-sonnet", "ensemble": True}

    def _threshold_for_level(self, level: EscalationLevel) -> float:
        thresholds = {
            EscalationLevel.DIRECT: 0.85,
            EscalationLevel.AUGMENTED: 0.80,
            EscalationLevel.REASONED: 0.75,
            EscalationLevel.ENSEMBLE: 0.70,
        }
        return thresholds[level]

    async def _evaluate_quality(self, task: str, result: dict) -> float:
        """Evaluate result quality using a lightweight scoring model."""
        eval_prompt = f"""Score this response from 0.0 to 1.0 for quality, accuracy, and completeness.

Question: {task}
Response: {result.get('output', '')}

Return ONLY a number between 0.0 and 1.0."""

        response = await self.model_client.chat(
            messages=[{"role": "user", "content": eval_prompt}],
            model="gpt-4o-mini",
            max_tokens=10,
        )
        try:
            return float(response["content"].strip())
        except ValueError:
            return 0.0
Enter fullscreen mode Exit fullscreen mode

The quality evaluator itself uses a cheap model, so the cost of evaluation is minimal. The key economic principle: spend $0.001 evaluating to decide whether to spend $0.05 or $0.50 on a more expensive approach.

7. Reference Architecture: Putting It All Together

All five patterns compose into a layered architecture:

┌──────────────────────────────────────────────────────────────┐
│                      Client / User                            │
└─────────────────────────┬────────────────────────────────────┘
                          │
┌─────────────────────────▼────────────────────────────────────┐
│              Request Gateway + Rate Limiter                   │
│  (per-user/per-org budgets, request dedup, auth)             │
└─────────────────────────┬────────────────────────────────────┘
                          │
┌─────────────────────────▼────────────────────────────────────┐
│              Task Classifier (Micro/Standard/Extended)        │
│  (lightweight model call, tier assignment)                    │
└─────────────────────────┬────────────────────────────────────┘
                          │
┌─────────────────────────▼────────────────────────────────────┐
│           Progressive Escalation Engine                       │
│  (starts cheap, escalates on quality failure)                 │
└─────────────────────────┬────────────────────────────────────┘
                          │
┌─────────────────────────▼────────────────────────────────────┐
│           Token Budget + Circuit Breaker                      │
│  (global budget, velocity limits, state machine)              │
└─────────────────────────┬────────────────────────────────────┘
                          │
          ┌───────────────┼───────────────┐
          ▼               ▼               ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Semantic     │ │ Sandboxed    │ │ Sub-Agent    │
│ Cache        │ │ Tool Registry│ │ Spawner      │
│ (dedup)      │ │ (limits)     │ │ (concurrency)│
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘
       │                │                │
       ▼                ▼                ▼
┌────────────────────────────────────────────────────────────┐
│              LLM Provider(s) + Tool Backends                │
└────────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The request flows through each layer, and each layer has the authority to short-circuit, reduce scope, or reject the request. This defense-in-depth approach ensures that no single failure mode can cause unbounded spending.

8. Implementation: The Sandbox Orchestrator

Here's the complete orchestrator that wires all patterns together:

import asyncio
import time
import uuid
from dataclasses import dataclass, field
from typing import Any, Callable, Optional


@dataclass
class AgentRequest:
    id: str = field(default_factory=lambda: str(uuid.uuid4()))
    user_id: str = ""
    task: str = ""
    context: dict = field(default_factory=dict)
    priority: str = "normal"  # low, normal, high, critical
    created_at: float = field(default_factory=time.time)


@dataclass
class AgentResult:
    request_id: str
    status: str  # completed, budget_exhausted, timeout, error
    output: Optional[str] = None
    tier: str = ""
    escalation_level: str = ""
    tokens_used: int = 0
    tools_called: int = 0
    elapsed_seconds: float = 0.0
    cost_estimate_usd: float = 0.0
    metadata: dict = field(default_factory=dict)


class SandboxOrchestrator:
    """Main orchestrator combining all budget-control patterns."""

    def __init__(
        self,
        model_client,
        retriever,
        tool_registry: ToolRegistry,
        semantic_cache: SemanticCache,
        org_budget_per_month: float = 1000.0,
    ):
        self.model_client = model_client
        self.retriever = retriever
        self.tool_registry = tool_registry
        self.semantic_cache = semantic_cache
        self.org_budget_per_month = org_budget_per_month
        self.tier_classifier = TierClassifier(model_client)
        self._monthly_spent = 0.0
        self._request_count = 0
        self._cost_log: list[dict] = []

    async def process(self, request: AgentRequest) -> AgentResult:
        start = time.time()

        # Step 1: Check organizational budget
        if self._monthly_spent >= self.org_budget_per_month:
            return AgentResult(
                request_id=request.id,
                status="org_budget_exhausted",
                metadata={"monthly_spent": self._monthly_spent, "limit": self.org_budget_per_month},
            )

        # Step 2: Check semantic cache
        cached = await self.semantic_cache.get_or_compute(
            prompt=request.task,
            model="cache-check",
            compute_fn=lambda: "CACHE_MISS",
        )
        if cached != "CACHE_MISS":
            return AgentResult(
                request_id=request.id,
                status="completed",
                output=cached,
                metadata={"cache_hit": True, "elapsed": time.time() - start},
            )

        # Step 3: Classify tier
        tier = self.tier_classifier.classify(request.task, request.context)
        tier_config = TIER_REGISTRY[tier]

        # Step 4: Create per-request token budget
        # Scale budget by priority
        priority_multiplier = {"low": 0.5, "normal": 1.0, "high": 1.5, "critical": 2.0}
        request_budget = TokenBudget(
            total_budget=int(tier_config.max_tokens_total * priority_multiplier.get(request.priority, 1.0)),
            max_velocity_tokens_per_second=200 if tier == ComputeTier.MICRO else 500,
        )

        # Step 5: Execute with progressive escalation
        escalation_agent = ProgressiveEscalationAgent(
            model_client=self.model_client,
            retriever=self.retriever,
            quality_evaluator=self.model_client,
            budget=request_budget,
        )

        try:
            result = await asyncio.wait_for(
                escalation_agent.execute(request.task),
                timeout=tier_config.timeout_seconds,
            )
        except asyncio.TimeoutError:
            result = {"status": "timeout", "output": None}

        # Step 6: Calculate and track cost
        elapsed = time.time() - start
        cost = self._estimate_cost(request_budget.consumed, tier)
        self._monthly_spent += cost
        self._request_count += 1
        self._cost_log.append({
            "request_id": request.id,
            "user_id": request.user_id,
            "cost": cost,
            "tokens": request_budget.consumed,
            "tier": tier.value,
            "elapsed": elapsed,
        })

        return AgentResult(
            request_id=request.id,
            status=result.get("status", "error"),
            output=result.get("output"),
            tier=tier.value,
            escalation_level=result.get("escalation_level", ""),
            tokens_used=request_budget.consumed,
            tools_called=self.tool_registry.get_all_stats(),
            elapsed_seconds=elapsed,
            cost_estimate_usd=cost,
            metadata=result.get("metadata", {}),
        )

    def _estimate_cost(self, tokens: int, tier: ComputeTier) -> float:
        """Rough cost estimation based on tier's typical model."""
        # Approximate pricing (update for your provider)
        cost_per_1k_tokens = {
            ComputeTier.MICRO: 0.0006,   # gpt-4o-mini
            ComputeTier.STANDARD: 0.005,  # gpt-4o
            ComputeTier.EXTENDED: 0.015,  # claude-sonnet
        }
        return (tokens / 1000) * cost_per_1k_tokens.get(tier, 0.005)

    def get_budget_report(self) -> dict:
        """Generate budget utilization report."""
        avg_cost = sum(c["cost"] for c in self._cost_log) / max(len(self._cost_log), 1)
        return {
            "monthly_spent": self._monthly_spent,
            "monthly_limit": self.org_budget_per_month,
            "utilization_pct": (self._monthly_spent / self.org_budget_per_month) * 100,
            "total_requests": self._request_count,
            "avg_cost_per_request": avg_cost,
            "cache_hit_rate": self.semantic_cache.hit_rate,
            "projected_monthly_cost": avg_cost * 30 * 24 * 60,  # naive projection
        }
Enter fullscreen mode Exit fullscreen mode

9. Operational Metrics and Alerting

Budget control is useless without observability. These are the metrics you should instrument and alert on:

Key Metrics

Metric Alert Threshold Action
agent.tokens_per_request.p95 > 2x median Investigate runaway loops
agent.budget_exhausted.rate > 5% of requests Review tier classification
agent.tool_call_failures.rate > 10% Check tool health
agent.circuit_breaker.open.count Any Immediate investigation
org.monthly_budget.utilization > 80% Warn stakeholders
semantic_cache.hit_rate < 10% Cache may need tuning
agent.velocity.throttled.count Any Check for retry storms

Alerting Configuration Example

# Prometheus alerting rules for agent budget monitoring
 groups:
 - name: agent-budget-alerts
   rules:
   - alert: HighBudgetUtilization
     expr: |
       agent_monthly_spend / agent_monthly_budget > 0.80
     for: 5m
     labels:
       severity: warning
     annotations:
       summary: "Agent monthly budget at {{ $value | humanizePercentage }}"

   - alert: RunawayAgentLoop
     expr: |
       rate(agent_tokens_consumed_total[5m]) > 2000
     for: 2m
     labels:
       severity: critical
     annotations:
       summary: "Token velocity exceeded: {{ $value }} tokens/sec"

   - alert: BudgetExhaustionRate
     expr: |
       rate(agent_budget_exhausted_total[1h]) / rate(agent_requests_total[1h]) > 0.05
     for: 10m
     labels:
       severity: warning
     annotations:
       summary: "5%+ of requests hitting budget limits"
Enter fullscreen mode Exit fullscreen mode

Cost Attribution Dashboard

For multi-tenant systems, cost attribution per user/team/project is essential. The _cost_log in the orchestrator provides the raw data. Aggregate it into a dashboard that shows:

  • Cost per user, per day
  • Cost per feature/workflow
  • Token consumption trends over time
  • Cache effectiveness by request type
  • Tier distribution (what % of requests hit each tier)

10. Frequently Asked Questions

How do I set the right token budget per tier?

Start by profiling your actual workloads. Run 100-200 representative requests through each tier and record the p50, p90, and p99 token consumption. Set your tier budgets at approximately the p95 of actual usage. This allows 95% of requests to complete within budget while containing the 5% of outliers that would otherwise spiral. Adjust over time as your workload distribution shifts.

What's the impact of semantic caching on answer quality?

With a similarity threshold of 0.92+, semantic caching has minimal quality impact because it only returns cached results for nearly identical queries. The main risk is returning a stale answer when the underlying data has changed. Mitigate this with TTL-based expiration (1 hour for volatile data, 24 hours for stable data) and by excluding queries that contain timestamps or version-specific references from caching.

Should I use these patterns for internal tools or only customer-facing apps?

Use them everywhere. Internal tools often have less budget scrutiny, which makes them the most likely place for silent cost overruns to accumulate. An internal agent that runs unattended overnight with no budget limits can generate more waste than a customer-facing agent with per-request limits. The same architectural patterns apply; the only difference is that internal tools can afford higher per-request budgets and more generous escalation thresholds since they're not competing for a shared organizational budget with paying customers.


The fundamental principle across all five patterns is defense in depth for compute resources. No single mechanism is sufficient — the circuit breaker catches velocity spikes, tier classification prevents over-provisioning, semantic caching eliminates redundancy, progressive escalation minimizes unnecessary complexity, and sandboxed tool execution contains individual failure modes. Together, they form a system where agent autonomy and budget safety coexist.

For more patterns and architectural deep dives on AI systems, see Tamiz's Insights and explore the broader engineering landscape at tamiz.pro.

Top comments (0)