DEV Community

Cover image for Sarvam Arya: Production Agent Orchestration Stack Exposes State, Routing, and Recovery Plumbing
mech.app
mech.app

Posted on Originally published at mech.app

Sarvam Arya: Production Agent Orchestration Stack Exposes State, Routing, and Recovery Plumbing

Sarvam just released Arya, an orchestration stack designed to move agent systems from prototype to production. The announcement centers on a concrete problem: frontier models can write compilers and rebuild rendering pipelines in demos, but they fail silently and inconsistently when you point them at real workloads like financial report extraction.

The Sarvam team ran a 24-hour sprint extracting 200 fields from public company filings. A single-agent approach with full context collapsed after 15 minutes, completing less than half the extraction. The same model, orchestrated through Arya's multi-agent primitives, finished the job. The difference was not model intelligence. It was structure.

The Orchestration Gap

Most agent frameworks treat orchestration as a state machine problem. You define nodes, edges, and transitions. LangGraph, for example, gives you a graph abstraction where each node is a function and edges represent control flow. This works for demos but breaks in production because:

  • State grows unbounded. Context windows fill with intermediate results, tool outputs, and retry logs.
  • Failures cascade. One tool timeout kills the entire workflow.
  • Observability is an afterthought. You get logs, not causal traces through nested agent calls.
  • Recovery is manual. Agents restart from scratch or hang indefinitely.

Arya addresses these gaps by treating orchestration as an infrastructure problem, not a prompt engineering problem.

Architecture: Layers and Boundaries

Arya separates concerns into four layers:

Layer Responsibility Failure Mode
Routing Decompose tasks, allocate to agent swarms Infinite planning loops, circular delegation
State Management Persist intermediate results, checkpoints State bloat, consistency violations
Tool Execution Invoke external APIs, databases, models Timeouts, rate limits, partial failures
Observability Trace causality, log decisions, measure latency Missing context, unlinked spans

The routing layer exposes primitives for task decomposition and work allocation. Instead of a single agent with full context, you define agent swarms where each agent handles a bounded subtask. The orchestrator manages work queues and result aggregation.

State management uses checkpointing to persist progress. When a tool call fails, the agent resumes from the last checkpoint instead of restarting. This requires explicit state boundaries: what gets persisted, what gets discarded, and how conflicts are resolved when concurrent agents write to the same state.

Tool execution is wrapped in retry logic with exponential backoff and circuit breakers. Arya tracks tool success rates and latency distributions, automatically degrading to fallback tools when primary tools fail consistently.

Observability is built around causal traces. Each agent call, tool invocation, and state transition emits a span with parent-child relationships. You can reconstruct the full execution graph after the fact, including retries and backtracking.

State Persistence and Recovery

The financial extraction example reveals how state persistence works in practice. Extracting 200 fields from 100 companies means 20,000 data points. A single-agent approach loads all context into one session. When the agent hits a rate limit or timeout, you lose all progress.

Arya's approach:

  1. Decompose the task into per-company subtasks.
  2. Allocate each subtask to an agent in the swarm.
  3. Checkpoint after each field extraction.
  4. Aggregate results into a shared state store.

When an agent fails mid-extraction, the orchestrator reassigns the subtask to another agent. The new agent loads the checkpoint and resumes from the last completed field. This requires a state schema that defines:

  • Granularity: What constitutes a checkpoint (per-field, per-section, per-document)?
  • Consistency: How do you handle concurrent writes to the same field?
  • Expiration: When do you garbage-collect old checkpoints?

Arya uses a key-value store with versioned writes. Each checkpoint includes a timestamp and parent checkpoint ID. Concurrent writes trigger a conflict resolution policy (last-write-wins, merge, or fail).

Routing Primitives: Beyond State Machines

Traditional agent frameworks model routing as a directed graph. You define nodes (agent actions) and edges (transitions). This works for linear workflows but breaks when you need dynamic routing based on intermediate results.

Arya introduces three routing primitives:

  • Swarm allocation: Distribute work across multiple agents with load balancing.
  • Conditional branching: Route based on tool outputs, not just state transitions.
  • Backtracking: Retry failed subtasks with different agents or tools.

Here's a simplified example of swarm allocation:

from arya import Orchestrator, Agent, Task

# Define agents with bounded responsibilities
extractor = Agent(
    name="field_extractor",
    tools=["parse_pdf", "extract_table", "validate_units"],
    max_context_tokens=8000
)

validator = Agent(
    name="data_validator", 
    tools=["cross_reference", "check_consistency"],
    max_context_tokens=4000
)

# Create orchestrator with routing policy
orchestrator = Orchestrator(
    agents=[extractor, validator],
    routing_policy="round_robin",
    checkpoint_granularity="per_field"
)

# Define task with decomposition strategy
task = Task(
    input_documents=company_filings,
    output_schema=extraction_schema,
    decomposition="per_company"
)

# Execute with automatic retry and recovery
result = orchestrator.execute(
    task=task,
    max_retries=3,
    timeout_per_subtask=300
)
Enter fullscreen mode Exit fullscreen mode

The orchestrator decomposes the task into per-company subtasks, allocates them to the extractor swarm, and routes validated results to the validator agent. If a subtask fails, the orchestrator retries with exponential backoff or routes to a fallback agent.

Observability: Tracing Causality

Production agent systems generate thousands of spans per workflow. Without causal tracing, you cannot debug failures or optimize performance. Arya's observability model links spans hierarchically:

  • Workflow span: Top-level task execution.
  • Agent spans: Individual agent invocations.
  • Tool spans: External API calls, database queries, model inference.
  • State spans: Checkpoint writes, reads, and conflicts.

Each span includes:

  • Parent span ID: Links to the calling span.
  • Timing: Start, end, and duration.
  • Metadata: Agent name, tool name, input/output sizes.
  • Errors: Exception type, stack trace, retry count.

This lets you answer questions like:

  • Which tool calls are the bottleneck?
  • How many retries did this workflow require?
  • Where did the agent backtrack?
  • What was the state at each checkpoint?

Arya exports traces in OpenTelemetry format, so you can ingest them into Jaeger, Honeycomb, or Datadog.

Isolation and Multi-Tenancy

Production deployments run multiple agent workflows concurrently. Without isolation, one workflow can starve others by consuming all tool quota or filling the state store.

Arya enforces isolation at three levels:

Boundary Mechanism Trade-off
Compute Per-workflow resource limits (CPU, memory) Overhead from containerization
State Namespaced key-value stores Increased storage cost
Tools Per-workflow rate limits and quotas Reduced throughput for bursty workloads

Compute isolation uses lightweight containers or process groups. Each workflow gets a dedicated execution environment with CPU and memory limits. This prevents runaway agents from crashing the orchestrator.

State isolation uses namespaced keys. Each workflow writes to a separate partition in the state store. This prevents cross-workflow conflicts but increases storage cost because checkpoints cannot be deduplicated.

Tool isolation enforces per-workflow rate limits. If a workflow exhausts its quota, subsequent tool calls fail fast instead of blocking other workflows. This requires coordination with external APIs that support per-tenant quotas.

Failure Modes and Recovery

Arya's design assumes failures are common, not exceptional. The orchestrator handles:

  • Tool timeouts: Retry with exponential backoff or route to fallback tool.
  • Rate limits: Queue requests and retry after cooldown.
  • Partial failures: Resume from last checkpoint instead of restarting.
  • State conflicts: Resolve using last-write-wins or merge policies.
  • Agent crashes: Reassign subtask to another agent in the swarm.

The financial extraction example hit all of these. PDF parsing timed out on large documents. Table extraction hit rate limits on the OCR API. Field validation found inconsistent units across documents. The orchestrator retried with smaller page ranges, queued OCR requests, and flagged validation errors for human review.

Recovery is not automatic. Arya provides hooks for custom recovery logic:

def handle_extraction_failure(subtask, error, retry_count):
    if isinstance(error, TimeoutError) and retry_count < 3:
        # Reduce page range and retry
        subtask.page_range = (subtask.page_range[0], subtask.page_range[0] + 10)
        return "retry"
    elif isinstance(error, RateLimitError):
        # Queue for later retry
        return "queue"
    else:
        # Escalate to human review
        return "escalate"

orchestrator.register_error_handler(handle_extraction_failure)
Enter fullscreen mode Exit fullscreen mode

This lets you encode domain-specific recovery strategies without modifying the orchestrator.

Deployment Shape

Arya runs as a stateful service with three components:

  • Orchestrator: Manages workflow execution, routing, and recovery.
  • State store: Persists checkpoints and intermediate results (Redis, PostgreSQL, or DynamoDB).
  • Observability backend: Ingests traces and metrics (OpenTelemetry collector).

The orchestrator is horizontally scalable. Each instance handles a subset of workflows using consistent hashing. The state store must support distributed transactions for checkpoint writes. The observability backend must handle high cardinality (thousands of spans per workflow).

Sarvam recommends deploying Arya on Kubernetes with:

  • 3+ orchestrator replicas for high availability.
  • Managed state store (AWS ElastiCache, Google Cloud Memorystore) for durability.
  • Dedicated observability cluster to avoid impacting workflow latency.

Technical Verdict

Use Arya when:

  • You need multi-step agent workflows that cannot fit in a single context window.
  • Failures are common and recovery must be automatic.
  • You need causal traces to debug and optimize agent behavior.
  • You are deploying multi-tenant agent systems with isolation requirements.

Avoid Arya when:

  • Your workflows fit comfortably in a single agent session.
  • You are prototyping and do not need production-grade recovery.
  • You cannot tolerate the operational overhead of running a stateful orchestration service.
  • Your workloads are batch-oriented and do not require real-time orchestration.

Arya fills the gap between toy agent frameworks and production systems. It exposes the plumbing (state, routing, recovery, observability) that most frameworks treat as an afterthought. The trade-off is operational complexity. You are running a distributed system with stateful components, not just calling an LLM API.

Source Links

Top comments (0)