DEV Community

Cover image for OpenRig: Multi-Agent Harness Architecture for Persistent Terminal Coordination
mech.app
mech.app

Posted on Originally published at mech.app

OpenRig: Multi-Agent Harness Architecture for Persistent Terminal Coordination

OpenRig is a harness-of-harnesses that runs multiple AI coding agents as a coordinated team inside tmux sessions. Instead of juggling separate terminal windows for Claude Code and Codex, you define a team in YAML, boot it with one command, and get persistent agent state at stable addresses.

The project hit #2 on GitHub Trending for TypeScript with 2,314 stars. It represents a shift from single-agent workflows to persistent team structures where a lead agent delegates to specialists and maintains shared context across restarts.

Architecture: Harness Wrapping and Session Management

OpenRig introduces three layers:

  1. Harness layer: Wraps individual agents (Claude Code, Codex) with tool boundaries and context management.
  2. Rig layer: Coordinates multiple harnesses as a team, defined in YAML.
  3. Session layer: Uses tmux to maintain persistent state and allow reconnection without losing work.

The core pattern is a lead agent that receives high-level instructions, delegates to specialist agents, and surfaces decisions that need human input. Each agent runs in its own tmux pane with isolated context but shared workspace access.

YAML Team Definition

Teams are declared in a rig.yaml file that specifies:

  • Seats: Named agent instances with assigned models and roles
  • Lead agent: The coordination point for human interaction
  • Specialists: Agents with specific skills (testing, documentation, refactoring)
  • Workspace boundaries: Shared repository access with tool restrictions

Example structure:

rig:
  name: build-team
  workspace: /path/to/repo

seats:
  - name: lead
    model: claude-3-5-sonnet
    role: coordinator
    tools:
      - delegate
      - summarize
      - request-decision

  - name: tester
    model: gpt-4
    role: specialist
    tools:
      - run-tests
      - write-tests
      - read-coverage

  - name: documenter
    model: claude-3-5-sonnet
    role: specialist
    tools:
      - read-code
      - write-docs
      - update-readme
Enter fullscreen mode Exit fullscreen mode

The lead agent can call delegate(task, specialist) to route work, then aggregate results before surfacing them to the human operator.

State Management and Context Boundaries

OpenRig solves three persistence problems:

  1. Session persistence: tmux keeps agents running even when you disconnect. Reattach to the same session and continue work.
  2. Context persistence: Each agent maintains its own conversation history and working memory in SQLite, keyed by seat name.
  3. Workspace persistence: All agents share the same repository but operate through tool boundaries that prevent conflicting file writes.

Context Isolation

Each seat gets:

  • A dedicated tmux pane
  • An isolated conversation history
  • A tool allowlist defined in YAML
  • Read access to the full workspace
  • Write access gated by tool permissions

When the lead agent delegates a task, it passes a scoped context (file paths, test results, error logs) rather than the full conversation history. This prevents context window overflow and keeps specialists focused.

Restart Behavior

If an agent crashes or hits a rate limit:

  1. The rig detects the failure via tmux pane exit status.
  2. It restarts the harness with the same seat configuration.
  3. The agent reloads its conversation history from SQLite.
  4. Work resumes from the last committed state (file changes, test runs).

This is different from stateless agent loops that lose context on restart. OpenRig treats each seat as a long-running process with durable state.

Coordination Primitives

The lead agent uses three coordination tools:

1. Delegate

Routes a task to a specialist with scoped context:

await delegate({
  specialist: 'tester',
  task: 'Run integration tests for auth module',
  context: {
    files: ['src/auth/**/*.ts'],
    previous_results: testOutput
  }
})
Enter fullscreen mode Exit fullscreen mode

The specialist executes, then returns a summary. The lead agent decides whether to surface the result to the human or delegate follow-up work.

2. Request Decision

Surfaces a choice to the human operator:

await requestDecision({
  question: 'Tester found 3 failing tests. Fix now or document as known issues?',
  options: ['fix', 'document', 'defer'],
  context: failureDetails
})
Enter fullscreen mode Exit fullscreen mode

This prevents agents from making architectural decisions or merging breaking changes without approval.

3. Summarize

Aggregates specialist outputs into a human-readable report:

await summarize({
  inputs: [testerResult, documenterResult],
  format: 'markdown',
  include_next_steps: true
})
Enter fullscreen mode Exit fullscreen mode

The summary becomes the primary interface for humans monitoring the rig.

Security and Tool Boundaries

OpenRig does not sandbox agents at the OS level. Instead, it uses tool allowlists to control what each seat can do.

Risk Mitigation Limitation
Arbitrary code execution Tool allowlist per seat Agents can still call allowed tools with malicious arguments
Conflicting file writes Lead agent coordinates write order No file locking; last write wins
Credential leakage No environment variable access by default Agents can still read .env files if given read-file tool
Runaway costs Per-seat token budgets in YAML Requires manual monitoring; no automatic cutoff
Prompt injection Context scoping limits cross-agent pollution Lead agent can still be tricked into delegating malicious tasks

The security model assumes you trust the models and are running in a development environment. Do not point OpenRig at production infrastructure or give it access to sensitive credentials.

Observability and Debugging

OpenRig exposes three observability layers:

  1. TUI (Terminal UI): Real-time view of all seats, their current task, token usage, and state.
  2. SQLite logs: Per-seat conversation history, tool calls, and errors.
  3. tmux session inspection: Direct access to each agent's terminal output.

To debug a stuck agent:

rig attach build-team  # Opens TUI
# Select the stuck seat
# Press 'd' to view detailed logs
# Press 't' to attach to the tmux pane
Enter fullscreen mode Exit fullscreen mode

The TUI shows:

  • Current task for each seat
  • Token usage and rate limit status
  • Last tool call and result
  • Pending delegations

This is more useful than tailing log files because it surfaces the coordination graph: which agent is waiting on which specialist.

Deployment Shape

OpenRig runs locally on macOS or Linux. It requires:

  • Node.js 22 or 24 (Node.js 22 on Apple Silicon due to a compatibility issue)
  • tmux
  • API keys for Claude and OpenAI

Installation:

npm install -g @openrig/cli
rig setup --dry-run  # Preview file changes
rig setup            # Write provider hooks and workspace trust settings
Enter fullscreen mode Exit fullscreen mode

The setup command modifies:

  • ~/.config/openrig/: Config files and SQLite databases
  • VS Code settings: Adds workspace trust for the rig directory
  • Shell environment: Adds rig command to PATH

There is no server component. Each rig runs as a local tmux session. To share a rig across a team, you would need to run it on a shared development machine and use tmux's multi-attach feature (not officially supported yet).

Failure Modes

1. Context Window Overflow

If the lead agent's conversation history grows too large, it will start dropping early context. This breaks long-running tasks that depend on decisions made hours ago.

Mitigation: Periodically call summarize and checkpoint the summary to a file. Restart the lead agent with the summary as initial context.

2. Delegation Loops

A lead agent can delegate to a specialist, which requests a decision, which the lead agent interprets as needing another delegation. This creates an infinite loop.

Mitigation: Set a max delegation depth in YAML (not yet implemented). Manually kill the loop via the TUI.

3. Conflicting File Writes

Two specialists modify the same file simultaneously. The second write clobbers the first.

Mitigation: The lead agent should sequence writes by delegating to one specialist at a time. No automatic conflict resolution exists.

4. Rate Limit Cascades

If one agent hits a rate limit, the lead agent may repeatedly delegate to it, causing all specialists to queue up and hit their own rate limits.

Mitigation: The rig detects rate limit errors and pauses the affected seat. The lead agent should check seat status before delegating (not yet automatic).

Technical Verdict

Use OpenRig when:

  • You need multiple AI agents working on the same codebase without losing context across sessions.
  • You want a lead agent to coordinate specialists rather than manually orchestrating tool calls.
  • You are comfortable with tmux and local development workflows.
  • You need to pause work, disconnect, and resume later without re-explaining the task.

Avoid OpenRig when:

  • You need production-grade sandboxing or multi-tenancy. The tool boundaries are not security boundaries.
  • You want a web UI or remote collaboration. This is a local terminal tool.
  • You need automatic conflict resolution for concurrent file writes.
  • You are on Windows without WSL2 (not tested) or need Node.js 24 on Apple Silicon (compatibility issue).

The harness-of-harnesses pattern is useful for complex, multi-day tasks where you want persistent agent state and human-in-the-loop decision points. It is not a replacement for single-agent tools like Cursor or Aider for quick edits.

Source Links

Top comments (0)