DEV Community

Cover image for Argo-Bench: Why Enterprise Data Agents Need Multi-Table Workflows, Not Just SQL Generation
mech.app
mech.app

Posted on Originally published at mech.app

Argo-Bench: Why Enterprise Data Agents Need Multi-Table Workflows, Not Just SQL Generation

Most text-to-SQL benchmarks test whether an agent can generate a single SELECT statement. Argo-Bench asks a harder question: can an agent navigate 235 tables, reconstruct hidden business facts, run statistical analyses, and execute actions that change the state of a simulated enterprise?

The answer is no. Frontier models score above 95 on only 34.8% of tasks and average 59.5 points. The gap reveals what breaks when you move from query generation to multi-stage data workflows.

The Problem with Existing Benchmarks

Text-to-SQL benchmarks like Spider and BIRD evaluate query generation in isolation. You get a natural language question, a schema, and a database. The agent writes SQL. The grader compares output to a reference answer.

This setup has three structural flaws:

  • Answer keys are frequently wrong. Audits of popular benchmarks find incorrect ground truth, making it unclear whether a failing agent is broken or correct.
  • Single-table bias. Public datasets fit business events into one table. Real enterprise warehouses spread a single transaction across dozens of normalized tables.
  • No action execution. Agents generate queries but never act on results. Real workflows require banning accounts, allocating budgets, or issuing refunds based on what the query reveals.

Argo-Bench addresses all three by grading actions in a simulator, not SQL strings against answer keys.

Architecture: Simulated World Plus Hidden Ground Truth

Argo-Bench models a food delivery platform in New York City with 81 million orders in 2024. The simulation includes:

  • Grounded economics (pricing, fees, discounts)
  • Fraud patterns (account takeovers, promo abuse)
  • Marketplace incentives (courier bonuses, restaurant promotions)

The simulator exports data to a 235-table ERP warehouse modeled on Oracle E-Business Suite. The warehouse contains 7.5 billion rows.

The critical design choice: the simulator's ground-truth state is withheld from the warehouse the agent sees. Tasks require reconstructing facts by navigating incomplete or denormalized data before acting.

For example, identifying fraudulent accounts requires joining order history, payment methods, device fingerprints, and geolocation logs across multiple tables. The agent must infer fraud patterns, not just query a fraud_flag column.

Task Structure and Grading

Each of the 210 tasks follows this flow:

  1. Natural language task description (e.g., "Ban accounts that placed more than 10 orders from different zip codes in 24 hours")
  2. Agent explores the warehouse (queries, statistical analysis, reasoning across tables)
  3. Agent files an action (ban list, budget allocation, refund batch)
  4. Grader scores the action by its consequences in the simulator

The grader does not compare SQL. It evaluates whether the agent's action produces the correct outcome in the hidden simulation state.

Every task includes an executable reference solution that demonstrates solvability using only the warehouse. This proves the task is not impossible and provides a correctness baseline.

What Breaks in Multi-Table Workflows

The benchmark exposes four failure modes:

Failure Mode Example Root Cause
Schema navigation Agent queries wrong table or misses join path 235 tables exceed context window or reasoning depth
Statistical reasoning Agent calculates median incorrectly or misinterprets percentile LLMs struggle with multi-step numerical logic
State reconstruction Agent assumes data completeness when records are missing No explicit signal that warehouse is incomplete
Action formulation Agent outputs correct analysis but files malformed action Disconnect between reasoning and execution interface

The strongest models fail most often on state reconstruction. They treat the warehouse as authoritative when it is actually a partial, denormalized view of the simulator's hidden state.

Observability and Debugging

Because the grader scores actions in a simulator, you get deterministic feedback. If an agent bans the wrong accounts, you can replay the simulation, inspect the agent's queries, and trace where reasoning diverged from ground truth.

This is harder with traditional benchmarks. If an agent's SQL output does not match the answer key, you cannot tell whether the agent is wrong or the answer key is wrong without manual review.

Argo-Bench provides:

  • Execution logs for every query the agent ran
  • Intermediate state snapshots showing what the agent knew at each step
  • Simulator diffs comparing the agent's action to the reference solution's action

This makes it possible to debug multi-step reasoning failures, not just query syntax errors.

State Management Requirements

Agents need to maintain context across multiple queries and analyses. A typical task requires:

  • Exploring 10-15 tables to understand schema relationships
  • Running 3-5 exploratory queries to validate assumptions
  • Performing statistical aggregations (percentiles, moving averages, cohort analysis)
  • Formulating an action based on the results

This implies:

  • Persistent memory to track which tables have been explored
  • Intermediate result storage for multi-step calculations
  • Hypothesis tracking to avoid re-querying the same data
  • Action staging to validate the action before submission

Most agents tested in the paper use stateless prompting with full conversation history. This works for simple tasks but fails when reasoning depth exceeds context window limits.

Security and Isolation Boundaries

The benchmark runs agents against a read-only warehouse. Agents cannot modify data, only file actions through a controlled API.

This separation is critical for production deployment. Enterprise data agents should never have write access to the warehouse. Instead, they should:

  • Query the warehouse read-only
  • Propose actions through a staging API
  • Wait for human or automated approval
  • Execute actions in a separate transaction layer

Argo-Bench enforces this boundary by design. The agent sees the warehouse but acts through the simulator's API, which validates and scores each action.

Code Example: Reference Solution Pattern

Here is a simplified reference solution for a fraud detection task:

# Step 1: Identify accounts with suspicious order patterns
suspicious_accounts = db.query("""
    SELECT customer_id, COUNT(DISTINCT zip_code) as zip_count
    FROM orders
    WHERE order_date >= CURRENT_DATE - INTERVAL '24 hours'
    GROUP BY customer_id
    HAVING COUNT(DISTINCT zip_code) > 10
""")

# Step 2: Cross-reference with payment anomalies
fraud_candidates = db.query("""
    SELECT DISTINCT o.customer_id
    FROM orders o
    JOIN payment_methods pm ON o.payment_id = pm.id
    WHERE o.customer_id IN ({})
    AND pm.card_country != o.delivery_country
""".format(','.join(str(id) for id in suspicious_accounts['customer_id'])))

# Step 3: File ban action
action = {
    "type": "ban_accounts",
    "account_ids": fraud_candidates['customer_id'].tolist(),
    "reason": "Multi-zip + cross-border payment anomaly"
}
submit_action(action)
Enter fullscreen mode Exit fullscreen mode

The reference solution demonstrates the workflow: explore, reason, act. The grader scores whether the banned accounts match the simulator's ground-truth fraud list.

Deployment Considerations

Running Argo-Bench requires:

  • Warehouse infrastructure (235 tables, 7.5 billion rows)
  • Simulator runtime (to score actions and provide ground truth)
  • Agent execution environment (isolated from production data)

The authors provide:

  • Parquet exports of the warehouse (downloadable)
  • Simulator code (open source)
  • Grading harness (deterministic scoring)

For production use, you would replace the simulated warehouse with your actual ERP schema and build a custom grader that validates actions against business rules instead of a simulator.

Likely Failure Modes

Agents will fail in production when:

  • Schema drift changes table relationships without updating the agent's schema knowledge
  • Data quality issues introduce nulls or duplicates that break statistical assumptions
  • Action APIs change and the agent continues using deprecated endpoints
  • Human approval delays cause the agent's analysis to become stale before the action executes

Argo-Bench does not model these failure modes. It assumes a static schema, clean data, and immediate action execution. Real deployments need monitoring for schema changes, data quality checks, and staleness detection.

Technical Verdict

Use Argo-Bench when:

  • You are building data agents that must navigate complex ERP schemas
  • You need to evaluate multi-step reasoning, not just query generation
  • You want deterministic grading based on action outcomes, not SQL string matching
  • You are designing eval harnesses for enterprise automation workflows

Avoid Argo-Bench when:

  • You only need to test single-query text-to-SQL generation
  • Your data fits in a few tables and does not require multi-step reasoning
  • You cannot run a 7.5 billion row warehouse locally or in CI
  • You need benchmarks that model schema drift, data quality issues, or approval workflows

The benchmark is most useful for teams building agents that orchestrate multi-stage analytics pipelines. It exposes the gap between generating SQL and executing data-driven actions at enterprise scale.

Source Links

Top comments (0)