DEV Community

Cover image for sprix-sage-router factors half-finished work into who finishes the task
Reno Lu
Reno Lu

Posted on

sprix-sage-router factors half-finished work into who finishes the task

A telling number in the sprix-sage-router README is not a utility score at all. It is the wasted work column. In the authors' own mid-execution replay, progress-aware SAGE edges its progress-masked twin on utility by 0.298 to 0.291, with both utility values reported as plus or minus 0.021, but it cuts wasted work from 0.104 to 0.059 and reports a 58.4% switch rate instead of 80.1%. That is the bet this project makes: once a task is underway, the work already done should change who finishes it.

The README describes the repository as an open-source research output of Sprix AI. It carries a Research Preview status badge and ships a Python 3.10+ reference implementation with no runtime dependencies.

The question discovery leaves open

The README frames the problem plainly. Agent discovery tells a system which agents exist. It does not say who should work with whom after execution has already begun. SAGE, short for State-Aware Graph Exchange, is pitched as a decision layer above the Agent2Agent (A2A) protocol, sitting between discovery and task execution.

It weighs three routes in one objective. SELF keeps the incumbent agent on the job. COLLABORATE keeps the incumbent as owner but recruits a small complementary team to cover missing requirements. HANDOFF gives full ownership to a peer, which the README says fits when specialist advantage exceeds context-transfer loss.

What makes it checkpoint-aware is the input. The router accounts for completed DAG nodes, reusable artifacts, observed partial quality, remaining work, failures, budget, and deadline. The central idea is a reuse fraction. If the current owner has finished part of a requirement and keeps it, that progress counts in full. If ownership changes, only the share that survives artifact portability carries over. The same fraction lowers projected remaining cost and duration. In the quick-start example, planning is complete with a transferability of 0.95, while coding is 35% in flight with a transferability of 0.40, so moving the coding work to the stronger coder does not come free.

Feasible routes are ranked by a linear utility that starts from a learned success probability and subtracts penalties including cost, latency, context-transfer loss, and coordination overhead, with terms for uncertainty-aware exploration. The authors say outright that noisy-OR, beam search, Beta beliefs, online logistic regression, and DAG-induced communication edges are not claimed as inventions. They call them replaceable implementation mechanisms.

What the evaluation shows, and what it does not

The README splits evaluation into three benchmark questions rather than leaning on one table, and it labels their nature.

The mid-execution replay in benchmark_dynamic.py covers 1,000 checkpoints over five seeds. The authors state that its evaluator scores artifact portability and remaining work without calling the router's switch-loss code, and that every policy gets the same registry, action space, budget, and deadline. The "Always hand off" baseline lands at 0.290 utility with a 100% switch rate and 0.130 wasted work. "Always continue" drops to 0.085 with a 34.4% deadline miss rate. A hidden-state dynamic oracle reaches 0.375, so by the authors' own numbers there is headroom above SAGE.

A controlled intervention changes only in-flight completion, from 0.0 to 0.9. The progress-aware minus progress-masked utility gap grows from 0.0000 to +0.0684. The caption on that figure says the values are synthetic five-seed means, not real-endpoint evidence.

A second benchmark tests requirement-conditioned trust. On the same evidence stream in a heterogeneous specialist setting, per-requirement trust reports a Brier score of 0.0125 and selection regret of 0.0094, against 0.0355 and 0.0310 for a single reputation score. In a homogeneous negative control both have zero routing regret, which the authors offer as a sign the benchmark does not manufacture an advantage when specialization is absent.

A third suite, benchmark.py, runs 2,500 independent tasks. The README describes it as regression and learning ablation, not support for the mid-execution claim.

Where the prototype stops

The A2A layer, sprix_a2a.py, validates declared skill IDs against locally calibrated evidence and turns a selected route into a transport-neutral ExecutionPlan covering ownership, assignments, DAG dependencies, communication edges, estimated resources, and rationale. The README says the current prototype intentionally does not transmit tasks, authenticate endpoints, or verify signatures. Your A2A client still handles message/send, streaming, polling, cancellation, and secure artifact handling.

Read it as the research preview its status badge says it is. If you run multi-agent pipelines where a reassignment throws away half-finished work, the reuse fraction and the wasted-work metric are the pieces worth borrowing for your own evaluation. The router exposes record_outcome to feed back execution results and export_state to persist learned evidence.


GitHub: https://github.com/wang2122/sprix-sage-router


Curated by Agent Palisade — practical AI for small and mid-sized businesses.

Top comments (0)