DEV Community

Cover image for An agent that refuses: building a payout agent with neatlogs, Entire and cfo.ai
Edy Cu
Edy Cu

Posted on

An agent that refuses: building a payout agent with neatlogs, Entire and cfo.ai

I won a $7,000 hackathon prize. The payout form offered me Wise. Wise's own help page says that personal accounts with an address in Indonesia "can no longer hold money or currencies in their account."

That is the kind of mistake I built Cair to prevent: an agent that plans how a prize or client payment reaches your bank account, and refuses routes that cannot deliver. It was built for neatHack, a 48-hour agent hackathon, with neatlogs (traces), Entire (a local code graph) and cfo.ai (the business plan). Code: github.com/edycutjong/cair. Live: app.cair.edycu.dev.

This post is about the part I did not expect to matter most: how little the model is allowed to decide.

The baseline is a confident paragraph

I started by measuring the obvious answer: one call to the model, no tools. I froze 48 scenarios (3 countries x 4 payers x 3 amounts, plus 12 leading prompts) and computed the correct labels from a sourced rules sheet before writing any planner code. Then I ran the one-call baseline three times.

v1: one call v2: Cair
Recommends a route its own provider rules out 18 / 48, every run 0 / 48
Contradicts a sourced tax fact 8 / 48, every run 0 / 48
Says UNCLEAR where no primary source exists 11-23 / 143 fields 143 / 143
Figures traceable to a tool output n/a, no tools 359 / 359

The baseline recommended Wise for every Indonesian Devpost winner in all three runs. A second baseline that gets the entire rules sheet pasted into its prompt still recommended a dead route twice and stated an unsupported tax position 11 times. The model had the facts in front of it and still wrote a confident answer around the gaps.

(The caveat that matters: this measures fidelity to the sourced sheet, not tax correctness. The labels and the agent read the same sheet.)

So the code decides, and the model only words it

Cair runs a fixed evidence plan before any prose: what the payer offers, whether each offered option can receive in your country, the US treaty position, the ECB reference rate, then Decimal arithmetic. Each step is a neatlogs span, so every plan is one trace:

@neatlogs.span(kind="TOOL", tool_name="fx_quote", description="ECB daily reference rate USD -> local currency")
def fx_quote(quote_ccy: str) -> dict:
    _maybe_fail("fx_quote")
    return fx.live(quote_ccy)
Enter fullscreen mode Exit fullscreen mode

The route and W-8BEN line 10 come from sourced rows, never from the model. The model writes the explanation, and then a provenance guardrail checks it: any money amount, rate or treaty article that is not in a tool output blocks the answer, which gets one retry and then a template built from the tool outputs.

with neatlogs.trace("provenance_guardrail", kind="GUARDRAIL") as sp:
    verdict = check(text, tool_outputs, card, question)
    sp.set_attribute("neatlogs.guardrail.passed", verdict["passed"])
    sp.set_attribute("neatlogs.guardrail.score", float(verdict["score"]))
Enter fullscreen mode Exit fullscreen mode

Because the guardrail is a span, every blocked answer is searchable in neatlogs. That turned out to be the debugging tool.

What trace search found

search_traces "provenance_guardrail BLOCKED" over the neatlogs MCP server returned the v2 plans whose explanation had been rejected. get_trace_context showed the model's text next to the guardrail verdict. Two real bugs came out of it:

  1. A correct answer was blocked because my own figure rule read 2000.00 USD as a year. Any amount from 1,900 to 2,099 written without a thousands separator was affected. A false block, not a false pass, but I would not have seen it without the span.
  2. Refusals that repeated the user's own "0%" ("I can't say the withholding is 0%") were blocked twice and fell back to the template. I considered allowing figures that appear in the question inside negated sentences, and rejected it: it would let "the withholding is not 30%, it is 0%" through.

Both are fixed, each pinned by a test named for the defect, and re-run. The write-up with the trace ids is in the repo (evidence/rca-01.md).

Entire's local code graph was the second tool in that loop: after the year fix, entire graph search given the bug in plain words ranked the corrected rule first and its test second. It was also wrong in an instructive way: impact on rail_eligibility reported 0 callers, probably because the planner passes tools as function values (att.call("fx_quote", tools.fx_quote, ...)), which the graph does not follow. I recorded that as found rather than hiding it.

When a tool fails

The hackathon asked for a failed-action recovery run, so I made one on purpose. scripts/eval.py --planner v2 --fail fx_quote makes the FX tool raise on every call:

for attempt in (1, 2):
    try:
        return fn(*args)
    except Exception as e:  # the failed span is already marked ERROR by the SDK
        ...
return fallback(last) if fallback else None
Enter fullscreen mode Exit fullscreen mode

The agent retries once, then falls back to the last ECB rate it fetched, stamped stale with its date, and finishes. In the 24 USD-payer scenarios that meant 24 fallbacks; all 48 plans completed, 0 undeliverable routes, 350 / 350 figures sourced. The neatlogs trace shows two ERROR spans and the fallback span, and the dashboard footer says "ended with error" because it counts the errored tool spans, while the workflow and agent spans succeeded. I kept that on the page instead of tidying it. The agent never asks the model for a rate.

What it costs

Measured tokens at list price: $0.0033 of model cost per plan, p50 5.50 s and p95 10.31 s on my laptop. In cfo.ai the base scenario (assumed prices, planned and not built) costs $18.887 a month to run, and the $25 AWS credit lasts 1.35 months before any revenue. The model is shared here: cfo.ai/s/PQ2HUGmiJl3P0RQB.

Limits

  • Three countries are fully sourced (Indonesia, India, the Philippines). Everything else is UNCLEAR by design.
  • The guardrail is rule-based. On a fresh held-out set it blocked 62 / 66 violations and 3 / 66 honest sentences, before I fixed its misses. Expect new phrasings to slip through at roughly that rate, which is why the route and every amount are decided by code.
  • Payoneer's receiving fee is not published as one number, so "what lands" is an upper bound.
  • Receipts so far are my own: one replay of a known payout (at most $4,900 promised, $4,897 arrived) and one test receipt. Forward receipts from real users: none yet.
  • Cair explains forms; it never files them. Not tax advice.

Repo, evidence and the commands to reproduce every number: github.com/edycutjong/cair.

Top comments (0)