DEV Community

Cover image for I built a Qwen 3.8 Max agent that decides which ERC-8004 agents to trust, and teaches itself to say 'not enough data'
umar idris
umar idris

Posted on

I built a Qwen 3.8 Max agent that decides which ERC-8004 agents to trust, and teaches itself to say 'not enough data'

AI agents are starting to hold wallets, take jobs and rate each other. ERC-8004 gives them an onchain identity and a public place to collect feedback. But when I looked at the real data on Monad testnet, I found a problem: the registry can tell you what people said about an agent, but not whether you should trust it.

So I built AgentCredit: a Qwen 3.8 Max agent that investigates an ERC-8004 agent, works out a trust score from the raw feedback, and can write its verdict onchain with a hash of its evidence. It was built for the Monad hackathon (Trust, Identity & AI Infrastructure track) and the Alibaba Cloud Best Builds with Qwen 3.8 Max bounty.

[

: the Trust Checker result for agent 10, with the reasoning steps and the score breakdown]

The problem: one number that means different things

ERC-8004's Reputation Registry lets anyone leave feedback for an agent: a number, plus tags saying what it is. There's no standard for what the number means. I scanned the first 100 registered agents on Monad testnet and looked at what the feedback really contains:

  • Agent #9 (ScavBot) has a registry summary of 1439. Its feedback is Elo ratings, which aren't a 0-100 score at all.
  • Agents #10 and #20 both show a summary of 0. But #10's feedback is 33 wins and 9 losses, and #20's is 2 wins and 21 losses. One is a decent performer and the other is a disaster, and the summary number hides it. The registry appears to average values that are +1 for a win and -1 for a loss, then round, so both land on 0.
  • Agent #1 had a summary of 52. In reality it had passed only 2 of 8 validation checks, and one friendly quality rating of 90 had pulled the blended average up.

If you build a trust layer on top of that summary, you approve the wrong agents. So the job isn't "ask an AI for a number". The job is to read the individual entries, work out what each one means, and refuse to score what can't be interpreted.

What I built

Here is the flow for a question like "Can I let agent 10 do a task that needs a trust score of at least 70?":

  1. Qwen 3.8 Max plans and calls tools. It confirms the agent exists, reads and classifies all of its feedback, computes the score, checks it against your threshold, and optionally records the verdict onchain.
  2. A contract on Monad stores the verdict together with a coverage figure and an evidence hash.
  3. Anyone can verify the record, and any app or contract can gate an action on it with one call.

[

: the live reasoning steps appearing one by one on the Trust Checker page]

How Qwen is used

The analyst is a function-calling loop on qwen3.8-max, using Alibaba Cloud's OpenAI-compatible endpoint. Qwen has five tools:

Tool What it does
get_agent_identity Reads the agent's ERC-8004 identity from the Identity Registry
get_feedback_breakdown Reads every feedback entry and sorts it into win/loss outcomes, pass/fail checks, percentage ratings and off-scale values
compute_trust_score Scores the agent, applying minimum-evidence rules
check_threshold Compares the score with your requirement
record_onchain Writes the verdict to the AgentCredit contract

The loop itself is short. The model asks for tools, my code runs them and returns the results, and the model decides what to do next, until it answers without asking for another tool:

for (let step = 1; step <= maxSteps; step++) {
  const res = await client.chat.completions.create({ model, messages, tools });
  const msg = res.choices[0].message;
  messages.push(msg);

  if (!msg.tool_calls?.length) return msg.content;   // final verdict

  for (const call of msg.tool_calls) {
    const args = JSON.parse(call.function.arguments || "{}");
    const result = await runTool(call.function.name, args, ctx);
    messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify(result) });
  }
}
Enter fullscreen mode Exit fullscreen mode

A real run takes four or five model rounds and 25 to 45 seconds. The website streams every tool call to the page as it happens, so you can watch the agent work instead of staring at a spinner.

[Screenshot: the verdict card with the "Recorded onchain" box and the transaction link]

The model can plan, but it can't write the numbers

This was the most important design decision. compute_trust_score takes only an agent ID. It derives every input from the verified feedback breakdown on its own, and record_onchain recomputes the score again before writing. Qwen never types a score. A model that hallucinated a number would have no way to put it onchain.

So what does Qwen actually do? It decides which tool to call and in what order, reads the results, stops when a tool says the evidence is too thin, and writes the verdict and the explanation. For agent #1, for example, it noticed on its own that the score came from "real failure data, not an absence of data", and that a single 90% rating "carries no weight because it does not clear the minimum-evidence bar". That kind of explanation is what makes the verdict usable by a person.

Where the model alone wasn't enough

I'll be upfront about this part, because it's the most useful thing I learned.

In my first version, the score was built from the registry's summary number. When I asked about agent #16, which has exactly one rating of 85 from a single client, Qwen answered APPROVE (with caution). It hedged well in its wording, but it still approved an agent on the strength of one review, and a single wallet could create that.

Telling the model to be careful isn't a safeguard. So I moved the rules into the tools, where the model can't argue with them:

  • A signal needs at least 3 feedback entries, and the agent needs feedback from at least 2 different clients.
  • Feedback on a non-percentage scale, like Elo, is ignored.
  • If nothing usable is left, the tool returns no score and the analyst must answer INSUFFICIENT_DATA.

After that, agent #16 ends as INSUFFICIENT_DATA, and ScavBot's Elo history is ignored with an explanation instead of being scored as a perfect 100. The model's judgment still matters for how it explains and what it recommends, but the line between "enough evidence" and "not enough" is code.

Results

These are the records the analyst wrote onchain, with a required score of 70:

Agent Evidence Score Coverage Verdict
#1 Monad Demo Agent 2 of 8 validation checks passed 25 25% Reject
#10 33 wins, 9 losses from 5 clients 79 40% Approve, low confidence
#20 Veridex Oracle Agent 2 wins, 21 losses from 4 clients 9 40% Reject
#9 ScavBot Elo ratings only none 0% Insufficient data
#16 one rating from one client none n/a Insufficient data

The score is a weighted average of five signals (task success 40%, validation 25%, reputation 20%, reliability 10%, recency 5%). Only the signals that have enough evidence are used, with the weights rescaled, and a coverage figure tells you how much of the model the evidence supports. Agent #10's 79 rests on one signal, so the analyst correctly calls its confidence low and recommends a supervised trial for anything high-stakes.

[Screenshot: the score breakdown panel for agent 10]

Putting it onchain

The AgentCredit contract on Monad testnet stores each verdict with its score, coverage, the attester and an evidence hash: keccak256 of the JSON of the identity, the feedback breakdown and the scoring result.

  • Only authorized attesters can write. The owner wallet and the analyst's wallet are separate, so a leaked analyst key can add scores but can't pause the contract or change who has access.
  • Only agents that really exist in the ERC-8004 Identity Registry can be scored.
  • The owner can pause the contract and remove a wrong record, and ownership transfers in two steps.
  • 15 Foundry tests cover these rules.

Two things make the records useful to other people:

Verification. The Verify onchain record button recomputes the hash from live registry data and compares it with the one stored on the contract. A match means the stored score was derived from exactly this evidence, and the full evidence JSON is shown so anyone can hash it themselves. If new feedback has arrived since, the hashes differ and the page says so.

Trust gating. The contract exposes isTrusted(agentId, minScore, maxAge, minCoverage). Any app or contract can require a minimum score, fresh data and enough coverage in a single call:

require(
  IAgentCredit(0x0b0792a328c2253e4F23f98875ebb7DEEa859971)
    .isTrusted(agentId, 70, 0, 20),
  "agent not trusted"
);
Enter fullscreen mode Exit fullscreen mode

The Agents page has a panel that calls this function straight from the browser, so you can move the sliders and watch the gate open and close.

[: the trust-gate panel showing GATE OPEN for agent 10]

What Qwen brought to the project

Honestly, the scoring math is simple, and I could have written it with no model at all. What Qwen added is the part that was hard to hard-code:

  • Handling the unexpected. Agents have different kinds of data, and some have none. The analyst follows a different path for each: score, no score, or stop early with an explanation. I didn't write that flow as a script. I gave it tools and rules, and it chose the path.
  • Explanations people can use. Every verdict comes with evidence and reasoning in plain language, including why a number deserves low confidence.
  • Useful caution. It recommends supervised trials, notes thin samples and points out missing metadata, which a bare score wouldn't.

What it didn't bring, and why I didn't ask it to: the safety rules. Those are code.

Running it as a public demo

A model-backed endpoint is an easy way to lose money, so the server has some limits: per-visitor rate limiting, a daily cap on analyses, an hourly cap on onchain writes, a 10-minute cache (replaying a cached result costs nothing and never writes onchain twice), and a CORS allowlist. The website never sends free text to the model. The question is built on the server from an agent number and a threshold. The Qwen key and the analyst wallet key live only in the server's environment.

Limitations

  • Feedback tags aren't standardized. Treating win and pass as positive and loss and fail as negative is a heuristic, and a win in a game is only a loose stand-in for task success.
  • Coverage tops out at 40% today. Reliability and recency have no onchain source yet.
  • Fake reviewers. Requiring 2 clients and 3 entries stops one-review scores, but not someone using several wallets. Real sybil resistance needs identity or stake weighting.
  • Verification runs on my server. The evidence JSON is there so you don't have to trust it, but a fully trustless check would hash the data in the browser.
  • Testnet. The caps live in memory and reset when the server restarts, and my registry scan stopped at 100 agents.

What's next

Weighting feedback by the reviewer's own reputation, reading ERC-8004 validation responses when agents start publishing them, and a small example contract that consumes isTrusted.

Try it at https://agentcredit-six.vercel.app: open the Agents page, pick an agent, and run the Trust Checker with "Record verdict onchain" ticked. The code is at https://github.com/Abbagigo13/agentcredit.

Top comments (0)