<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: umar idris</title>
    <description>The latest articles on DEV Community by umar idris (@abbagigo).</description>
    <link>https://dev.to/abbagigo</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173283%2F626f65a7-bce9-4fb9-98aa-e3b9496f3de5.jpg</url>
      <title>DEV Community: umar idris</title>
      <link>https://dev.to/abbagigo</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/abbagigo"/>
    <language>en</language>
    <item>
      <title>I built a Qwen 3.8 Max agent that decides which ERC-8004 agents to trust, and teaches itself to say 'not enough data'</title>
      <dc:creator>umar idris</dc:creator>
      <pubDate>Fri, 09 Oct 2026 11:48:22 +0000</pubDate>
      <link>https://dev.to/abbagigo/i-built-a-qwen-38-max-agent-that-decides-which-erc-8004-agents-to-trust-and-teaches-itself-to-say-37fd</link>
      <guid>https://dev.to/abbagigo/i-built-a-qwen-38-max-agent-that-decides-which-erc-8004-agents-to-trust-and-teaches-itself-to-say-37fd</guid>
      <description>&lt;p&gt;AI agents are starting to hold wallets, take jobs and rate each other. ERC-8004 gives them an onchain identity and a public place to collect feedback. But when I looked at the real data on Monad testnet, I found a problem: the registry can tell you &lt;em&gt;what&lt;/em&gt; people said about an agent, but not whether you should trust it.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;AgentCredit&lt;/strong&gt;: a Qwen 3.8 Max agent that investigates an ERC-8004 agent, works out a trust score from the raw feedback, and can write its verdict onchain with a hash of its evidence. It was built for the Monad hackathon (Trust, Identity &amp;amp; AI Infrastructure track) and the Alibaba Cloud &lt;strong&gt;Best Builds with Qwen 3.8 Max&lt;/strong&gt; bounty.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Live demo: &lt;a href="https://agentcredit-six.vercel.app" rel="noopener noreferrer"&gt;https://agentcredit-six.vercel.app&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/Abbagigo13/agentcredit" rel="noopener noreferrer"&gt;https://github.com/Abbagigo13/agentcredit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Contract on Monad testnet: &lt;code&gt;0x0b0792a328c2253e4F23f98875ebb7DEEa859971&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;[&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F31gm5yo2ukvw93skvsmv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F31gm5yo2ukvw93skvsmv.png" alt=" " width="800" height="414"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;: the Trust Checker result for agent 10, with the reasoning steps and the score breakdown]&lt;/p&gt;
&lt;h2&gt;
  
  
  The problem: one number that means different things
&lt;/h2&gt;

&lt;p&gt;ERC-8004's Reputation Registry lets anyone leave feedback for an agent: a number, plus tags saying what it is. There's no standard for what the number means. I scanned the first 100 registered agents on Monad testnet and looked at what the feedback really contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent #9 (ScavBot)&lt;/strong&gt; has a registry summary of &lt;strong&gt;1439&lt;/strong&gt;. Its feedback is Elo ratings, which aren't a 0-100 score at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents #10 and #20&lt;/strong&gt; both show a summary of &lt;strong&gt;0&lt;/strong&gt;. But #10's feedback is 33 wins and 9 losses, and #20's is 2 wins and 21 losses. One is a decent performer and the other is a disaster, and the summary number hides it. The registry appears to average values that are +1 for a win and -1 for a loss, then round, so both land on 0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent #1&lt;/strong&gt; had a summary of &lt;strong&gt;52&lt;/strong&gt;. In reality it had passed only 2 of 8 validation checks, and one friendly quality rating of 90 had pulled the blended average up.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you build a trust layer on top of that summary, you approve the wrong agents. So the job isn't "ask an AI for a number". The job is to read the individual entries, work out what each one means, and refuse to score what can't be interpreted.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;Here is the flow for a question like "Can I let agent 10 do a task that needs a trust score of at least 70?":&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Qwen 3.8 Max plans and calls tools.&lt;/strong&gt; It confirms the agent exists, reads and classifies all of its feedback, computes the score, checks it against your threshold, and optionally records the verdict onchain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A contract on Monad stores the verdict&lt;/strong&gt; together with a coverage figure and an evidence hash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anyone can verify the record&lt;/strong&gt;, and any app or contract can gate an action on it with one call.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;[&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdw6x39vw2veb9rft2r3k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdw6x39vw2veb9rft2r3k.png" alt=" " width="800" height="411"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;: the live reasoning steps appearing one by one on the Trust Checker page]&lt;/p&gt;
&lt;h2&gt;
  
  
  How Qwen is used
&lt;/h2&gt;

&lt;p&gt;The analyst is a function-calling loop on &lt;code&gt;qwen3.8-max&lt;/code&gt;, using Alibaba Cloud's OpenAI-compatible endpoint. Qwen has five tools:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_agent_identity&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reads the agent's ERC-8004 identity from the Identity Registry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_feedback_breakdown&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reads every feedback entry and sorts it into win/loss outcomes, pass/fail checks, percentage ratings and off-scale values&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;compute_trust_score&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Scores the agent, applying minimum-evidence rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;check_threshold&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Compares the score with your requirement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;record_onchain&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Writes the verdict to the AgentCredit contract&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The loop itself is short. The model asks for tools, my code runs them and returns the results, and the model decides what to do next, until it answers without asking for another tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;maxSteps&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tools&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// final verdict&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;call&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;{}&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;runTool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;tool_call_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A real run takes four or five model rounds and 25 to 45 seconds. The website streams every tool call to the page as it happens, so you can watch the agent work instead of staring at a spinner.&lt;/p&gt;

&lt;p&gt;[Screenshot: the verdict card with the "Recorded onchain" box and the transaction link]&lt;/p&gt;

&lt;h3&gt;
  
  
  The model can plan, but it can't write the numbers
&lt;/h3&gt;

&lt;p&gt;This was the most important design decision. &lt;code&gt;compute_trust_score&lt;/code&gt; takes &lt;strong&gt;only an agent ID&lt;/strong&gt;. It derives every input from the verified feedback breakdown on its own, and &lt;code&gt;record_onchain&lt;/code&gt; recomputes the score again before writing. Qwen never types a score. A model that hallucinated a number would have no way to put it onchain.&lt;/p&gt;

&lt;p&gt;So what does Qwen actually do? It decides which tool to call and in what order, reads the results, stops when a tool says the evidence is too thin, and writes the verdict and the explanation. For agent #1, for example, it noticed on its own that the score came from "real failure data, not an absence of data", and that a single 90% rating "carries no weight because it does not clear the minimum-evidence bar". That kind of explanation is what makes the verdict usable by a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the model alone wasn't enough
&lt;/h2&gt;

&lt;p&gt;I'll be upfront about this part, because it's the most useful thing I learned.&lt;/p&gt;

&lt;p&gt;In my first version, the score was built from the registry's summary number. When I asked about &lt;strong&gt;agent #16&lt;/strong&gt;, which has exactly one rating of 85 from a single client, Qwen answered &lt;em&gt;APPROVE (with caution)&lt;/em&gt;. It hedged well in its wording, but it still approved an agent on the strength of one review, and a single wallet could create that.&lt;/p&gt;

&lt;p&gt;Telling the model to be careful isn't a safeguard. So I moved the rules into the tools, where the model can't argue with them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A signal needs at least &lt;strong&gt;3 feedback entries&lt;/strong&gt;, and the agent needs feedback from at least &lt;strong&gt;2 different clients&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Feedback on a non-percentage scale, like Elo, is &lt;strong&gt;ignored&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;If nothing usable is left, the tool returns no score and the analyst must answer &lt;code&gt;INSUFFICIENT_DATA&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After that, agent #16 ends as &lt;code&gt;INSUFFICIENT_DATA&lt;/code&gt;, and ScavBot's Elo history is ignored with an explanation instead of being scored as a perfect 100. The model's judgment still matters for how it explains and what it recommends, but the line between "enough evidence" and "not enough" is code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;These are the records the analyst wrote onchain, with a required score of 70:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;#1 Monad Demo Agent&lt;/td&gt;
&lt;td&gt;2 of 8 validation checks passed&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;Reject&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;#10&lt;/td&gt;
&lt;td&gt;33 wins, 9 losses from 5 clients&lt;/td&gt;
&lt;td&gt;79&lt;/td&gt;
&lt;td&gt;40%&lt;/td&gt;
&lt;td&gt;Approve, low confidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;#20 Veridex Oracle Agent&lt;/td&gt;
&lt;td&gt;2 wins, 21 losses from 4 clients&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;40%&lt;/td&gt;
&lt;td&gt;Reject&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;#9 ScavBot&lt;/td&gt;
&lt;td&gt;Elo ratings only&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;Insufficient data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;#16&lt;/td&gt;
&lt;td&gt;one rating from one client&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;Insufficient data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The score is a weighted average of five signals (task success 40%, validation 25%, reputation 20%, reliability 10%, recency 5%). Only the signals that have enough evidence are used, with the weights rescaled, and a &lt;strong&gt;coverage&lt;/strong&gt; figure tells you how much of the model the evidence supports. Agent #10's 79 rests on one signal, so the analyst correctly calls its confidence low and recommends a supervised trial for anything high-stakes.&lt;/p&gt;

&lt;p&gt;[Screenshot: the score breakdown panel for agent 10]&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting it onchain
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;AgentCredit&lt;/code&gt; contract on Monad testnet stores each verdict with its score, coverage, the attester and an &lt;strong&gt;evidence hash&lt;/strong&gt;: &lt;code&gt;keccak256&lt;/code&gt; of the JSON of the identity, the feedback breakdown and the scoring result.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Only authorized &lt;strong&gt;attesters&lt;/strong&gt; can write. The owner wallet and the analyst's wallet are separate, so a leaked analyst key can add scores but can't pause the contract or change who has access.&lt;/li&gt;
&lt;li&gt;Only agents that really exist in the ERC-8004 Identity Registry can be scored.&lt;/li&gt;
&lt;li&gt;The owner can pause the contract and remove a wrong record, and ownership transfers in two steps.&lt;/li&gt;
&lt;li&gt;15 Foundry tests cover these rules.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two things make the records useful to other people:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification.&lt;/strong&gt; The &lt;strong&gt;Verify onchain record&lt;/strong&gt; button recomputes the hash from live registry data and compares it with the one stored on the contract. A match means the stored score was derived from exactly this evidence, and the full evidence JSON is shown so anyone can hash it themselves. If new feedback has arrived since, the hashes differ and the page says so.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trust gating.&lt;/strong&gt; The contract exposes &lt;code&gt;isTrusted(agentId, minScore, maxAge, minCoverage)&lt;/code&gt;. Any app or contract can require a minimum score, fresh data and enough coverage in a single call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;require(
  IAgentCredit(0x0b0792a328c2253e4F23f98875ebb7DEEa859971)
    .isTrusted(agentId, 70, 0, 20),
  "agent not trusted"
);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Agents page has a panel that calls this function straight from the browser, so you can move the sliders and watch the gate open and close.&lt;/p&gt;

&lt;p&gt;[: the trust-gate panel showing GATE OPEN for agent 10]&lt;/p&gt;

&lt;h2&gt;
  
  
  What Qwen brought to the project
&lt;/h2&gt;

&lt;p&gt;Honestly, the scoring math is simple, and I could have written it with no model at all. What Qwen added is the part that was hard to hard-code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Handling the unexpected.&lt;/strong&gt; Agents have different kinds of data, and some have none. The analyst follows a different path for each: score, no score, or stop early with an explanation. I didn't write that flow as a script. I gave it tools and rules, and it chose the path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explanations people can use.&lt;/strong&gt; Every verdict comes with evidence and reasoning in plain language, including &lt;em&gt;why&lt;/em&gt; a number deserves low confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Useful caution.&lt;/strong&gt; It recommends supervised trials, notes thin samples and points out missing metadata, which a bare score wouldn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it didn't bring, and why I didn't ask it to: the safety rules. Those are code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it as a public demo
&lt;/h2&gt;

&lt;p&gt;A model-backed endpoint is an easy way to lose money, so the server has some limits: per-visitor rate limiting, a daily cap on analyses, an hourly cap on onchain writes, a 10-minute cache (replaying a cached result costs nothing and never writes onchain twice), and a CORS allowlist. The website never sends free text to the model. The question is built on the server from an agent number and a threshold. The Qwen key and the analyst wallet key live only in the server's environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Feedback tags aren't standardized.&lt;/strong&gt; Treating &lt;code&gt;win&lt;/code&gt; and &lt;code&gt;pass&lt;/code&gt; as positive and &lt;code&gt;loss&lt;/code&gt; and &lt;code&gt;fail&lt;/code&gt; as negative is a heuristic, and a win in a game is only a loose stand-in for task success.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coverage tops out at 40% today.&lt;/strong&gt; Reliability and recency have no onchain source yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fake reviewers.&lt;/strong&gt; Requiring 2 clients and 3 entries stops one-review scores, but not someone using several wallets. Real sybil resistance needs identity or stake weighting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification runs on my server.&lt;/strong&gt; The evidence JSON is there so you don't have to trust it, but a fully trustless check would hash the data in the browser.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Testnet.&lt;/strong&gt; The caps live in memory and reset when the server restarts, and my registry scan stopped at 100 agents.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Weighting feedback by the reviewer's own reputation, reading ERC-8004 validation responses when agents start publishing them, and a small example contract that consumes &lt;code&gt;isTrusted&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Try it at &lt;a href="https://agentcredit-six.vercel.app:" rel="noopener noreferrer"&gt;https://agentcredit-six.vercel.app:&lt;/a&gt; open the Agents page, pick an agent, and run the Trust Checker with "Record verdict onchain" ticked. The code is at &lt;a href="https://github.com/Abbagigo13/agentcredit" rel="noopener noreferrer"&gt;https://github.com/Abbagigo13/agentcredit&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>crypto</category>
      <category>marketing</category>
    </item>
  </channel>
</rss>
