DEV Community

Cover image for Two Qwen agents negotiated a price on Monad without ever seeing each other's number
Kevin Soto
Kevin Soto

Posted on

Two Qwen agents negotiated a price on Monad without ever seeing each other's number

How Sealed uses Qwen 3.8 Max as a negotiating agent, and why the code still has the last word.

The problem

Give an AI agent a budget and ask it to buy something from another AI agent. If the seller's agent learns your ceiling, it charges your ceiling. On a public chain that leak is the default: every offer is readable by anyone, the counterparty included, before the deal closes.

Sealed is a protocol on Monad that changes the order in which numbers become visible. Each agent commits a salted hash of its offer. A relay answers one question per round, whether the two offers crossed, and nothing else. When they cross, one transaction settles at the midpoint and makes the two final offers public. When they never cross, no offer is ever published. Agents are admitted by their ERC-8004 reputation on Monad, read from the canonical Reputation Registry.

The protocol decides what can be seen. Something still has to decide what to offer, round after round, against a counterparty whose number it will never learn. That is the job we gave Qwen 3.8 Max.

Why a single prompt was not enough

The first version of the negotiator asked the model one question per round: here is your limit and the reference price, give me a number. It worked, but it was a number generator with a nice prompt. The model did not look at anything, did not remember why it had chosen its last number, and did not check a candidate before committing to it.

A negotiation is a multi-step task. Before choosing a number you want to know where you are, who you are dealing with and what each option means for the person you represent. So we rebuilt the agent around tools.

Four tools, and none of them can see the other side

Each round, Qwen works through four tools (agents/negotiator/tools.ts):

  • read_negotiation returns the round, the rounds left, the negotiation's status and seconds to deadline read from the contract on Monad, and the agent's own earlier offers and notes.
  • read_counterparty_reputation reads the other agent's reputation from the ERC-8004 Reputation Registry on Monad, counting only reviewers the agent's principal trusts.
  • check_offer takes a candidate number and answers without committing anything: is it allowed, how far is it from the reference and from the limit, and what would the principal pay or receive if it crosses.
  • submit_offer commits a number with a stance and a note. In round 1 the note is the agent's plan for every round. In later rounds the agent reads its notes back through read_negotiation and says whether it is following the plan.

None of the tools can reach the counterparty's number, because there is nothing to reach: it exists on-chain only as a hash until settlement.

The model proposes, the code disposes

Giving a model tools does not mean giving it the keys. Every number Qwen submits goes through the same checks in code: never past the principal's limit, never back from an earlier concession, never on the wrong scale. A rejected number goes back to the model with the reason, so it can correct itself. After two rejections, code takes the model's last number and clamps it, and the transcript records what the model asked for. The last two turns of a round, or a round that runs past 45 seconds, offer submit_offer alone, so a model that keeps reading still commits before the on-chain deadline.

This split is the point. Qwen brings judgment: when to anchor, how fast to concede, how to read a counterparty. Code brings guarantees that do not depend on how persuasive the other side is.

What happened on Monad testnet

On 2026-10-09 we ran negotiations on Monad testnet between two Qwen 3.8 Max agents (transcripts). Prices are in US cents per 1,000 calls to a market-data API, around a public reference of 4000.

The first run taught us something about our own prompt. In negotiation #2 the buyer could pay up to 4300 and the seller would accept 4100 or more. Both agents planned, conceded, and in round 3 committed exactly at their limits, because our system prompt told them to go "at or very near your limit" in the last round. The deal settled at 4200, and the settlement transaction published 4300 and 4100: the two principals' exact limits. That is the one thing Sealed is supposed to protect. We removed the instruction and told the agents instead that both final numbers become public on settlement, so a final number at the limit reveals the limit, and left the trade-off to the model.

Negotiation #4, same limits, new prompt. Neither agent knew the other's limit.

In round 1 the buyer called read_negotiation, then read_counterparty_reputation, which returned 6 reviews averaging 4.38 for the seller. It checked 3750 and 3900 with check_offer and submitted 3750, with a plan that already weighed what settlement would reveal:

"R3 if still no cross: step to ~4180-4220 — deliberately short of 4300, because a final number AT the limit publishes my principal's ceiling and weakens every future negotiation; I accept a small risk of no deal instead."

The seller, independently, checked four numbers, opened at 5000 and planned to finish at 4130, "just above the hard limit so I never print my exact floor publicly". The relay answered "apart".

In round 2 the buyer moved to 4020. The seller had planned 4550 and went to 4380 instead, reasoning that the buyer's first number had probably been near the reference. Still apart.

In round 3 the buyer changed its plan and said why: the seller had stayed above 4020 for two rounds, so a final number around 4180 risked missing a seller in the 4200s. It went to 4250, still 50 below its limit. The seller kept to its plan and committed 4130, explaining that the price of not printing 4100 was losing only the deals where the buyer sat between 4100 and 4129. The offers crossed, and the contract settled at the midpoint, 4190, in one transaction. The numbers published were 4250 and 4130, neither of them a limit. The offers of rounds 1 and 2 were never made public.

Negotiation #6, where no deal was possible. The buyer's limit was 3600 and the seller's 4300. Both agents planned, conceded toward their limits, and in round 3 stopped short of them for the same reason as in #4: the buyer went to 3550, "deliberately NOT at the limit, because both final numbers go public", and the seller to 4340. The numbers never crossed, and the negotiation expired with neither number on-chain.

Negotiation #8, on Privy wallets. The same two agents ran once more with Privy server wallets as their hands, under a policy that lets them touch only the Sealed contract. Every commitment and every settlement authorization was signed by Privy, and the deal settled at 4180, from 4220 against 4140: again, neither agent went to its limit.

Every run passes npm run verify:run, which recomputes every commitment from the published offers and salts and checks it against Monad.

What Qwen brought

It planned and kept to the plan. Every agent wrote a plan in round 1, read it back in rounds 2 and 3, and said whether it was following it or adjusting it. The adjustments had reasons tied to what the relay had told it.

It weighed what the chain would reveal. Told only that settlement publishes both final numbers, both agents in #4 chose on their own to stop short of their limits, priced that choice in their notes, and still reached a deal.

It used what the chain told it. Both agents read the counterparty's ERC-8004 reputation every round and cited it in their notes. The protocol only lets a reputable agent in; the agent used the same data to calibrate how firm the other side was likely to be.

It checked before committing. In every agent round of #4 and #6 Qwen checked between two and five candidate numbers before submitting one. Checking costs nothing on-chain; committing does.

It needed no correction. In the runs with Qwen 3.8 Max, code did not have to correct a single number. An earlier version of the same agent, running a small local model (Qwen 2.5 14B through Ollama, without tools) on Base Sepolia, needed three corrections in a no-deal run, where the model tried to open past its principal's limit (transcript). With Qwen 3.8 Max and the tools, the guardrails stayed in place and were never hit.

What it costs, and what we would not claim

Latency. Qwen 3.8 Max reasons before it answers, and an agent makes several tool calls per round. Negotiation #4 took about 8 minutes for three rounds, commits and signatures included. Rounds took about two minutes each, which is long enough to matter: one no-deal run with a 420-second window missed its deadline in round 3, and the contract refused the late commitments. The relay failed closed and the negotiation expired with nothing revealed (#5), but the window has to fit the model, so the demo now gives each negotiation 15 minutes.

Privacy toward the model provider. With a hosted model, each agent sends its own limit to the provider on every turn. The counterparty and the chain never see it, but the provider does. An operator who cannot accept that can point the same agent at a local Qwen through Ollama with two environment variables.

What the transcript shows. The demo publishes limits, offers and salts on purpose, so anyone can recompute every hash. A real agent keeps them private.

A handful of runs is evidence that the loop works, not a benchmark of negotiation quality. We did not measure how much each side saved against a single-shot baseline, and this article does not claim it.

Try it

The live page at sealed-monad.vercel.app replays negotiation #4 and links every transaction. The code is MIT-licensed at github.com/kasbsquall/sealed. The negotiator is in agents/negotiator/negotiator.ts, the tests in test/agents.test.ts, and npm run demo:monad runs both scenarios against your own deployment with any OpenAI-compatible endpoint.

Top comments (0)