DEV Community

Cover image for Deliberation Dashboard: see and analyze your consensus results
Anton Babenko for AWS Heroes

Posted on

Deliberation Dashboard: see and analyze your consensus results

Note: I use Deliberation heavily every day, in the harnesses I picked for my own work: Claude Code, Codex and Antigravity, running in T3 Code and Claude Code on the web at the same time. It has been tested in real work much harder than the version I wrote about in June.


In June I wrote that every model on a panel has to prove it earns its seat, and I shipped a debug log and an analyze tool to do the proving. Both worked. The log on my laptop is 2.4 MB of JSONL today. Measuring was the easy part. A number I have to go and dig for is a number I do not act on.

So most of the work since the June post went into making the panel visible: a local dashboard that draws every run while it happens, an analyzer that says which model to drop and why, and a record of which config each run used. Around that, a lot of smaller fixes for the ways a multi-model run goes wrong when nobody is watching. Deliberation went from v3.5.2 to v3.28.0 in that time.

If you are new here: Deliberation is an open-source MCP server that lets the agent you already use (Claude Code, Codex, Cursor, Kiro, Antigravity, OpenCode) ask GPT, Gemini, Grok and any OpenRouter model for a second opinion, or make them argue until they agree. The June post explains the consensus loop. This one is about watching it.

Install in Claude Code is still three lines:

/plugin marketplace add antonbabenko/agent-plugins
/plugin install deliberation@antonbabenko
/deliberation:setup
Enter fullscreen mode Exit fullscreen mode

Other hosts have their own guides in docs/hosts.

You cannot fix a panel you cannot see

Turn it on with one line in ~/.config/deliberation/config.json:

"dashboard": { "enabled": true }
Enter fullscreen mode Exit fullscreen mode

Then run /deliberation:dashboard (or deliberation-mcp dashboard in a terminal). It prints a URL on 127.0.0.1 and opens it. There are five tabs: Live, Runs, Config, Stats and Analyzer. I generated most of the dashboard UI in a single shot with Impeccable.

Runs tab filtered to consensus runs

The Runs tab is the history. Every /consensus, /ask-all and /ask-* call is one row with its status, the project it came from, the providers, rounds, tokens and wall time. You can filter by project, tool, provider, status and date. The yellow UNRESOLVED rows are the panel saying it could not agree in five rounds.

Click a run and you get the part I wanted from the start:

One consensus run: lanes per model, debate trajectory, provider latency

This is a real consensus-step run on the Deliberation repo itself, with Claude as the arbiter and five reviewers. Each model gets a lane, and the bars show when it was thinking and what it voted. The trajectory below tells the story of the run in three lines. Round 1: Codex and Grok asked for changes. Round 2: only Codex still objected. Round 3: all five approved, 21 minutes 46 seconds and 330k tokens after the start.

The latency table under it is where the panel stops being abstract. Grok spent 14 minutes 04 seconds in total over three calls. Kimi K3 spent 1 minute 20 seconds and voted the same way every round. In a parallel fan-out the slowest member sets the clock, so that table is effectively a bill.

Live shows the same graph while a run is still going, so you can see which lane a slow run is waiting on.

The dashboard is read-only and it is built to stay on the machine. It listens on loopback only, accepts GET requests only, and the URL carries a token that becomes a cookie for that one browser. By default it records metadata only, without prompts or responses. If you switch dashboard.capture to content, secrets are scrubbed before anything is written and PII is redacted when a page is served, unless you turn on dashboard.showPII.

The analyzer tells me which model to drop

The June version of analyze was a command that printed a report. It is still there (/deliberation:analyze, now with a since window and only the models you have configured), but the useful version is the Analyzer tab.

Analyzer: agreement vs findings, projects, request size vs latency

The first table answers the question from the June post: which model agrees more than it adds. It reads only consensus rounds, since those are the ones with verdicts. "Adds nothing" means every issue category the model raised in a round was also raised by another voice. A model becomes a drop candidate only when all of these hold: at least 10 rounds, adds nothing in 80% or more of them, is the lone objector in 10% or fewer, rarely raises an issue the arbiter accepts, and is either slow or errors a lot.

My own numbers over the last 30 days, 388 runs: Gemini sat in 273 rounds, added nothing in 95% of them and was slower than the other responders' median every time. Grok added nothing in 87%. Codex looks worse on paper at 85%, but it has the highest lone-dissent rate on the panel (15%), and the arbiter accepted its issue in all 6 rounds where a decision was recorded. The lone objector who is usually right is exactly the model you keep, and the table says "keep".

Gemini and Grok are marked UNCONFIRMED, with the config line that would turn them off. They are still on in my config. Six rounds with arbiter decisions is too few for me to trust the verdict, and the analyzer says the same thing by not calling them drop candidates yet.

The second table groups runs by project (by git remote, so every clone and worktree of a repo lands on one row) with runs, calls, timeouts, errors, p50 and p95. The third one plots request size against latency per model. For Codex on text-only prompts the fit is about 5.5 seconds per thousand characters, and the advice column refuses to suggest a timeout until a bucket has 20 calls. For Kimi K3 with files, 1 of 9 calls timed out, and it already points at models.kimi-k3.timeout as probably too low.

To get these numbers, every call now records its request size, the attached file bytes, the timeout it was actually granted and which setting granted it, and every run records the project it came from.

Every run remembers the config it ran with

In June, runs made under different configs all landed in the same log, with nothing saying which config produced which number. Now each run keeps a snapshot of the config it ran with, and the Stats tab compares configs side by side.

Stats: configuration comparison

Two rows from my own history: config 8e33794253ec had 68 runs, 23 errors and a p95 of 25 minutes 37 seconds. Config dd82d2de6943, the one the run above used, has 20 runs, 1 error and a p95 of 7 minutes 14 seconds. I would not call that a controlled experiment, but it is the first time I could see that a config change did something.

This works because more of the config is pinned in one file. GPT's model and reasoning effort (providers.codex.model, providers.codex.reasoningEffort), the Gemini and Grok models, a default OpenRouter model, per-provider timeouts and a separate reasoning effort for consensus rounds (providers.<name>.consensusReasoningEffort) all live in config.json. The Config tab shows the effective values and where each one came from.

Config: provider health

Slow is fine, unbounded is not

The June post said a full consensus loop can take minutes and that is fine for the right job. That held until one /consensus run took more than 45 minutes, because a single Codex call inside it ran for about 38. Nothing was broken. It was just waiting.

Since then a run has limits:

  • consensus.maxWallMs is a time budget for the whole run (default 30 minutes), and a peer that times out twice in a row is dropped from the remaining rounds.
  • consensus.quorumFloor (default 2) means a run cannot converge when too few reviewers answered. Two models dropping out no longer leaves one model "agreeing" with the arbiter.
  • Real disagreement resets the error streak, so a model that keeps objecting is never cut by the circuit breaker that is meant for broken providers.
  • On rounds with dissent, the arbiter's adjudication and revision overlap instead of running one after another.
  • Later rounds get a digest of earlier ones instead of the full history, and static files go first in the prompt so OpenRouter can cache them.

The other half is about calls that look like answers and are not. A 429 is retried once, honoring Retry-After. A reply like "I will begin by reviewing the files..." counts as a failure instead of a vote. One broken provider no longer costs the whole round. And timeouts are per provider now, after I watched a DeepSeek call succeed at 250 seconds while its three siblings were killed at 180.

Models argue about today, not their training cutoff

Every delegate prompt now starts with today's UTC date and one rule: when a model does not recognize something (a new model name, a recent release, a CLI flag), it marks the claim [unverified] instead of saying it does not exist.

Smaller things that removed friction

  • Claude Code on the web works: every call fits under the host's 60-second tool-call cap, and the plugin declares a 30-minute per-server timeout where the host allows it.
  • Codex login is a device login (/deliberation:codex-login): it shows a link and a code in about a second, prefers your ChatGPT login and never uses OPENAI_API_KEY.
  • Antigravity is a native host, and Codex and Gemini CLIs start on Windows.
  • /deliberation:help shows real prompts you can paste, /deliberation:doctor checks config, CLIs and paths without changing anything, and /deliberation:reload-mcp cycles the server after an update.
  • Read-only is enforced for advisory calls: Codex runs with --sandbox read-only, Gemini under a read-only sandbox.
  • Secret scrubbing understands syntax now: private key blocks, JWTs, API key formats and assignments in JSON, shell and YAML are removed before anything is written.
  • Grok and OpenRouter cannot read your repo, so Deliberation can attach a small orientation bundle (orientation.enabled). It has a byte budget now (orientation.maxBytes, 16 KB by default), which cut that part of each call from about 26K tokens to about 4K in my runs.

What I would still not trust it with

The dashboard is for the machine it runs on. There is no hosted version and no team view.

The analyzer works on coarse issue categories and needs volume. With a few dozen consensus rounds it will mostly say UNCONFIRMED, and that is the correct answer. Treat it as advice, as the page itself says.

And agreement is still not truth. The dashboard makes it much easier to see when five models agree. It does not make them right.

Try it

Update the plugin, add "dashboard": { "enabled": true } to your config, run a couple of /ask-all or /consensus calls, then /deliberation:dashboard. The Analyzer tab gets useful after a week of real runs.

Source, issues and the star button are at github.com/antonbabenko/deliberation. I want to hear the feedback (good and bad), especially from anyone running a panel I have not tried.

Top comments (0)