DEV Community

Cover image for Glyph: I made two AI models fight inside real terminals and the only honest judge is a PTY
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

Glyph: I made two AI models fight inside real terminals and the only honest judge is a PTY

Model benchmarks are numbers on a leaderboard. Nobody feels them. I wanted the opposite: send one prompt to two models, watch both answers animate live in terminals next to each other, and keep a replayable artifact as proof. That is Glyph — a "model duel" app where the scoreboard is two CRT panes rendering real ANSI bytes captured from real pseudoterminals.

No canvas animation. No simulation. If the bytes didn't come out of a PTY, they don't go on screen.

The idea in one diagram

  one prompt ─┬─▶ Model A ─▶ script A ─▶ PTY ─▶ cast A ─▶ CRT A (cyan)
              └─▶ Model B ─▶ script B ─▶ PTY ─▶ cast B ─▶ CRT B (magenta)
Enter fullscreen mode Exit fullscreen mode

The prompt is fixed and deliciously constrained:

"Write a single Node.js script that animates the terminal for 5 seconds using ONLY raw ANSI escape codes. No libraries, no dependencies, no external processes. Output only code."

Raw ANSI only — so model quality shows up visually: escape-code fluency, frame pacing, flicker control. A weak model prints garbage; a strong one draws a spinning cube. You see the capability gap instead of reading about it.

Architecture

Vite + React + TypeScript on :5173 (xterm.js for rendering), Node + Express + TypeScript on :3001, node-pty for execution, plain fetch for model calls. The browser never calls a model — everything goes through the backend, which is what makes local providers (Ollama, LM Studio) work with zero CORS setup and keeps API keys out of the client.

BROWSER :5173 ──/api proxy──▶ BACKEND :3001 ──fetch──▶ Particle.ai / Ollama /
                                                      LM Studio / OpenRouter
Enter fullscreen mode Exit fullscreen mode

Composing: POST /api/glyph

Both slots fire concurrently with the identical system prompt
("You are a precise assistant. Answer the user's request directly.") and user prompt — plus a fresh random nonce appended to each, so response caches can't fake determinism across repeated runs:

const promptWithNonce = `${userPrompt}\n\n<!-- nonce:${nonce()} -->`;
const [a, b] = await Promise.all([callModel(slotA, ...), callModel(slotB, ...)]);
Enter fullscreen mode Exit fullscreen mode

Code extraction prefers a

```js

fence and falls back to a reply containing process.stdout.write or \x1b[ — recording which path was used:

export function extractCode(reply: string) {
  const fence = reply.match(/```
{% endraw %}
(?:js|javascript|node)?\s*\n([\s\S]*?)
{% raw %}
```/i);
  if (fence) return { code: fence[1].trim(), path: "fence" };
  if (reply.includes("process.stdout.write") || reply.includes("\u001b["))
    return { code: reply.trim(), path: "heuristic" };
  return null;
}
Enter fullscreen mode Exit fullscreen mode

Three capability rules I wish every multi-provider app followed:

  1. Reasoning tokens are detected, never assumed. Only usage.completion_tokens_details.reasoning_tokens counts. Ollama and LM Studio don't send it → UI renders n/a and hides the thinking toggle. Never print 0.
  2. Provider-specific fields stay provider-specific. chat_template_kwargs: { enable_thinking: false } goes out only for Particle.ai deepseek-* models — other providers ignore or reject unknown fields.
  3. The CoT text is radioactive. reasoning_content may arrive in the message; it is stripped and never logged, displayed, or persisted. Only its token count may be shown. And since a hidden chain-of-thought can eat the whole budget, an HTTP 200 with empty content triggers one retry with a doubled budget (cap 4000) — never misread as a refusal.

Executing: POST /api/run — the honest core

This is the part that makes Glyph provable. Before anything runs, an AST scan rejects child_process, sockets (net, fetch, WebSocket…), and any require/import — flagged VIOLATION, never executed. Then the script runs under node --no-warnings inside a real 100×30 PTY, and every byte is timestamped on arrival:

const proc = pty.spawn("node", ["--no-warnings", file], {
  name: "xterm-256color", cols: 100, rows: 30, cwd: dir, ...
});
proc.onData((data: string) => {
  cast.push({ t: (Date.now() - t0) / 1000, data });  // real arrival time
  bytes += Buffer.byteLength(data);
});
setTimeout(() => finish("TIMEOUT", 124), 6000);     // 6s hard kill
Enter fullscreen mode Exit fullscreen mode

The capture is also sealed into a standard asciinema v2 .cast file, downloadable per corner and convertible with agg casts/glyph-A-<ts>.cast out.gif.

Replaying: the browser is a VCR, not a renderer

The frontend never generates a single escape code. It replays the cast with the recorded timing — term.write(chunk) per frame — so both CRTs animate exactly as they did on the server. Play/Pause/Restart, per-frame stepping, and a scrubber let you freeze one frame for a screenshot; record mode hides all chrome except the two terminals for the thumbnail.

Proving it's real

npm run verify asserts the whole chain without any API key:

  • byte counts in the thousands (44k), not tens,
  • two distinct .cast files,
  • an infinite loop is killed at the 6s cap and reported TIMEOUT,
  • the AST scan blocks child_process and socket usage.

What I'd build next

Tournament brackets with Elo, blind audience voting, a token-level diff of the two scripts, and in-backend agg so the app serves the GIF directly.

The invariant stays: if it didn't come out of a real PTY, it doesn't go on screen.

How it looks

Code & more: https://www.dailybuild.xyz/project/263-glyph

Top comments (0)