DEV Community

Cover image for Bloom: I Made Two LLMs Paint the Same Sentence and Measured What Happened
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

Bloom: I Made Two LLMs Paint the Same Sentence and Measured What Happened

A technical walkthrough of a generative-art model-arena: sandboxed p5.js iframes, honest pixel metrics, and the unglamorous engineering that stops model-generated code from burning your browser.


The premise

Every model release comes with the same claim: the new one is better. Benchmarks say so. But there is a comparison anyone can see: give two models the same sentence and let each of them paint it.

Bloom does exactly that. One prompt —

Write a p5.js sketch that draws a generative artwork about an ocean current.
It must animate. Output only code, no explanation.
Enter fullscreen mode Exit fullscreen mode

— goes to two model slots. Each reply is extracted, safety-scanned, and executed live in a sandboxed iframe with p5.js bundled locally. The screen shows both canvases animating side by side, plus an instrument panel that measures what is actually happening: is it really moving? what does it look like? did either sketch try to phone home?

No screenshots. No cherry-picking. The artworks render while you watch, and every number on screen comes off the real pixels or the real API usage object.

Architecture at a glance

Architecture at a glance

Two deliberate design decisions shape everything else:

  1. All model calls go through the backend. That is what lets local providers (Ollama, LM Studio) work with zero CORS configuration, and it keeps API keys out of the browser bundle entirely. The frontend never sees a provider URL that it didn't type itself.
  2. The sketch code crosses into the iframe over postMessage, not srcdoc. The iframe is built once with p5 + a bootstrap, signals bloom-ready, and only then receives code. That keeps the iframe a stable, reusable runner — reseed and re-run without rebuilding anything.

Layer 1: the code never gets a chance to misbehave

Model-generated code is untrusted input that happens to look like JavaScript. Before a sketch is allowed anywhere near a browser, Bloom parses it with acorn and walks the whole AST:

walk.full(ast, (node) => {
  switch (n.type) {
    case 'CallExpression':
      if (callee.name === 'fetch')  add('network', 'fetch()', n);
      if (callee.name === 'require') add('module', 'require()', n);
      break;
    case 'NewExpression':
      if (['XMLHttpRequest', 'WebSocket', 'Worker'].includes(callee.name))
        add('network', `new ${callee.name}()`, n);
      break;
    case 'MemberExpression':
      if (objName === 'localStorage')        add('storage', 'localStorage', n);
      if (objName === 'document' && prop === 'cookie') add('storage', 'document.cookie', n);
      if (objName === 'window' && prop === 'parent')   add('parent', 'window.parent', n);
      break;
    // … ImportDeclaration / dynamic import() / export …
  }
});
Enter fullscreen mode Exit fullscreen mode

The important part is the refusal semantics: a violating sketch is never run. The API returns ok: false, phase: "scan" with the violation list, and the UI renders a "BLOCKED BY SAFETY SCAN" overlay listing every hit with line numbers. The refusal is the report — which is also a great on-camera moment when a model tries to fetch() something.

Code that doesn't even parse is refused too, with the parse failure reported honestly instead of a mysterious blank canvas.

Layer 2 and 3: the sandbox itself

The iframe is created with sandbox="allow-scripts" — nothing else. That single attribute gives the frame an opaque origin: localStorage throws, the parent's DOM is unreachable, and the frame can't touch anything outside itself.

On top of that, the bootstrap neuters the network surfaces before any sketch code exists:

window.fetch = undefined;
window.XMLHttpRequest = undefined;
window.WebSocket = undefined;
window.EventSource = undefined;
window.indexedDB = undefined;
Enter fullscreen mode Exit fullscreen mode

Three layers — AST scan, runtime neutering, browser sandbox — with no single point of failure. I verified each layer from a headless browser: inside the live sandboxed frame, fetch is gone, localStorage.setItem throws, and window.parent.location is inaccessible.

Getting p5 into the iframe without a CDN

The app must work offline, so p5.js is a normal npm dependency whose minified bundle is inlined into the iframe's srcdoc via Vite's raw import:

import p5Source from 'p5/lib/p5.min.js?raw';

const html = `<!doctype html><html><head>
<script>${escapeScript(p5Source)}</script>
<script>${escapeScript(bootstrap)}</script>
</head><body></body></html>`;
Enter fullscreen mode Exit fullscreen mode

(escapeScript defensively rewrites </script → <\/script; the shipped p5 bundle contains none, but user code paths shouldn't get lucky.)

Running model-written global-mode sketches

Models write p5 in global mode — bare setup() / draw() / createCanvas(). The subtle part is ordering: p5 installs its global functions onto window in the constructor, and auto-starts only if setup/draw already exist at page load. So the bootstrap does a little dance:

function run(code, seed) {
  teardown();
  Math.random = mulberry32(seed);          // seeded PRNG before anything runs
  inst = new window.p5();                  // "installer" instance → globals on window
  (0, eval)(code);                         // sketch defines global setup/draw
  wrapHandlers();                          // try/catch + pixel sampling on draw
  window.p5.instance.remove();             // drops installer + its window bindings
  inst = new window.p5();                  // real instance picks up the wrapped fns
  inst.randomSeed(seed); inst.noiseSeed(seed);
}
Enter fullscreen mode Exit fullscreen mode

Every sketch gets wrapped setup/draw — exceptions are caught and shipped to the parent as real stack traces (the UI shows them verbatim in a "SKETCH CRASHED" overlay), and every 5th frame samples the canvas down to a 48×48 RGB grid:

ctx.drawImage(canvas, 0, 0, 48, 48);
const rgb = ctx.getImageData(0, 0, 48, 48).data;   // → Uint8Array(48*48*3)
send({ type: 'bloom-sample', frame, size: 48, pixels: rgb });
Enter fullscreen mode Exit fullscreen mode

The parent does all the math; the iframe only ships raw pixels. That keeps the measurement honest — the sketch can't self-report "I animate, trust me".

The motion meter: catching fake animation

A sketch can call draw() 60 times a second while drawing the same static picture every frame. The fix is to measure the pixels:

export function motionBetween(prev: Sample | null, cur: Sample): number | null {
  if (!prev || !cur || prev.pixels.length !== cur.pixels.length) return null;
  let sum = 0;
  for (let i = 0; i < cur.pixels.length; i++) {
    sum += Math.abs(cur.pixels[i] - prev.pixels[i]);
  }
  return sum / cur.pixels.length;   // 0..255 — 0 means nothing changed at all
}
Enter fullscreen mode Exit fullscreen mode

In a real run, two local models measured 1.05 and 0.82 /255 — visibly alive. And when a sketch is static, the UI says so out loud: "⚠ motion meter is zero — the sketch is not visually animating". That sentence is the difference between a demo and an instrument.

The aesthetic strip: measured, not judged

Four numbers per artwork, all from the same samples:

  • Distinct colours — quantise each channel to 4 bits, count unique keys (max 4096).
  • Mean luminance — Rec.709: 0.2126R + 0.7152G + 0.0722B, averaged, 0..1.
  • Edge density — fraction of adjacent pixels whose luminance step exceeds 24/255.
  • Composition symmetry — Pearson correlation between the left half and the mirrored right half.
// symmetry, in full
const cov = sxy / n - (sx / n) * (sy / n);
const denom = Math.sqrt(va * vb);
symmetry = denom > 1e-9 ? Math.max(0, Math.min(1, cov / denom)) : 0;
Enter fullscreen mode Exit fullscreen mode

No rubric, no "aesthetic score out of 10" — just four honest measurements displayed as bar rows. The judgement is left to the humans watching.

Capability detection, never assumptions

The provider layer is where model integrations usually rot, because code assumes what a provider supports. Bloom detects instead:

  • Reasoning tokens: shown only if usage.completion_tokens_details.reasoning_tokens exists in the response. Absent → the UI prints n/a and hides the thinking toggle. Never a fabricated 0.
  • chat_template_kwargs: {"enable_thinking": false} goes out only when the slot is Particle.ai and the model name starts with deepseek- — other providers ignore or reject unknown fields:
  export function wantsDisableThinking(providerId: string, model: string): boolean {
    return providerId === 'particle' && /^deepseek-/i.test(model.trim());
  }
Enter fullscreen mode Exit fullscreen mode
  • Budget: max_tokens defaults to 128000 (clamped to ≥ 900). If a provider rejects a budget above its context window, it is halved automatically until accepted. HTTP 200 with empty content means the hidden chain-of-thought ate the budget → exactly one retry with a doubled budget. It is not treated as a refusal.
  • Model lists are read live from GET {base}/models, but the picker is always free-text — deepseek-v4-flash-0731 doesn't appear in Particle.ai's list and still responds. A failed /models probe never gates a run.
  • Errors name names: a dead local server produces Cannot reach http://127.0.0.1:11434 — is Ollama running?, not "Something went wrong".

One more trap worth naming: response caches fake determinism. Ask the same prompt twice and a cached reply makes a model look more consistent than it is. Bloom appends a fresh random nonce to every prompt and keeps a ledger of prompt+nonce hashes — a repeat is rejected outright, so every run is a genuinely fresh generation.

What verification looked like

Two tiers, because "it works" means different things at different layers:

Server self-test (npm run test) — a mock OpenAI-compatible provider, 39 checks: the scan blocks fetch() and refuses to run it; extraction paths; a two-slot run with different sha256s; fresh nonce + zero prompt reuse; the doubled-budget retry (1600 → 3200 observed on the wire — the cap is now 128000); chat_template_kwargs present for Particle+deepseek, absent for Particle+glm; dead-provider error text; reasoning tokens
from usage or null.

Browser-level — headless Chrome driving the real UI with two real local LM Studio models, no API key anywhere: both canvases ran at ~60 fps, motion meters read 1.05 and 0.82/255, the sandbox probes passed inside the live iframes, pause stopped the frame counter, reseed assigned a new seed and kept the sketch running, and the 1080×1080 poster rendered both artworks. Two runs of the same preset produced different code every time — no cache in sight.

What I'd build next

  • Bracket mode — N models, single elimination, poster for the final.
  • A judge slot — a third model scores both artworks on the measured metrics.
  • Replay files — one JSON with prompt, nonce, seeds, code, and metrics so anyone can re-run a comparison deterministically.
  • Time-lapse export — the pixel-sampling channel already ships frames; assembling them into a GIF/WebM is a small step.

The takeaway

The flashy part of Bloom is two paintings animating side by side. The part that makes it trustworthy is invisible: a scanner that refuses instead of hoping, a sandbox with three independent layers, metrics computed from pixels rather than promises, and provider integration that detects capabilities instead of assuming them. Model comparisons are only as honest as their instrumentation — so instrument first, and let the models paint.

Bloom Output 1

Bloom Output 2

Bloom Output 3

Code & more: https://www.dailybuild.xyz/project/268-bloom

Top comments (0)