DEV Community

Cover image for Can we use a System One model in the browser?
Andrey Gubanov
Andrey Gubanov

Posted on Originally published at Medium

Can we use a System One model in the browser?

I ran Laya, an open 421M-parameter decision model, in a browser tab and measured what it costs: download, speed and accuracy.

In mid-September, TypeSafe AI released Jev, a “System One” model. It doesn’t write text. You give it some input and a few typed questions, and it gives back answers with probabilities. A few days later Convai Innovations released an open alternative, Laya: 421M parameters, Apache 2.0, free to download.

A model that small raises an obvious question for anyone who builds web apps. Can it run right in the browser, with no server, no API key, and nothing leaving the page?

I tested it.

Setup

  • Model: convaiinnovations/laya, the general English checkpoint. ModernBERT-large plus a decision head, 421,293,830 parameters.
  • Browser build: layaForWeb, a community port to ONNX Runtime Web 1.30. It has two weight-only quantized builds. The 8-bit one (q8e8) uses int8 block-128 MatMulNBits weights. The 4-bit one (q4e8) uses int4 block-32 weights. Both keep the embeddings in int8.
  • Machine: Apple Silicon Mac (14 cores, Apple GPU with Metal 3), macOS 26.5, ~105 Mbit/s connection.
  • Browsers: an embedded Chromium 152 browser inside a desktop app for loads, speed and accuracy. Google Chrome 153 for memory and the background-tab test.
  • Input: the port’s support-ticket preset (“App crashes on launch… I have a demo in one hour!”) with three questions: which team (a choice of 3), how urgent (a 4-level score), and is the customer angry (yes/no). Each question becomes its own sequence, and all of them go through one batched session.run(): 220 tokens for three questions, 75 for one.
  • Timing: session.run() only. Tokenizing and softmax add 0.2–1.4 ms (median). CPU means WebAssembly with 4 threads (the port caps it at 4). Speed numbers are medians of 20 calls after 5 warm-up calls. Every load is a fresh page, 3–4 per case.
  • MB means 10⁶ bytes.

1. The download is the price

Bar chart: first visit 35.4 s and 478 MB for the 8-bit build on CPU, 25.9 s and 327 MB for the 4-bit build on GPU; repeat visit 4.0 s and 3.4 s from cache

First visit, empty cache Download Ready after (median of 3)
8-bit build, CPU 478 MB 35.4 s (34.4–36.8)
4-bit build, GPU 327 MB 25.9 s (25.2–26.2)

Both totals include the 28 MB ONNX Runtime runtime and the 3.6 MB tokenizer. The download is 85–96% of the wait. On your connection, expect about 8 × MB ÷ Mbit/s seconds for the download, plus 1.4–3.8 s to start the session and make a warm-up call. A repeat visit reads the weights from Cache Storage and is ready in 2.4–4.6 s.

So it suits a tool people come back to, not a landing page.

2. After that, it’s quick

Bar chart of session.run() medians: 1 question 0.37 s on 8-bit CPU, 0.39 s on 4-bit CPU, 0.19 s on 4-bit GPU; 3 questions 0.92 s, 0.93 s and 0.36 s

session.run(), median of 20 1 question 3 questions
8-bit, CPU 370 ms 923 ms
4-bit, CPU 389 ms 931 ms
4-bit, GPU (WebGPU) 191 ms 358 ms

The GPU numbers moved between sessions: an earlier run on the same day gave 132 ms and 293 ms. The CPU numbers stayed within about 5%.

One surprise: the 4-bit build is not faster on the CPU. It only makes the weights 34% smaller. What it really buys you is the GPU, because the 8-bit build can’t create a WebGPU session in this runtime. The port blocks that combination, and when I bypassed the check, ONNX Runtime refused on its own:

Only 2b and 4b quantization is supported for MatMulNBits op, additional bits support is planned.
Enter fullscreen mode Exit fullscreen mode

3. 4-bit moves some answers

Dot plot of the largest probability shift per answer: the 8-bit build stays under 0.06, the 4-bit build reaches 0.274 on the guardrail question needs a human, which flips from 0.34 to 0.61

The port includes a check against the original PyTorch model: 12 answers across 6 saved cases. The 8-bit build kept 11 of 12 top answers, with a worst probability shift of 0.059. The 4-bit build kept 10 of 12, worst shift 0.274. The 4-bit results were identical on GPU and CPU, and the same in both sessions.

The miss that matters is a guardrail question, “does this need a human?”. The original said 0.337, which means no. The 4-bit build said 0.611, which means yes.

The port author’s larger check (48 questions) shows the same pattern. Both builds keep 47 of 48 top answers, but the average shift is 0.013 for 8-bit and 0.063 for 4-bit.

4-bit is fine for routing and tagging. For guardrails, or anything that acts on a probability threshold, keep a human or a server-side check in the loop.

4. Keep the tab in front

Bar chart: Chrome visible tab 0.92 s, Chrome background tab median 7.1 s with runs from 6.2 to 13.2 s, embedded browser with a hidden pane 0.92 s

This one surprised me most. When I switched to another tab in Chrome, the same 3-question call took 6.2 to 13.2 seconds (4 runs) instead of 0.92 s (median of 10). Chrome deprioritises background tabs. The embedded browser didn’t slow down with its pane hidden, so this is browser policy, not a limit of the model. I didn’t isolate which part of Chrome’s policy causes it.

If your feature needs to run while the tab is in the background, a Chrome tab is the wrong place for it.

So, can we?

Yes, for quick decisions that a person triggers, in apps they come back to.

Good fit: routing tickets, tagging messages, scoring leads, “should a human look at this?” hints, and anything with private data that shouldn’t leave the device.

Not yet: first impressions, work on every keystroke, background jobs in Chrome, and unsupervised guardrails on the 4-bit build.

It also costs memory. With the 8-bit model loaded, Chrome’s performance.measureUserAgentSpecificMemory() reported 1.54 GB for the page (JavaScript and WebAssembly memory). I didn’t measure the 4-bit build.


Notes. Measured on 25 and 30 September 2026 on one Mac, in Chromium-based browsers only. Phones, Safari, Firefox and low-end laptops are untested. The browser build is an unofficial community port, not Convai’s own runtime. I tested the general checkpoint; the port also offers the fine-tuned laya-typed-decisions one. The model card reports 0.362 accuracy for the general checkpoint on its typed-decisions benchmark, close to chance, against 0.766 for the fine-tuned one. So fine-tune before trusting the answers. That’s a separate question from whether it runs in a browser.

Model: huggingface.co/convaiinnovations/laya · Browser build: vishalmysore.github.io/layaForWeb

Top comments (0)