I ran Laya, an open 421M-parameter decision model, in a browser tab and measured what it costs: download, speed and accuracy.
In mid-September, TypeSafe AI released Jev, a “System One” model. It doesn’t write text. You give it some input and a few typed questions, and it gives back answers with probabilities. A few days later Convai Innovations released an open alternative, Laya: 421M parameters, Apache 2.0, free to download.
A model that small raises an obvious question for anyone who builds web apps. Can it run right in the browser, with no server, no API key, and nothing leaving the page?
I tested it.
Setup
-
Model:
convaiinnovations/laya, the general English checkpoint. ModernBERT-large plus a decision head, 421,293,830 parameters. -
Browser build: layaForWeb, a community port to ONNX Runtime Web 1.30. It has two weight-only quantized builds. The 8-bit one (
q8e8) uses int8 block-128MatMulNBitsweights. The 4-bit one (q4e8) uses int4 block-32 weights. Both keep the embeddings in int8. - Machine: Apple Silicon Mac (14 cores, Apple GPU with Metal 3), macOS 26.5, ~105 Mbit/s connection.
- Browsers: an embedded Chromium 152 browser inside a desktop app for loads, speed and accuracy. Google Chrome 153 for memory and the background-tab test.
-
Input: the port’s support-ticket preset (“App crashes on launch… I have a demo in one hour!”) with three questions: which team (a choice of 3), how urgent (a 4-level score), and is the customer angry (yes/no). Each question becomes its own sequence, and all of them go through one batched
session.run(): 220 tokens for three questions, 75 for one. -
Timing:
session.run()only. Tokenizing and softmax add 0.2–1.4 ms (median). CPU means WebAssembly with 4 threads (the port caps it at 4). Speed numbers are medians of 20 calls after 5 warm-up calls. Every load is a fresh page, 3–4 per case. - MB means 10⁶ bytes.
1. The download is the price
| First visit, empty cache | Download | Ready after (median of 3) |
|---|---|---|
| 8-bit build, CPU | 478 MB | 35.4 s (34.4–36.8) |
| 4-bit build, GPU | 327 MB | 25.9 s (25.2–26.2) |
Both totals include the 28 MB ONNX Runtime runtime and the 3.6 MB tokenizer. The download is 85–96% of the wait. On your connection, expect about 8 × MB ÷ Mbit/s seconds for the download, plus 1.4–3.8 s to start the session and make a warm-up call. A repeat visit reads the weights from Cache Storage and is ready in 2.4–4.6 s.
So it suits a tool people come back to, not a landing page.
2. After that, it’s quick
session.run(), median of 20 |
1 question | 3 questions |
|---|---|---|
| 8-bit, CPU | 370 ms | 923 ms |
| 4-bit, CPU | 389 ms | 931 ms |
| 4-bit, GPU (WebGPU) | 191 ms | 358 ms |
The GPU numbers moved between sessions: an earlier run on the same day gave 132 ms and 293 ms. The CPU numbers stayed within about 5%.
One surprise: the 4-bit build is not faster on the CPU. It only makes the weights 34% smaller. What it really buys you is the GPU, because the 8-bit build can’t create a WebGPU session in this runtime. The port blocks that combination, and when I bypassed the check, ONNX Runtime refused on its own:
Only 2b and 4b quantization is supported for MatMulNBits op, additional bits support is planned.
3. 4-bit moves some answers
The port includes a check against the original PyTorch model: 12 answers across 6 saved cases. The 8-bit build kept 11 of 12 top answers, with a worst probability shift of 0.059. The 4-bit build kept 10 of 12, worst shift 0.274. The 4-bit results were identical on GPU and CPU, and the same in both sessions.
The miss that matters is a guardrail question, “does this need a human?”. The original said 0.337, which means no. The 4-bit build said 0.611, which means yes.
The port author’s larger check (48 questions) shows the same pattern. Both builds keep 47 of 48 top answers, but the average shift is 0.013 for 8-bit and 0.063 for 4-bit.
4-bit is fine for routing and tagging. For guardrails, or anything that acts on a probability threshold, keep a human or a server-side check in the loop.
4. Keep the tab in front
This one surprised me most. When I switched to another tab in Chrome, the same 3-question call took 6.2 to 13.2 seconds (4 runs) instead of 0.92 s (median of 10). Chrome deprioritises background tabs. The embedded browser didn’t slow down with its pane hidden, so this is browser policy, not a limit of the model. I didn’t isolate which part of Chrome’s policy causes it.
If your feature needs to run while the tab is in the background, a Chrome tab is the wrong place for it.
So, can we?
Yes, for quick decisions that a person triggers, in apps they come back to.
Good fit: routing tickets, tagging messages, scoring leads, “should a human look at this?” hints, and anything with private data that shouldn’t leave the device.
Not yet: first impressions, work on every keystroke, background jobs in Chrome, and unsupervised guardrails on the 4-bit build.
It also costs memory. With the 8-bit model loaded, Chrome’s performance.measureUserAgentSpecificMemory() reported 1.54 GB for the page (JavaScript and WebAssembly memory). I didn’t measure the 4-bit build.
Notes. Measured on 25 and 30 September 2026 on one Mac, in Chromium-based browsers only. Phones, Safari, Firefox and low-end laptops are untested. The browser build is an unofficial community port, not Convai’s own runtime. I tested the general checkpoint; the port also offers the fine-tuned laya-typed-decisions one. The model card reports 0.362 accuracy for the general checkpoint on its typed-decisions benchmark, close to chance, against 0.766 for the fine-tuned one. So fine-tune before trusting the answers. That’s a separate question from whether it runs in a browser.
Model: huggingface.co/convaiinnovations/laya · Browser build: vishalmysore.github.io/layaForWeb




Top comments (0)