DEV Community

LYR profile picture

LYR

Building LYR. Engineering notes on AI performance, latency, inference optimization, and production AI systems.

Location Tokyo, Japan Joined Joined on  Personal website https://lyr.jp/en/
Are you still picking that setting by gut feel?

Are you still picking that setting by gut feel?

Comments
5 min read
A general-purpose model is too wasteful a vessel for subtitles

A general-purpose model is too wasteful a vessel for subtitles

Comments
18 min read
Generation speed is decided by the byte count of the weights, not the parameter count

Generation speed is decided by the byte count of the weights, not the parameter count

Comments
7 min read
Let another AI write your teacher data

Let another AI write your teacher data

Comments
12 min read
A small specialist drew level with the strongest model in service

A small specialist drew level with the strongest model in service

Comments
8 min read
Let the AI pick a translation better than the "right answer"

Let the AI pick a translation better than the "right answer"

Comments
10 min read
Swahili from 0% to 30%. But the value wasn't the score — it was that the failures changed

Swahili from 0% to 30%. But the value wasn't the score — it was that the failures changed

Comments
11 min read
I cut the bottom-rank rate from 64% to 24%, and held the top-rank loss to 12pt

I cut the bottom-rank rate from 64% to 24%, and held the top-rank loss to 12pt

Comments
9 min read
"Add context and the model gets smarter" was half a lie

"Add context and the model gets smarter" was half a lie

Comments
9 min read
The "97% improvement" I threw away

The "97% improvement" I threw away

Comments
7 min read
The AI that was smart on the bench suddenly started making mistakes in production

The AI that was smart on the bench suddenly started making mistakes in production

Comments
7 min read
Taking a small AI's accuracy from 42% to 84% — and on some tasks, level with the giants

Taking a small AI's accuracy from 42% to 84% — and on some tasks, level with the giants

Comments
7 min read
Compress a model to a quarter of its size and its smarts barely drop

Compress a model to a quarter of its size and its smarts barely drop

Comments
6 min read
When "which one is faster?" is the wrong question to begin with

When "which one is faster?" is the wrong question to begin with

Comments
5 min read
Five rebuilds and it still never beat chance — the cause was the labels, not the model

Five rebuilds and it still never beat chance — the cause was the labels, not the model

Comments
7 min read
Don't stop at measuring a single mode

Don't stop at measuring a single mode

Comments
7 min read
Don't explain the "why" until you've taken a control

Don't explain the "why" until you've taken a control

Comments
6 min read
That "accuracy" is a definition you chose

That "accuracy" is a definition you chose

Comments
7 min read
A duplicated benchmark lies to you quietly

A duplicated benchmark lies to you quietly

Comments
8 min read
If you let an AI do the scoring, start by doubting the scores

If you let an AI do the scoring, start by doubting the scores

Comments
7 min read
False fires on UI elements: 44% 22%, without adding a single outer if

False fires on UI elements: 44% 22%, without adding a single outer if

Comments
9 min read
I cut OCR's power budget by an order of magnitude — not by making it faster, but by calling it less

I cut OCR's power budget by an order of magnitude — not by making it faster, but by calling it less

Comments
8 min read
AI apps get easier once you think in CPU and RAM

AI apps get easier once you think in CPU and RAM

Comments
11 min read
I moved one stateful component and perceived latency went from 469ms to 252ms

I moved one stateful component and perceived latency went from 469ms to 252ms

Comments
9 min read
Geography is the latency floor you can't move

Geography is the latency floor you can't move

Comments
4 min read
Don't let a benchmark decide what "fast enough" means

Don't let a benchmark decide what "fast enough" means

Comments
6 min read
You are the bottleneck

You are the bottleneck

Comments
6 min read
I brought AI in-house: E2E ~500ms 120ms, cost down to 1/3–1/10

I brought AI in-house: E2E ~500ms 120ms, cost down to 1/3–1/10

Comments
10 min read
Borrowed AI disappears overnight

Borrowed AI disappears overnight

Comments
6 min read
Some AI models are disqualified before you ever measure quality

Some AI models are disqualified before you ever measure quality

Comments
6 min read
Last year's right answer becomes this year's wrong one

Last year's right answer becomes this year's wrong one

Comments
6 min read
One setting took output tokens to 1/27, latency to 1/4, and cost to 1/5

One setting took output tokens to 1/27, latency to 1/4, and cost to 1/5

Comments
6 min read
loading...