DEV Community

Cover image for Best LLM for Translation? I Tested Six Models on 46 Manga Pages
Peter
Peter

Posted on

Best LLM for Translation? I Tested Six Models on 46 Manga Pages

Short answer: in my test, gemini-3.8-flash and claude-opus-5-5 made no clear mistake translating 46 manga pages (278 lines) into English. gpt-6-sol made one, gpt-6-luna three, claude-sonnet-5-5 five and claude-haiku-5-5 nine. A bigger model was not always the better translator: a light model beat a middle one.

I build Serifu, a manga translator, so how models handle comics is something I look at closely. I ran this on 10 October 2026. Nobody sponsored it.

Model Clear mistakes on 46 pages Seconds a page
gemini-3.8-flash 0 about 15
claude-opus-5-5 0 10 to 20, at times over a minute
gpt-6-sol 1 about 7
gpt-6-luna 3 about 8
claude-sonnet-5-5 5 about 7
claude-haiku-5-5 9 about 6

One run on 46 pages shows the mistakes I saw. It is not an error rate.

Why manga is a hard translation test

Comics are a good stress test for an LLM translator, because the text alone is not enough:

  • Japanese drops "he" and "she". Often only the drawing says who is being talked about.
  • One sentence is split across two or three balloons, and each balloon needs its own part.
  • A name on page 5 has to be the same name on page 150.
  • Sound effects, dialect and puns have no dictionary answer.

A benchmark on clean paragraphs measures none of that.

How the test was set up

  • The pages. 46 pages, 278 lines: 35 pages of Give My Regards to Black Jack vol. 1 by Shuho Sato (Japanese, free for secondary use), 9 pages of the Korean webtoon 「동백꽃」 by 이호윤 (CC BY), and 2 comic pages in Traditional Chinese. All translated into English.
  • Two rounds. First 25 ordinary pages. Then 21 pages picked because translators get them wrong.
  • The rubric came first. For the 21 hard pages I wrote down what a right translation must do before any model ran. That stops you from grading on which output you happen to like.
  • The same input for every model. The lines' text, the image of the page, the two pages before it, and the same instructions: translate the way an official English release would read.
  • One run each, nothing edited, each model at its low or no-thinking setting.
  • What counts as a mistake. A wrong meaning, the wrong person or sex, a name or term broken, or a word left untranslated. A plainer or livelier voice is not a mistake.

What the 18 mistakes were

Sorted by kind, not by model:

Kind of mistake Count Models
Meaning turned around, or who does what reversed 4 sonnet 1, haiku 3
Sound effects 4 sonnet 2, haiku 2
Names, terms and numbers 3 sonnet 1, haiku 2
Needed the picture 2 gpt-6-sol, gpt-6-luna
A rare or old word 2 gpt-6-luna, haiku
A thought added or misread 2 sonnet, haiku
One sentence across two balloons 1 gpt-6-luna

Almost none of these is a vocabulary problem. They are context problems.

1. The answer was in the image

On this page the patient is a woman, and only the drawing says so. Both GPT-6 models wrote "him". The other four read the picture.

A patient who is a woman: gemini, opus, sonnet and haiku write

Every model was given the page image. Having the image and using it are two different things.

2. Structure: one sentence, two balloons

Each balloon needs its own half of the sentence. gpt-6-luna wrote the whole sentence in both.

One sentence across two balloons: five models split it, gpt-6-luna repeats it in both

If your output is keyed line by line, check for this. The translation is correct and the page is still wrong.

3. The same name twice on one page

第一外科, the First Surgery department, is named twice on one page. claude-sonnet-5-5 called it "General Surgery" the first time and "First Surgery" the second. claude-haiku-5-5 dropped "First" both times.

A department named twice: sonnet writes

If a model drifts inside one page, it will drift across a 200-page volume. This is why a glossary has to go in with every request.

4. Numbers and drug names

A resident reports heparin at 800 units an hour and a drug called Millisrol. claude-haiku-5-5 lost "an hour" and spelled the drug "Milislon". The other five were right.

A medical line: five models right, haiku drops

5. Dialect turned the meaning around

The Korean webtoon is set in a mountain village and its characters speak Gangwon dialect. Asked "are you working alone?", the boy snaps back: of course alone, would I do it in a crowd? Two Claude models turned the retort around.

A line in Gangwon dialect: four models right, sonnet and haiku turn it around

6. Sound effects

ハァ is a sigh. claude-sonnet-5-5 wrote "Pant" for a man who is standing still. Again the picture had the answer.

The sound effect ハァ:

Things I did not count

Two cases have no single right answer, so nobody lost a point.

A drunk man slurs "that's bad for you" into the word for liver. gpt-6-luna kept both meanings in "It's ba-ad for your li-ver". gpt-6-sol and claude-opus-5-5 wrote English puns of their own.

A pun on

The webtoon's title flower, 동백꽃, means camellia in standard Korean. In the story's dialect it is the yellow spicebush. Five models wrote "yellow camellias". Only gemini-3.8-flash wrote "yellow ginger flowers".

What I took from it

  1. Model size is not a ranking. claude-sonnet-5-5, the middle model of its family, made five mistakes. gemini-3.8-flash, a light model, made none. Test the exact model, not the family.
  2. Send the image, and the pages before it. Without the picture, nothing says the patient is a woman. Without the earlier pages, names and speakers drift.
  3. Write the rubric before you run anything. "Which reads better" is taste. "Did it call the woman him" is a count.
  4. Keep the misses. A table of scores tells you who won. The list of what each model got wrong tells you what to put in the prompt.
  5. Translation is half the job. A model gives you text. The page still has to be cleaned and the translation lettered back into each balloon. That part is what Serifu does around the model.

Limits

One run per model, 46 pages, one judge (me), three source languages, English only as the target. Different pages or a second run could move a model up or down by a mistake or two. Treat the order of the top three as close.

The full write-up, with every mistake listed for every model and all ten example pictures, is on the Serifu blog: Best AI Model for Manga Translation in 2026.

If you have run a translation test of your own on another kind of text, I would like to hear which mistakes showed up there.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to