<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Peter</title>
    <description>The latest articles on DEV Community by Peter (@hx23840).</description>
    <link>https://dev.to/hx23840</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174973%2F6ae4d6fb-82fd-4531-8d66-849e284021bc.jpg</url>
      <title>DEV Community: Peter</title>
      <link>https://dev.to/hx23840</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hx23840"/>
    <language>en</language>
    <item>
      <title>Best LLM for Translation? I Tested Six Models on 46 Manga Pages</title>
      <dc:creator>Peter</dc:creator>
      <pubDate>Sat, 10 Oct 2026 10:33:10 +0000</pubDate>
      <link>https://dev.to/hx23840/best-llm-for-translation-i-tested-six-models-on-46-manga-pages-a5h</link>
      <guid>https://dev.to/hx23840/best-llm-for-translation-i-tested-six-models-on-46-manga-pages-a5h</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer: in my test, gemini-3.8-flash and claude-opus-5-5 made no clear mistake translating 46 manga pages (278 lines) into English.&lt;/strong&gt; gpt-6-sol made one, gpt-6-luna three, claude-sonnet-5-5 five and claude-haiku-5-5 nine. A bigger model was not always the better translator: a light model beat a middle one.&lt;/p&gt;

&lt;p&gt;I build &lt;a href="https://serifu.net/" rel="noopener noreferrer"&gt;Serifu&lt;/a&gt;, a manga translator, so how models handle comics is something I look at closely. I ran this on 10 October 2026. Nobody sponsored it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Clear mistakes on 46 pages&lt;/th&gt;
&lt;th&gt;Seconds a page&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.8-flash&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;about 15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;claude-opus-5-5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;10 to 20, at times over a minute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-6-sol&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;about 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-6-luna&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;about 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;claude-sonnet-5-5&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;about 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;claude-haiku-5-5&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;about 6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One run on 46 pages shows the mistakes I saw. It is not an error rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why manga is a hard translation test
&lt;/h2&gt;

&lt;p&gt;Comics are a good stress test for an LLM translator, because the text alone is not enough:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Japanese drops "he" and "she". Often only the drawing says who is being talked about.&lt;/li&gt;
&lt;li&gt;One sentence is split across two or three balloons, and each balloon needs its own part.&lt;/li&gt;
&lt;li&gt;A name on page 5 has to be the same name on page 150.&lt;/li&gt;
&lt;li&gt;Sound effects, dialect and puns have no dictionary answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A benchmark on clean paragraphs measures none of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the test was set up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The pages.&lt;/strong&gt; 46 pages, 278 lines: 35 pages of &lt;em&gt;Give My Regards to Black Jack&lt;/em&gt; vol. 1 by Shuho Sato (Japanese, free for secondary use), 9 pages of the Korean webtoon 「동백꽃」 by 이호윤 (CC BY), and 2 comic pages in Traditional Chinese. All translated into English.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two rounds.&lt;/strong&gt; First 25 ordinary pages. Then 21 pages picked because translators get them wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The rubric came first.&lt;/strong&gt; For the 21 hard pages I wrote down what a right translation must do before any model ran. That stops you from grading on which output you happen to like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same input for every model.&lt;/strong&gt; The lines' text, the image of the page, the two pages before it, and the same instructions: translate the way an official English release would read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One run each, nothing edited&lt;/strong&gt;, each model at its low or no-thinking setting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What counts as a mistake.&lt;/strong&gt; A wrong meaning, the wrong person or sex, a name or term broken, or a word left untranslated. A plainer or livelier voice is not a mistake.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the 18 mistakes were
&lt;/h2&gt;

&lt;p&gt;Sorted by kind, not by model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Kind of mistake&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Models&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Meaning turned around, or who does what reversed&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;sonnet 1, haiku 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sound effects&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;sonnet 2, haiku 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Names, terms and numbers&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;sonnet 1, haiku 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Needed the picture&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;gpt-6-sol, gpt-6-luna&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A rare or old word&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;gpt-6-luna, haiku&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A thought added or misread&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;sonnet, haiku&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One sentence across two balloons&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;gpt-6-luna&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Almost none of these is a vocabulary problem. They are context problems.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The answer was in the image
&lt;/h3&gt;

&lt;p&gt;On this page the patient is a woman, and only the drawing says so. Both GPT-6 models wrote "him". The other four read the picture.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06ejfhk0qs3utibzxuo9.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06ejfhk0qs3utibzxuo9.webp" alt="A patient who is a woman: gemini, opus, sonnet and haiku write " width="800" height="461"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every model was given the page image. Having the image and using it are two different things.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Structure: one sentence, two balloons
&lt;/h3&gt;

&lt;p&gt;Each balloon needs its own half of the sentence. gpt-6-luna wrote the whole sentence in both.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftc2enoq4pc3xhmb6z1bj.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftc2enoq4pc3xhmb6z1bj.webp" alt="One sentence across two balloons: five models split it, gpt-6-luna repeats it in both" width="800" height="477"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If your output is keyed line by line, check for this. The translation is correct and the page is still wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The same name twice on one page
&lt;/h3&gt;

&lt;p&gt;第一外科, the First Surgery department, is named twice on one page. claude-sonnet-5-5 called it "General Surgery" the first time and "First Surgery" the second. claude-haiku-5-5 dropped "First" both times.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fen0a3ri1ns6lhl0z1f37.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fen0a3ri1ns6lhl0z1f37.webp" alt="A department named twice: sonnet writes " width="800" height="461"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If a model drifts inside one page, it will drift across a 200-page volume. This is why a glossary has to go in with every request.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Numbers and drug names
&lt;/h3&gt;

&lt;p&gt;A resident reports heparin at 800 units an hour and a drug called Millisrol. claude-haiku-5-5 lost "an hour" and spelled the drug "Milislon". The other five were right.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmzzu0vjfq9p6qmbhx6i.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmzzu0vjfq9p6qmbhx6i.webp" alt="A medical line: five models right, haiku drops " width="800" height="461"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Dialect turned the meaning around
&lt;/h3&gt;

&lt;p&gt;The Korean webtoon is set in a mountain village and its characters speak Gangwon dialect. Asked "are you working alone?", the boy snaps back: of course alone, would I do it in a crowd? Two Claude models turned the retort around.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3gujnvfe9rq0790hq16r.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3gujnvfe9rq0790hq16r.webp" alt="A line in Gangwon dialect: four models right, sonnet and haiku turn it around" width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Sound effects
&lt;/h3&gt;

&lt;p&gt;ハァ is a sigh. claude-sonnet-5-5 wrote "&lt;em&gt;Pant&lt;/em&gt;" for a man who is standing still. Again the picture had the answer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg6ctnvkqeftlhsumokmf.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg6ctnvkqeftlhsumokmf.webp" alt="The sound effect ハァ: " width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Things I did not count
&lt;/h2&gt;

&lt;p&gt;Two cases have no single right answer, so nobody lost a point.&lt;/p&gt;

&lt;p&gt;A drunk man slurs "that's bad for you" into the word for liver. gpt-6-luna kept both meanings in "It's ba-ad for your li-ver". gpt-6-sol and claude-opus-5-5 wrote English puns of their own.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp8eah4djfgvfgepg1qip.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp8eah4djfgvfgepg1qip.webp" alt="A pun on " width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The webtoon's title flower, 동백꽃, means camellia in standard Korean. In the story's dialect it is the yellow spicebush. Five models wrote "yellow camellias". Only gemini-3.8-flash wrote "yellow ginger flowers".&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model size is not a ranking.&lt;/strong&gt; claude-sonnet-5-5, the middle model of its family, made five mistakes. gemini-3.8-flash, a light model, made none. Test the exact model, not the family.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send the image, and the pages before it.&lt;/strong&gt; Without the picture, nothing says the patient is a woman. Without the earlier pages, names and speakers drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the rubric before you run anything.&lt;/strong&gt; "Which reads better" is taste. "Did it call the woman him" is a count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the misses.&lt;/strong&gt; A table of scores tells you who won. The list of what each model got wrong tells you what to put in the prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Translation is half the job.&lt;/strong&gt; A model gives you text. The page still has to be cleaned and the translation lettered back into each balloon. That part is what Serifu does around the model.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;One run per model, 46 pages, one judge (me), three source languages, English only as the target. Different pages or a second run could move a model up or down by a mistake or two. Treat the order of the top three as close.&lt;/p&gt;

&lt;p&gt;The full write-up, with every mistake listed for every model and all ten example pictures, is on the Serifu blog: &lt;a href="https://serifu.net/blog/best-ai-model-for-manga-translation" rel="noopener noreferrer"&gt;Best AI Model for Manga Translation in 2026&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you have run a translation test of your own on another kind of text, I would like to hear which mistakes showed up there.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
