My wife reads faster than books get translated into Bulgarian. Some of the ones she wants don’t exist in Bulgarian and probably never will, because the market is small and the translation doesn’t pay for itself.
So I tried machine translation. It came out unreadable: sentences that are grammatically correct and mean nothing, and dialogue punctuated the English way. My wife got to page thirty and gave up.
I started with DeepL, which everyone calls the best. It turned out to be all talk, and later in the test it came last of ten.
The same happened with qwen3.8-max. I only included it because it is supposedly the best model for translation. It came second to last and was the only one that dropped parts of the text. All talk again.
How I tested the AI translators
So I decided to run my own study and find out which AI literary translator really does the job best. I took a Robert E. Howard story of about 12,000 words that also has a published human translation, and gave it to eight language models and to DeepL. Every one of them got the same instructions, word for word.
For the scoring I picked four models from four different makers. Each has a relative among the translators and may go a little easy on it, so with four the bias gets diluted, and I can also see whether their scores agree. I set a few rules as well. The judges don’t know who translated what, and they score all the translations at once, so they measure on the same scale. I also checked whether the models had seen the published translation during training, because then they would simply be remembering it. I found no sign of that.
Which AI literary translator came out on top
Here is how they ranked. One judge gave only an order without points, so the chart shows three.
Points are one thing. The other is what translating the same story costs with each model.
What the pipeline adds
Based on these results I added a few steps from the pipeline of wenyi, a tool for translating books. The model first reads the whole book and builds a summary and a glossary of names, and after translating it reviews its own text and fixes what it finds.
I ran the four best models again, this time through the full pipeline. In the chart below each of them has three versions: the plain translation, the same translation with a review added, and a new translation through the full pipeline. gpt-6-astra and gemini-3.1-pro gain one or two points, while gemini-3.8-flash and deepseek-v4-pro lose one, and each set of three was scored separately, so compare it only with itself.
The human translation, and the judge as a translator
I also scored the published human translation. It came in below the best machines, and the main reason is completeness. The translator cuts on purpose: of 125 departures from the original, 66 are deliberate choices, and the scoring counts every one of them as a loss, even when the cut works. It also loses points on language. It keeps using the short definite article where Bulgarian grammar needs the full one: “кимериеца”, “вашия капитан”, “пътя ти”. There is wrong verb government too, such as “пусни ни борда”, and that one appears twice. About ten typesetting leftovers remain, such as a line-break hyphen in the middle of a word (“удивле-ние”), doubled words and Latin letters inside Cyrillic words. For Bulgarian it scores 6 out of 10, against 9 for the best model.
Finally I had one of the judges, Fable 5.1, translate the story too, in the same three versions. Two outside models scored it together with all the other translations at once, so the scale was shared. It came right after gpt-6-astra, but I don’t count that as a win, because it was the only one that knew in advance what it would be judged on.
So which model translates fiction best
Overall, gpt-6-astra showed the best results in every round, but at a price. It is the most expensive model, more than 12 times the price of the cheapest one.
gemini-3.8-flash gave the best value for money: it works fast, cheap and well enough.
I also saw that using a pipeline during translation improves the results significantly, and if I set out to translate fiction myself, I would definitely do it with something like wenyi, even if only written up as a skill or a custom agent in the tool I use.
Here you can read the detailed article about the research: In Search of the Best A.I. Literary Translator – Detailed.
The post In Search of the Best A.I. Literary Translator first appeared on Valcheff Net.






Top comments (0)