DEV Community

Peter's Lab
Peter's Lab

Posted on

Keeping Character Names Consistent Across 200 Chapters: A Glossary Layer for LLM Translation

LLMs are surprisingly good at translating a single page of dialogue. Tone, slang, implied subjects, they handle a lot of it well.

The problem shows up later. Translate chapter 1 and chapter 40 separately, and the same character can come out with two different spellings. A sect name gets transliterated in one chapter and translated literally in the next. A nickname quietly turns into the character's full name halfway through the series.

Each page looks fine on its own. Read them in sequence and the story falls apart.

This post is about a simple fix that works well in practice: a small glossary layer that sits between your text and the model.

Why this happens

Most translation pipelines process one page or one chapter at a time. The model has no memory of what it decided before, so every page is a fresh guess.

Names are especially fragile because:

Japanese names have multiple readings. The same kanji can be read several ways, and the model has to guess.
Romanization varies. Shō, Sho, Shou and Shoh are all "correct" for the same name.
Chinese and Korean names often look like ordinary words. A name made of characters meaning "forest" and "wind" can get translated literally.
Order varies. Family name first or last? The model may switch depending on context.

Giving the model the whole series as context isn't realistic. What it actually needs is the decisions, not the text.

The glossary

A glossary is just a list of source terms and how they should be rendered. Keep it boring and explicit:

python
from dataclasses import dataclass, field

@dataclass
class Term:
source: str # as it appears in the original text
target: str # how it should appear in the translation
kind: str = "name" # name, place, technique, title, ...
note: str = "" # optional hint for the model
aliases: list[str] = field(default_factory=list)

glossary = [
Term("田中", "Tanaka", note="family name, keep family-name-first"),
Term("翔", "Shou", note="given name, use 'Shou' not 'Sho' or 'Shō'"),
Term("林风", "Lin Feng", note="person's name, do not translate the meaning"),
Term("九天雷掌", "Nine Heavens Thunder Palm", kind="technique"),
]

The note field matters more than it looks. "Do not translate the meaning" or "keep the honorific" gives the model the why, which helps it handle variations it hasn't seen before.

Only send what's relevant

You don't want to stuff a 500-entry glossary into every prompt. It costs tokens, and long irrelevant lists can actually make the model less reliable.

Instead, scan the source text and include only the terms that appear in this chunk:

python
def relevant_terms(text: str, glossary: list[Term]) -> list[Term]:
hits = []
for term in glossary:
if term.source in text or any(a in text for a in term.aliases):
hits.append(term)
return hits

For Japanese and Chinese, plain substring matching works surprisingly well because there are no spaces to worry about. For languages with inflection or particles attached to names, you may need a light normalization step first.

One catch: short terms can produce false matches inside longer words. Sorting terms by length and matching longest first reduces that.

Put it in the prompt as a hard constraint

Format the matched terms as a clear table and tell the model they are non-negotiable:

python
def build_prompt(text: str, terms: list[Term], target_lang: str) -> str:
lines = [f"- {t.source} → {t.target}" + (f" ({t.note})" if t.note else "")
for t in terms]
glossary_block = "\n".join(lines) if lines else "(none)"

return f"""Translate the following manga dialogue into {target_lang}.
Enter fullscreen mode Exit fullscreen mode

Use these renderings exactly. Do not change spelling, order or form:
{glossary_block}

Dialogue:
{text}
"""

Short and explicit beats clever. Models follow a small, clearly formatted list far more reliably than a paragraph of instructions.

Check the output anyway

Even with the glossary in the prompt, models occasionally drift. A cheap post-check catches most of it:

python
def missing_terms(source: str, translation: str, terms: list[Term]) -> list[Term]:
"""Terms that appear in the source but whose target is missing from the output."""
return [t for t in terms
if t.source in source and t.target.lower() not in translation.lower()]

If something is missing, you have options: retry with a stronger instruction, flag the line for review, or log it so you can see which terms cause trouble.

This check isn't perfect. A translation might legitimately use a pronoun instead of repeating a name. But as a signal, it's very useful, and it costs almost nothing.

Growing the glossary over time

The interesting part is how the glossary gets built. A few approaches that work:

Seed it manually for the main cast. Ten to twenty entries cover most of the dialogue in a typical series.
Ask the model to extract candidates. After translating a chapter, ask it to list proper nouns it encountered and how it rendered them. New entries go to a review queue rather than straight into the glossary.
Promote on repetition. A name that appears in three or more chapters is worth locking in.
Let humans override. If a reader or translator corrects a name, that correction should win everywhere, retroactively if possible.

Scanlation teams have done this by hand for years: a shared document of every name, term and spelling decision for the series. The glossary layer is really just that practice, made machine-readable.

What it doesn't solve
Readings that only appear once. If a name's reading is given only in a tiny furigana on the first appearance and OCR misses it, the glossary starts with a wrong entry. Garbage in, consistent garbage out.
Context-dependent rendering. Sometimes a character is called by a nickname in casual scenes and by their full name in formal ones. That's a style decision, and a flat glossary can flatten it.
Honorifics. Whether "-san" or "-kun" stays attached to a name is a series-wide convention, best handled as a separate rule rather than per-name entries.

Wrapping up

The core idea is simple: LLMs don't need your whole series as context. They need a short, explicit list of the decisions you've already made, scoped to the text in front of them, plus a cheap check that they actually followed it.

It's a small piece of code, but in long-running translation work, consistency is often what separates something readable from something that feels machine-made. This is one of the problems I keep coming back to while working on AI Manga Translator.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The alias handling has a useful test case: relevant_terms can select an entry through an alias, but missing_terms only checks whether the canonical source appears. An alias-only sentence could therefore bypass the output check even though the glossary entry was included in the prompt.

I would carry the matched source span or alias into validation, then distinguish a missing rendering from an intentionally permitted pronoun. A glossary revision recorded with each translated chapter would also make human corrections traceable, so retroactive fixes can target affected chapters instead of silently mixing decisions from different glossary versions.