DEV Community

Cover image for Teaching a Computer to Hear Japanese Pitch Accent
Raihan
Raihan

Posted on

Teaching a Computer to Hear Japanese Pitch Accent

はし can be chopsticks or bridge — the difference is pitch, and most apps ignore it. I built an open scorer that hears it plus a TTS voice you can steer mora by mora.


akusento results

Scope note, first paragraph as promised: Tokyo-dialect dictionary accents (OpenJTalk, not human-checked — neither I nor the owner is a Japanese speaker, disclosed throughout); one JSUT speaker for training; JVNV base voice for synthesis; rule-based F0 measurement, no LLM judge.

Why pitch accent is invisible

Japanese textbooks teach vocabulary and grammar. Pitch accent — the high/low melody that distinguishes 雨 (rain, HL) from 飴 (candy, LH) — is rarely taught and almost never measured. So I measured it: F0 tracking (pyworld) aligned to morae (Julius forced alignment) against dictionary labels on 500 JSUT utterances. Predicted falls materialize 58% of the time, rises 74%; a third of directional transitions contradict the dictionary (peak delay is real — the pitch peak lands a mora late). All numbers from results/acoustic_check.jsonl in the repo.

The label pipeline

OpenJTalk full-context labels → Kurihara-method mora parsing (A1/A2/A3 fields) → H/L per mora. Three independent sources must agree per phrase (NJD chain chunks, label F-fields, A2 groups) or the utterance errors loudly: 500/500 pass. Cross-checking two OpenJTalk vintages (2017 jsut-lab vs current) agrees on 95% of utterances; training uses only agreed ones.

The scorer: 0.789 vs 0.504

Twelve per-mora F0 features (phrase-median normalized for downstep, plus neighbor-mora context) into logistic regression: mora H/L accuracy 0.789 on 2,020 held-out morae (majority baseline 0.504; v1 with utterance normalization was 0.736). Word-level exact pattern match is 0.460 — harsh metric, roughly what 0.79⁴ predicts. The weights are on Hugging Face; the model cannot reproduce anyone's voice (statistics plus a linear layer).

The controllable voice

Style-Bert-VITS2 takes per-mora tone input — but its default (line_split=True) silently ignores it, which my first experiment "proved" with a zero effect. With the flag off: rendering 雨が with 飴が's tones follows the tones (contour correlation 0.82–0.92), not the text. Across 22 minimal-pair groups: auto renders agree with dictionary tones 72% directionally; forced renders follow forced tones 59% vs 54% baseline — control is real but partial (BERT semantics compete). Kana-level ASR roundtrip CER is 0.105.

What shipped, honestly

Open: scorer code + weights, eval code + numbers, Gradio demo (type a word, hear correct vs flipped; record yourself, see the pattern). Local-only: nothing — no fine-tune was needed, so there are no JSUT-derived TTS weights to gate (the base JVNV voice is CC BY-SA). No human ever checked an accent label here; the write-up says so everywhere it matters.

Repo: github.com/raihan-js/akusento · Scorer: huggingface.co/raihan-js/akusento-scorer

Top comments (0)