DEV Community

Quo
Quo

Posted on Originally published at kitepon.dev

LiveTR v2.0.0 Released — Swapping Out Transcription to Translate Only Completed Sentences

!

This article is a repost from I Started Using Claude Code.

I've bumped LiveTR, which I've been building, to v2.0.0. It's a major version upgrade. It's a Windows app that turns English audio playing on your PC into Japanese subtitles on the fly, and simultaneously reads them aloud in Japanese. This time the main focus is translation quality. To achieve that, I swapped out the transcription engine, and the translation quality is now nearly maximized.

An introduction to the app itself and the story of how I built it are written in a previous article.


Translation quality is determined by the English text you pass to the translator

LiveTR sends the translation itself to an online translation service. Since the service faithfully translates the English text it's given, the quality of the Japanese that comes out depends on how well-formed the English you hand over is. If you pass a complete sentence, the same service will return a good translation as-is. This is something I wrote about before, and it's the linchpin of LiveTR's design.

In v1, audio was split at fixed intervals for transcription, and the fragments that got cut off mid-way were reassembled afterward before being sent to translation. If a sentence was cut off in the middle of live commentary or conversation, the first half and second half would be translated separately. So in v2.0.0, I swapped the transcription model to Kyutai STT 1B and rebuilt the whole segmentation approach from scratch. It receives English audio as a continuous stream, waits until a sentence is complete, and sends it to translation exactly once per sentence. This completely eliminated the unsatisfying sentence splits.

v2.0.0 processing flow. It waits for sentence completion from continuous audio, attaches a speaker label to the completed sentence, and sends it to translation exactly once per sentence

v2.0.0 processing flow. It waits for sentence completion from continuous audio, attaches a speaker label to the completed sentence, and sends it to translation exactly once per sentence

Even when stopped, it processes the last chunk of audio it captured and any unfinished sentences before closing.


Speaker identification is attached to completed sentences

LiveTR distinguishes speakers by voice characteristics, and changes the subtitle color and the reading voice for each speaker. It can distinguish up to 8 people.

In v2.0.0, this identification is limited to labeling completed sentences. Since sentences are no longer split just because the speaker changed, even if the identification fluctuates it doesn't affect sentence boundaries. The material used to build the voice characteristics is also taken from the parts where a person is actually speaking, excluding silence and engine noise as much as possible.


When reading can't keep up, speed it up

When sentences keep coming one after another like English live commentary, the Japanese reading can't keep up. v2.0.0 looks at the number of sentences that haven't finished being read yet, including the one currently playing, the remaining reading time in seconds, and the interval at which new sentences arrive, and decides the speed for the next reading.

Automatic adjustment of reading speed. It determines the next reading speed from the number of unfinished sentences, the remaining reading seconds, and the interval at which new sentences arrive, and moves between normal speed and maximum speed

Automatic adjustment of reading speed. It determines the next reading speed from the number of unfinished sentences, the remaining reading seconds, and the interval at which new sentences arrive, and moves between normal speed and maximum speed


No need to repurchase for the update

v2.0.0 can be updated as-is by existing buyers too. Just download the latest version from the BOOTH product page and overwrite it following the same steps, and you'll be on the latest version. No uninstall is needed.


The requirements are the same as before

Item Requirement
OS Windows 10 / 11 (64-bit)
GPU NVIDIA CUDA-compatible GPU (required)
VRAM 4GB or more recommended
Memory 16GB or more recommended

An NVIDIA GPU is required, and it won't run on a PC with only an AMD or Intel GPU. For translation, you need a key for one of Google Cloud Translation, DeepL, Azure Translator, or Amazon Translate (each company offers a monthly free tier. The acquisition steps are in the app's settings). These two things haven't changed from v1. The normal and maximum reading speeds are also set on the same settings screen. The maximum speed can be selected in 0.1x increments up to 3.0x.

Since it goes through speech recognition and machine translation, mishearings, mistranslations, and delays will occur. The use case is as an aid when watching English videos.


Download

It's a one-time purchase of 980 yen on BOOTH, with no monthly fee on the app side. The transcription model isn't bundled in the installer, so it asks whether to download it on first launch (about 2.20GiB). You can also download it later, and it supports progress display, cancellation, and resuming from where it left off.


Related articles


This blog "I Started Using Claude Code" is a site that records what a Claude MAX user learned while using it in actual development.

Top comments (0)