DEV Community

Arthur031221
Arthur031221

Posted on

I built a Mandarin to English caption pipeline

I wanted to speak Mandarin while drafting English text and keep the original words visible. Finchling is my first version of that workflow. It reads microphone blocks or a WAV file, shows the recognition hypothesis beside an English preview, and emits completed clauses as commit events.

Recorded Finchling caption replay

Where a clause becomes output

A speech recognizer can change its answer after more audio arrives. That makes typing every partial translation risky. I kept the Chinese source, English preview, and committed output separate. A new hypothesis can replace the preview. Text already committed stays intact.

The default path commits at recognizer endpoints or a released push to talk key. The optional Jev integration asks typed questions about completion, intent, commands, and register. It does not produce translated text. I bind its responses to the segment, revision, and committed source position. An old response cannot authorize a new clause.

I allow speculative completion only with an explicit early commit option. If later recognition revises the committed prefix, the tool reports a conflict. It preserves the output rather than attempting to edit an application whose cursor and text ownership it does not know.

Try the local path

The README provides one command to install the speech extras and download models. Python 3.10 or newer and uv are required. Speech recognition uses sherpa-onnx. Translation uses Argos and CTranslate2 on the CPU. After setup, that path works without a cloud decision service.

finchling --live --ui
finchling --push-to-talk --cursor
finchling recording.wav --realtime --ui
finchling recording.wav --trace session.jsonl
finchling --live --jev --early-commit --ui
Enter fullscreen mode Exit fullscreen mode

What I measured

Three upstream recorded clips measured a 95.79 ms median from the end of each audio file to its final English commit. Replay followed recorded timing and included translation previews. Models were loaded before replay. The repository documents the method and raw results.

Those clips contain recognition errors and missing translation details. I treat the measurement as a small timing check, not a quality score or a live microphone benchmark. I tested Jev through the official SDK with a mock transport because this host had no API key. The 18 tests also cover streaming delivery, stale results, deadlines, and immediate cursor commits.

I used AI coding assistance and reviewed the implementation. I would like feedback on Mandarin technical vocabulary, clause boundaries, and desktop permissions before expanding the workflow.

Read the code and setup instructions.

Top comments (0)