<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 王俊俊</title>
    <description>The latest articles on DEV Community by 王俊俊 (@_d713d35749a9d64072798).</description>
    <link>https://dev.to/_d713d35749a9d64072798</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173527%2F3c957472-e77e-46b1-8da0-93544eb80695.jpg</url>
      <title>DEV Community: 王俊俊</title>
      <link>https://dev.to/_d713d35749a9d64072798</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_d713d35749a9d64072798"/>
    <language>en</language>
    <item>
      <title>Transcribing Video in the Browser Without Uploading the Film</title>
      <dc:creator>王俊俊</dc:creator>
      <pubDate>Fri, 09 Oct 2026 23:03:59 +0000</pubDate>
      <link>https://dev.to/_d713d35749a9d64072798/transcribing-video-in-the-browser-without-uploading-the-film-5gif</link>
      <guid>https://dev.to/_d713d35749a9d64072798/transcribing-video-in-the-browser-without-uploading-the-film-5gif</guid>
      <description>&lt;p&gt;I needed timed subtitles for language study videos that already lived on my disk. Uploading a two-hour lecture to a transcription SaaS felt wrong for privacy and for bandwidth. The shape that worked was: keep the media element local, demux audio in the browser, and send only short mono chunks to a Worker.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint
&lt;/h2&gt;

&lt;p&gt;Playback stays in the page. Local files and direct links (MP4, WebM, HLS) go straight into a &lt;code&gt;&amp;lt;video&amp;gt;&lt;/code&gt; / media element. A signed-out library can live in IndexedDB. Sign-in can sync &lt;em&gt;library metadata and subtitle tracks&lt;/em&gt;; the film itself never becomes an upload.&lt;/p&gt;

&lt;p&gt;That split forces the transcription path to be audio-only.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunked audio instead of a full upload
&lt;/h2&gt;

&lt;p&gt;When a file has no usable subtitles, the client:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Demuxes audio in the browser&lt;/li&gt;
&lt;li&gt;Skips long silences when possible&lt;/li&gt;
&lt;li&gt;Encodes short 16 kHz mono segments (roughly 20â€“30 seconds)&lt;/li&gt;
&lt;li&gt;POSTs each chunk to a Cloudflare Worker&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Worker runs speech-to-text (Workers AI / Whisper-class model), returns timed text, and the client turns that into cues. Existing SRT, VTT, and ASS imports follow the same cue model, so "I already have subtitles" and "generate them" share one player.&lt;/p&gt;

&lt;p&gt;Keeping requests small matters more than clever batching. Failed chunks retry without restarting a multi-gigabyte upload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a Worker on the same host helps
&lt;/h2&gt;

&lt;p&gt;A single host serving the Vue app and the API avoids CORS theater for cookie/session flows and keeps the mental model simple: one origin for the player shell. Hono on a Worker is enough for auth, credit accounting, and the STT proxy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The learning loop is the product
&lt;/h2&gt;

&lt;p&gt;None of the model work matters if seeking is clumsy. The UI unit is a subtitle cue: previous / next, replay, loop one line, or loop an Aâ€“B range. Bilingual cues sit on the video. Word lookup opens against &lt;em&gt;that sentence&lt;/em&gt;, not a generic dictionary homepage.&lt;/p&gt;

&lt;p&gt;If you study from files you already have, that architecture is what I shipped in &lt;a href="https://langscribe.online" rel="noopener noreferrer"&gt;LangScribe&lt;/a&gt; local playback, optional AI captions, and a player built around the line you are trying to hear.&lt;/p&gt;

</description>
      <category>browser</category>
      <category>javascript</category>
      <category>privacy</category>
    </item>
    <item>
      <title>I built a language-learning player</title>
      <dc:creator>王俊俊</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:32:27 +0000</pubDate>
      <link>https://dev.to/_d713d35749a9d64072798/i-built-a-language-learning-player-1hi3</link>
      <guid>https://dev.to/_d713d35749a9d64072798/i-built-a-language-learning-player-1hi3</guid>
      <description>&lt;p&gt;I wanted the thing I actually do when I study: open a lecture or a show, see the line in two languages, hit replay until the sentence sticks, and click a word I do not know. Desktop players can do pieces of this. Most of them also want an install, a GPU stack, and an API key before the first video plays.&lt;/p&gt;

&lt;p&gt;So I built LangScribe as a website. The interesting constraint was privacy: the video file should stay on the machine. Subtitles and a library can sync. The film should not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What runs where
&lt;/h2&gt;

&lt;p&gt;The app is one Vue 3 frontend. A Cloudflare Worker (Hono) serves the built site and the API on the same host, so the player and the account API are not two deployments.&lt;/p&gt;

&lt;p&gt;Playback never leaves the browser. Local files and direct links (MP4, WebM, HLS) are handed to the media element. A signed-out library lives in IndexedDB. Sign in with Google and the library plus subtitle tracks sync; the video file still does not get uploaded.&lt;/p&gt;

&lt;p&gt;That split is the whole product. Guest mode has to be useful. Cloud mode has to be optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  Subtitles without shipping the film
&lt;/h2&gt;

&lt;p&gt;If a video has no subtitles, the browser demuxes the audio, skips silence, and uploads short 16 kHz mono chunks, about 20–30 seconds each. Speech-to-text is Cloudflare Workers AI, Whisper large-v3-turbo. The worker returns timed text; the client turns that into cues and can export SRT. Existing SRT, VTT, and ASS files import the same way.&lt;/p&gt;

&lt;p&gt;I did not want a second "upload your course to our servers" step. Chunked audio is enough for transcription, and it keeps the request small.&lt;/p&gt;

&lt;h2&gt;
  
  
  Translation that can see the previous line
&lt;/h2&gt;

&lt;p&gt;Free translation works without an account, line by line, as you watch.&lt;/p&gt;

&lt;p&gt;AI translation is the path that cares about context. Lines go up in small batches with a few preceding subtitle pairs, to a Llama 3.3 70B model on Workers AI. A batch that fails to parse falls back to single-line calls. The player only asks for lines near the playhead, then refills when you are about to run out of translated text. Word lookup is a separate, cached call: part of speech, a short explanation in that sentence, pronunciation.&lt;/p&gt;

&lt;p&gt;Credits apply to AI subtitles, AI translation, and AI lookups. Free translation does not. The server records usage; the page does not decide that a feature is free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The player loop
&lt;/h2&gt;

&lt;p&gt;None of the model work matters if seeking feels clumsy. The UI is built around the cue list: previous line, next line, replay, loop one line, or loop an A–B range. Bilingual cues sit on the video. Click a word and the card opens on that sentence, not on a dictionary homepage.&lt;/p&gt;

&lt;p&gt;That is the architecture I would keep if I rewrote the UI: local media, optional account, audio-only transcription, translation with a little context, and a player that treats a subtitle line as the unit of study.&lt;/p&gt;

&lt;p&gt;If you learn from video, you can try it at &lt;a href="https://langscribe.online" rel="noopener noreferrer"&gt;LangScribe&lt;/a&gt;. &lt;/p&gt;

</description>
      <category>showdev</category>
      <category>vue</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
