DEV Community

Cover image for Engineering Real-Time Video Pronunciation Search: Subtitle Synchronization, Syllable Stress Parsing & Phonetic Alignment
SayItVid
SayItVid

Posted on

Engineering Real-Time Video Pronunciation Search: Subtitle Synchronization, Syllable Stress Parsing & Phonetic Alignment

Traditional online dictionaries treat pronunciation as an isolated acoustic unit. You look up a word, see a static International Phonetic Alphabet (IPA) string, and press a speaker button to play a 1.5-second pre-recorded audio snippet captured in a silent sound studio.

While useful for elementary vocabulary, this model breaks down in conversational speech. In the wild, spoken English is a stress-timed, connected stream. Native speakers constantly elide unstressed vowels, flap intervocalic alveolar stops, and link consonant-vowel boundaries across phrases.

To bridge this gap, we built SayItVid — a real-time video pronunciation search engine that indexes thousands of authentic conversational video moments and provides synchronized subtitles, IPA transcriptions, visual syllable stress markers, and word origins in under 50 milliseconds.

In this article, we break down the engineering architecture behind SayItVid: from subtitle temporal alignment and phonetic mapping to sub-second search indexing and front-end video synchronization.


1. High-Level Architecture Overview

At a high level, the SayItVid ingestion and retrieval engine consists of four primary subsystems:

[ Video Corpus (Lectures, Talks, Interviews) ]
                     │
                     ▼
       [ Transcript / VTT Ingestion ]
                     │
       ┌─────────────┴─────────────┐
       ▼                           ▼
[ Temporal Word Sync ]    [ Phonetic & IPA Parser ]
(Start/End Timestamps)    (CMU Dict + Stress Indices)
       └─────────────┬─────────────┘
                     │
                     ▼
          [ Search Engine Index ]
          (SQLite FTS5 / BM25)
                     │
                     ▼
       [ SayItVid Client Runtime ]
     (Sub-50ms Player Sync & UI)
Enter fullscreen mode Exit fullscreen mode
  1. Ingestion & Timestamp Extraction: Ingesting timestamped WebVTT and subtitle tracks, parsing sentence boundaries, and establishing word-level temporal anchors.
  2. Phonetic & Syllable Stress Mapping: Aligning English lexicon tokens with international phonetic standards, parsing vowel nuclei to detect primary (ˈ) and secondary (ˌ) stress positions.
  3. Full-Text Inverted Indexing: Fast indexing with SQLite FTS5 for sub-millisecond retrieval of lexical phrases, collocations, and idioms.
  4. Interactive Video Player Synchronization: A lightweight, zero-dependency browser runtime that binds video playback directly to subtitle cues and phonetic cards.

2. Temporal Subtitle Alignment & Sentence Windowing

A common failure mode in video clip search is the "fragmentation problem": if a player jumps strictly to the start timestamp of a target word, the user hears an abrupt sound without the acoustic runway necessary to recognize the sentence rhythm.

To solve this, our pipeline calculates an acoustic context window around each token:

interface VideoSegment {
  videoId: string;
  word: string;
  startTime: number;      // Exact timestamp of the word
  duration: number;       // Duration of the clip window
  leadInSeconds: number;  // Pre-roll context (typically 0.8s - 1.5s)
  sentenceText: string;   // Full subtitle sentence
  textBefore: string;     // Contextual runway
  textAfter: string;      // Following clause
}

function calculatePlaybackWindow(wordStart: number, wordEnd: number, sentenceStart: number, sentenceEnd: number): { playAt: number; stopAt: number } {
  // Guarantee a comfortable cognitive buffer without bleed-over into unrelated dialogue
  const playAt = Math.max(sentenceStart, wordStart - 1.2);
  const stopAt = Math.min(sentenceEnd + 0.5, wordEnd + 1.8);
  return { playAt, stopAt };
}
Enter fullscreen mode Exit fullscreen mode

This guarantees that when a user searches for a difficult word like literally, the video begins just before the clause starts, allowing the brain's auditory processing to calibrate to the speaker's cadence before the target word is uttered.


3. Algorithmic Syllable Stress & IPA Parsing

Understanding pronunciation requires more than hearing the sound—it requires visualizing where acoustic energy is concentrated.

In English phonology, syllables with primary stress feature longer vowel duration, higher pitch, and greater intensity, while unstressed syllables undergo vowel reduction to schwa (/ə/).

To surface this clearly to learners, our phonetic parser converts raw dictionary entries into visual badge arrays:

interface SyllableStructure {
  syllables: string[];
  stressedIndex: number;
  ipa: string;
  displayWord: string;
}

function parseSyllableStress(word: string, rawIpa: string): SyllableStructure {
  // Split IPA into discrete syllable blocks based on stress markers and syllable boundaries
  const cleanedIpa = rawIpa.replace(/[\[\]\/]/g, '');
  const syllableParts = word.split(/[\s·•\-]+/);

  // Locate the primary stress indicator (ˈ) in the acoustic transcription
  let primaryStressIndex = 0;
  const ipaSyllables = cleanedIpa.split('.');
  ipaSyllables.forEach((syl, idx) => {
    if (syl.includes('ˈ')) {
      primaryStressIndex = idx;
    }
  });

  return {
    syllables: syllableParts,
    stressedIndex: primaryStressIndex,
    ipa: `/${cleanedIpa}/`,
    displayWord: word
  };
}
Enter fullscreen mode Exit fullscreen mode

When rendered in the client UI, the stressed syllable is highlighted with distinct contrast and accent color, instantly communicating the word's acoustic center of gravity.


4. Case Studies: Analyzing Real-World Phonetics in Action

To demonstrate why video indexing is critical compared to synthetic text-to-speech, let’s look at four live examples indexed on SayItVid:

1. Intervocalic Flapping: "Literally"

2. Consonant Transitions: "Thoroughly"

  • Dictionary IPA: /ˈθɜːr.ə.li/
  • Spoken Phenomenon: Transitioning from the voiceless dental fricative /θ/ straight into the rhotic retroflex vowel /ɜːr/ without inserting an artificial glottal break requires specific articulatory momentum.
  • 👉 Explore synchronized video clips for "thoroughly" on SayItVid

3. Syllable Weight & Reduction: "Vulnerable"

4. Colloquial Rhythm & Idioms: "Mum's the Word"


5. Sub-50ms Search Performance: SQLite FTS5 + Edge Caching

To keep the platform responsive, we optimized the database and search queries around lightweight inverted indices using SQLite’s native FTS5 (Full-Text Search 5):

  • BM25 Ranking: Video subtitle segments are ranked by relevance, speaker authority, and clip acoustic clarity.
  • Zero Client Bloat: The entire front-end application runs on vanilla modern JavaScript without heavy client frameworks, ensuring fast Time-to-Interactive (TTI) on mobile devices and lower-bandwidth cellular connections.
  • Edge Pre-fetching: When a user navigates between video examples, consecutive video segments and subtitle streams are pre-buffered asynchronously.

Conclusion & Try It Out

Language is fundamentally a multimodal human phenomenon. By indexing real speech moments and pairing video context with rigorous phonetic IPA data and syllable stress parsing, we can provide language learners with an authentic reflection of how spoken English actually operates.

Have questions about video alignment, subtitle parsing, or linguistic tech? Drop a comment below!

Top comments (0)