Traditional online dictionaries treat pronunciation as an isolated acoustic unit. You look up a word, see a static International Phonetic Alphabet (IPA) string, and press a speaker button to play a 1.5-second pre-recorded audio snippet captured in a silent sound studio.
While useful for elementary vocabulary, this model breaks down in conversational speech. In the wild, spoken English is a stress-timed, connected stream. Native speakers constantly elide unstressed vowels, flap intervocalic alveolar stops, and link consonant-vowel boundaries across phrases.
To bridge this gap, we built SayItVid — a real-time video pronunciation search engine that indexes thousands of authentic conversational video moments and provides synchronized subtitles, IPA transcriptions, visual syllable stress markers, and word origins in under 50 milliseconds.
In this article, we break down the engineering architecture behind SayItVid: from subtitle temporal alignment and phonetic mapping to sub-second search indexing and front-end video synchronization.
1. High-Level Architecture Overview
At a high level, the SayItVid ingestion and retrieval engine consists of four primary subsystems:
[ Video Corpus (Lectures, Talks, Interviews) ]
│
▼
[ Transcript / VTT Ingestion ]
│
┌─────────────┴─────────────┐
▼ ▼
[ Temporal Word Sync ] [ Phonetic & IPA Parser ]
(Start/End Timestamps) (CMU Dict + Stress Indices)
└─────────────┬─────────────┘
│
▼
[ Search Engine Index ]
(SQLite FTS5 / BM25)
│
▼
[ SayItVid Client Runtime ]
(Sub-50ms Player Sync & UI)
- Ingestion & Timestamp Extraction: Ingesting timestamped WebVTT and subtitle tracks, parsing sentence boundaries, and establishing word-level temporal anchors.
-
Phonetic & Syllable Stress Mapping: Aligning English lexicon tokens with international phonetic standards, parsing vowel nuclei to detect primary (
ˈ) and secondary (ˌ) stress positions. - Full-Text Inverted Indexing: Fast indexing with SQLite FTS5 for sub-millisecond retrieval of lexical phrases, collocations, and idioms.
- Interactive Video Player Synchronization: A lightweight, zero-dependency browser runtime that binds video playback directly to subtitle cues and phonetic cards.
2. Temporal Subtitle Alignment & Sentence Windowing
A common failure mode in video clip search is the "fragmentation problem": if a player jumps strictly to the start timestamp of a target word, the user hears an abrupt sound without the acoustic runway necessary to recognize the sentence rhythm.
To solve this, our pipeline calculates an acoustic context window around each token:
interface VideoSegment {
videoId: string;
word: string;
startTime: number; // Exact timestamp of the word
duration: number; // Duration of the clip window
leadInSeconds: number; // Pre-roll context (typically 0.8s - 1.5s)
sentenceText: string; // Full subtitle sentence
textBefore: string; // Contextual runway
textAfter: string; // Following clause
}
function calculatePlaybackWindow(wordStart: number, wordEnd: number, sentenceStart: number, sentenceEnd: number): { playAt: number; stopAt: number } {
// Guarantee a comfortable cognitive buffer without bleed-over into unrelated dialogue
const playAt = Math.max(sentenceStart, wordStart - 1.2);
const stopAt = Math.min(sentenceEnd + 0.5, wordEnd + 1.8);
return { playAt, stopAt };
}
This guarantees that when a user searches for a difficult word like literally, the video begins just before the clause starts, allowing the brain's auditory processing to calibrate to the speaker's cadence before the target word is uttered.
3. Algorithmic Syllable Stress & IPA Parsing
Understanding pronunciation requires more than hearing the sound—it requires visualizing where acoustic energy is concentrated.
In English phonology, syllables with primary stress feature longer vowel duration, higher pitch, and greater intensity, while unstressed syllables undergo vowel reduction to schwa (/ə/).
To surface this clearly to learners, our phonetic parser converts raw dictionary entries into visual badge arrays:
interface SyllableStructure {
syllables: string[];
stressedIndex: number;
ipa: string;
displayWord: string;
}
function parseSyllableStress(word: string, rawIpa: string): SyllableStructure {
// Split IPA into discrete syllable blocks based on stress markers and syllable boundaries
const cleanedIpa = rawIpa.replace(/[\[\]\/]/g, '');
const syllableParts = word.split(/[\s·•\-]+/);
// Locate the primary stress indicator (ˈ) in the acoustic transcription
let primaryStressIndex = 0;
const ipaSyllables = cleanedIpa.split('.');
ipaSyllables.forEach((syl, idx) => {
if (syl.includes('ˈ')) {
primaryStressIndex = idx;
}
});
return {
syllables: syllableParts,
stressedIndex: primaryStressIndex,
ipa: `/${cleanedIpa}/`,
displayWord: word
};
}
When rendered in the client UI, the stressed syllable is highlighted with distinct contrast and accent color, instantly communicating the word's acoustic center of gravity.
4. Case Studies: Analyzing Real-World Phonetics in Action
To demonstrate why video indexing is critical compared to synthetic text-to-speech, let’s look at four live examples indexed on SayItVid:
1. Intervocalic Flapping: "Literally"
-
Dictionary IPA:
/ˈlɪt.ər.əl.i/ -
Spoken Phenomenon: In natural North American speech, the intervocalic
/t/becomes an alveolar tap[ɾ], and the medial vowel is elided, compressing four syllables into three:[ˈlɪɾrəli]. - 👉 Inspect live native video clips for "literally" on SayItVid
2. Consonant Transitions: "Thoroughly"
-
Dictionary IPA:
/ˈθɜːr.ə.li/ -
Spoken Phenomenon: Transitioning from the voiceless dental fricative
/θ/straight into the rhotic retroflex vowel/ɜːr/without inserting an artificial glottal break requires specific articulatory momentum. - 👉 Explore synchronized video clips for "thoroughly" on SayItVid
3. Syllable Weight & Reduction: "Vulnerable"
-
Dictionary IPA:
/ˈvʌl.nər.ə.bəl/ - Spoken Phenomenon: Demonstrates how primary stress on the initial syllable forces reduction across the remaining unstressed suffixes in professional lectures.
- 👉 Analyze syllable stress markers for "vulnerable" on SayItVid
4. Colloquial Rhythm & Idioms: "Mum's the Word"
- Linguistic Context: Originating from Middle English momme, meaning closed lips and silence. Idiomatic delivery involves conspiratorial pacing and comedic pause contours that are absent in audio-only dictionary snippets.
- 👉 Watch authentic usage of "mum's the word" across video scenes on SayItVid
5. Sub-50ms Search Performance: SQLite FTS5 + Edge Caching
To keep the platform responsive, we optimized the database and search queries around lightweight inverted indices using SQLite’s native FTS5 (Full-Text Search 5):
- BM25 Ranking: Video subtitle segments are ranked by relevance, speaker authority, and clip acoustic clarity.
- Zero Client Bloat: The entire front-end application runs on vanilla modern JavaScript without heavy client frameworks, ensuring fast Time-to-Interactive (TTI) on mobile devices and lower-bandwidth cellular connections.
- Edge Pre-fetching: When a user navigates between video examples, consecutive video segments and subtitle streams are pre-buffered asynchronously.
Conclusion & Try It Out
Language is fundamentally a multimodal human phenomenon. By indexing real speech moments and pairing video context with rigorous phonetic IPA data and syllable stress parsing, we can provide language learners with an authentic reflection of how spoken English actually operates.
- Try the search engine live: SayItVid.com
- Explore the pronunciation index: SayItVid Pronunciation Guides
Have questions about video alignment, subtitle parsing, or linguistic tech? Drop a comment below!
Top comments (0)