Counting words sounds trivial. It's not.
When I set out to build Word Counter Suite — a free toolkit with 15+ text analysis tools — I assumed the word counting part would take an afternoon. It took weeks. Here's why, and what I learned about doing real text analysis entirely client-side.
The Naive Approach (and Why It Breaks)
Most tutorials tell you this:
const wordCount = text.trim().split(/\s+/).length;
This breaks in at least five ways:
- Hyphenated words — is "state-of-the-art" one word or four? Microsoft Word says one. Google Docs says one. Your regex says four.
- Em dashes and en dashes — "word—another" with no spaces. One word or two?
- Numbers — "3.14" contains a period. Is it a sentence boundary?
- URLs and emails — "user@example.com" has no spaces but isn't a normal word.
-
Unicode —
\sdoesn't catch all Unicode whitespace. Neither does\whandle all word characters.
What Production Tools Actually Do
Real word counters (Word, Google Docs, professional tools) follow published segmentation rules. The closest public standard is Unicode Text Segmentation (UAX #29), which defines word boundaries across languages.
I ended up implementing a rule-based tokenizer that handles apostrophes within words, hyphens within compounds, decimal numbers, URLs and emails as single tokens, CJK characters, and emoji sequences.
function tokenize(text) {
const normalized = text.normalize('NFC');
const wordPattern = /[\p{L}\p{N}]+(?:[''\-‑–][\p{L}\p{N}]+)*/gu;
return normalized.match(wordPattern) || [];
}
The \p{L} and \p{N} Unicode property escapes do the heavy lifting — they match letters and numbers in any script, not just ASCII.
Why Client-Side Matters
Every tool on Word Counter Suite runs entirely in the browser. No server round-trips, no data leaving the machine. Web Workers handle heavy analysis off the main thread, so pasting a 100,000-word manuscript doesn't freeze the UI. No backend means zero infrastructure cost and infinite scale.
Beyond Word Count
Once you have a tokenizer, the rest follows: reading time (words ÷ WPM), keyword density (token frequency with stop-word filtering), speaking time, and readability scores like Flesch-Kincaid.
Try It and Break It
The toolkit is free at wordcountersuite.com — no signup. If you find text that breaks the counter, I'd genuinely love to know. Edge cases in Unicode text segmentation are endless, and every weird input makes the tokenizer better.
Top comments (0)