A word count is only useful when you know what the counter considers a word or a character. Two tools can read the same draft and show different totals without either one being broken.
Words depend on segmentation rules
A counter has to decide which parts of a string are words. Punctuation affects that decision. Hyphenated terms can also be treated differently by different editors and forms. Text in another writing system may produce different results again.
The free Firm Beacon word and character counter has a language selector. You choose the language of your text, and when the browser supports it, the page runs the native Intl.Segmenter with that locale. The counter counts the segments that the browser marks as word-like.
That gives the page a clear rule, but it does not promise to match every editor or submission form, and those tools can count differently. When a form has a binding word limit, check the final draft in that form. Its counter is the one that decides whether the submission is accepted.
Visible characters are not always code units
JavaScript strings use UTF-16 code units. A visible character can use more than one code unit. A letter with a combining accent and an emoji sequence are common examples. They may look like one character while occupying several code units in the string.
The counter therefore reports grapheme clusters for its character count. This is closer to what a person sees. It also reports a second count with whitespace removed. Those values can differ from a website that enforces a UTF-16 limit.
The tool itself limits input to 20,000 UTF-16 code units. That limit keeps the browser work bounded. It is separate from the visible-character count shown in the results.
Paragraphs need a definition too
The counter treats a blank line as a paragraph separator. A single line break inside a paragraph does not create another paragraph. Other editors may use a different document model, especially when they distinguish visual lines from paragraphs.
Use a target without sending the draft
You can enter a whole-number target from 1 to 1,000,000. The page shows how many words remain, or how far the draft is over the target.
The text stays in the browser tab. The counter does not upload or save the text, target, or counts. Loading the page still makes normal website and analytics requests. The analytics event records that the counter was used, not the text or its results.
A short checklist
Before submitting a draft:
- Check the word count in the destination form if it has a hard limit.
- Check whether punctuation and hyphenated terms affect the result.
- Treat visible characters and UTF-16 limits as different measurements.
- Use a blank line when you want the counter to start a new paragraph.
- Keep a copy of the final text before making last-minute edits.
The free word and character counter is available without an account or email for quick browser-local checks.
Top comments (2)
The server side has its own version of this. PHP's str_word_count() only knows a-z, so a note written in Cyrillic or Greek counted as zero words, and "naïve" came out as two.
I swapped it for a Unicode regex that keeps well-known and don't as one word each, and counts every Chinese or Japanese character as one, since those scripts don't put spaces between words.
Swapping str_word_count for a Unicode regex is the right fix for the ASCII problem, and keeping well-known and don't as single words matches what most writers expect. Counting each Chinese or Japanese character as one is a different measure, though. It's a character count rather than word segmentation, and a Japanese sentence usually has fewer words than characters. If you need word-like units, Intl.Segmenter with the right locale can give you those. For limits that are actually enforced, trust the destination's submission form over any local count. Is this for enforcing a limit before people submit, or for a rough count?