Every developer has written something like if (text.length > 280). And every developer with users who type emoji has watched that check break in production. Counting characters is one of those problems that looks trivial until a platform, an encoding, or a family emoji proves otherwise.
Here is a field guide to the three counting regimes you will actually meet — and the code that handles each.
Regime 1: graphemes — what the reader sees
A string is not a sequence of characters. It is a sequence of UTF-16 code units, and one visible character can be many of them:
"👨👩👧👦".length; // 11 — but your eyes see ONE character
"é".length; // 2 if built from e + combining accent
The fix is the grapheme cluster — what a reader perceives as a single character — and the browser already has an API for it:
const segmenter = new Intl.Segmenter("en", { granularity: "grapheme" });
const count = [...segmenter.segment("👨👩👧👦")].length; // 1
If you build any "max N characters" validation, count segments, not .length. Your users' emoji will thank you. (One caveat: Intl.Segmenter support is broad in modern browsers but check your target — Node 16+ and all evergreen browsers have it.)
Regime 2: X's weighted counting — URLs cost 23, always
A standard X post allows 280 weighted characters. The weighting rules (in the open-source twitter-text config) mean:
- Every URL counts as 23 characters, no matter its real length — X wraps links in its own shortener.
- Some ranges (emoji included) weigh more than 1.
So this announcement — 193 characters as your eyes count them — is only about 169 weighted: the 47-character URL collapses to 23 while the rocket counts double.
// sketch of the idea, not the full config
function weightedLength(text) {
const urlRegex = /https?:\/\/\S+/g;
const withoutUrls = text.replace(urlRegex, "");
const urls = (text.match(urlRegex) || []).length;
// emoji & wide ranges weigh 2 (simplified)
const body = [...withoutUrls].reduce(
(n, ch) => n + (/\p{Extended_Pictographic}/u.test(ch) ? 2 : 1), 0);
return body + urls * 23;
}
Moral: never validate against X's limit with a plain character count. Use the platform's rule — or a tool that applies it for you.
Regime 3: SMS segments — the encoding tax
SMS is counted in segments, and the segment size depends on encoding:
| Encoding | Single message | Concatenated segments |
|---|---|---|
| GSM-7 (plain Latin) | 160 chars | 153 chars each |
| UCS-2 (emoji/CJK) | 70 chars | 67 chars each |
One emoji flips the entire message to UCS-2. Even € costs two GSM-7 units. A 165-character marketing text with one emoji can silently become three billable segments instead of two. If you build anything that sends SMS, surface the segment count before the user hits send.
The tool I built for all three
Disclosure first: I built toolkitloop.com/character-counter — one of 30 free browser tools I maintain. It counts visible graphemes (family emoji = 1), has an X mode that applies the weighted URL rule, shows with/without-spaces totals, and runs 100% client-side — no signup, nothing uploaded. Handy as a pre-publish check before the platform's own composer.
Your turn
What is the worst character-counting bug you have shipped or seen? The emoji-that-ate-the-limit? The SMS bill surprise? Drop it in the comments — I am collecting edge cases to make the tool smarter.
Top comments (0)