DEV Community

Cover image for Character Counting Is Harder Than It Looks: A Dev's Guide
bowen Tse
bowen Tse

Posted on Originally published at toolkitloop.blogspot.com

Character Counting Is Harder Than It Looks: A Dev's Guide

Every developer has written something like if (text.length > 280). And every developer with users who type emoji has watched that check break in production. Counting characters is one of those problems that looks trivial until a platform, an encoding, or a family emoji proves otherwise.

Here is a field guide to the three counting regimes you will actually meet — and the code that handles each.

Regime 1: graphemes — what the reader sees

A string is not a sequence of characters. It is a sequence of UTF-16 code units, and one visible character can be many of them:

"👨‍👩‍👧‍👦".length; // 11 — but your eyes see ONE character
"é".length;        // 2 if built from e + combining accent
Enter fullscreen mode Exit fullscreen mode

The fix is the grapheme cluster — what a reader perceives as a single character — and the browser already has an API for it:

const segmenter = new Intl.Segmenter("en", { granularity: "grapheme" });
const count = [...segmenter.segment("👨‍👩‍👧‍👦")].length; // 1
Enter fullscreen mode Exit fullscreen mode

If you build any "max N characters" validation, count segments, not .length. Your users' emoji will thank you. (One caveat: Intl.Segmenter support is broad in modern browsers but check your target — Node 16+ and all evergreen browsers have it.)

Regime 2: X's weighted counting — URLs cost 23, always

A standard X post allows 280 weighted characters. The weighting rules (in the open-source twitter-text config) mean:

  • Every URL counts as 23 characters, no matter its real length — X wraps links in its own shortener.
  • Some ranges (emoji included) weigh more than 1.

So this announcement — 193 characters as your eyes count them — is only about 169 weighted: the 47-character URL collapses to 23 while the rocket counts double.

// sketch of the idea, not the full config
function weightedLength(text) {
  const urlRegex = /https?:\/\/\S+/g;
  const withoutUrls = text.replace(urlRegex, "");
  const urls = (text.match(urlRegex) || []).length;
  // emoji & wide ranges weigh 2 (simplified)
  const body = [...withoutUrls].reduce(
    (n, ch) => n + (/\p{Extended_Pictographic}/u.test(ch) ? 2 : 1), 0);
  return body + urls * 23;
}
Enter fullscreen mode Exit fullscreen mode

Moral: never validate against X's limit with a plain character count. Use the platform's rule — or a tool that applies it for you.

Regime 3: SMS segments — the encoding tax

SMS is counted in segments, and the segment size depends on encoding:

Encoding Single message Concatenated segments
GSM-7 (plain Latin) 160 chars 153 chars each
UCS-2 (emoji/CJK) 70 chars 67 chars each

One emoji flips the entire message to UCS-2. Even € costs two GSM-7 units. A 165-character marketing text with one emoji can silently become three billable segments instead of two. If you build anything that sends SMS, surface the segment count before the user hits send.

The tool I built for all three

Disclosure first: I built toolkitloop.com/character-counter — one of 30 free browser tools I maintain. It counts visible graphemes (family emoji = 1), has an X mode that applies the weighted URL rule, shows with/without-spaces totals, and runs 100% client-side — no signup, nothing uploaded. Handy as a pre-publish check before the platform's own composer.

Your turn

What is the worst character-counting bug you have shipped or seen? The emoji-that-ate-the-limit? The SMS bill surprise? Drop it in the comments — I am collecting edge cases to make the tool smarter.

Top comments (0)