DEV Community

FUG
FUG

Posted on

Character Counts in JavaScript: Code Units, Code Points, Graphemes, and UTF-8 Bytes

As of October 8, 2026

"\u{1F44D}\u{1F3FD}".length is 4. A reader looking at the resulting single modified thumbs-up symbol sees one symbol. JavaScript sees four UTF-16 code units. Neither observation is broken; they answer different questions. The bug starts when a product requirement says “characters” without naming the unit.

We make browser tools, including a counter that shows counts with and without spaces and UTF-8 bytes. That interface is deliberately more specific than a single number: a tweet field, a storage column, and a visible counter may each have a different definition of length.

Four units hiding behind one word

JavaScript strings are sequences of UTF-16 code units. String.prototype.length reports that representation. Basic Latin letters normally occupy one code unit. Many code points outside the Basic Multilingual Plane occupy a surrogate pair, so they take two. Emoji sequences can contain multiple code points: a base symbol, a skin-tone modifier, a variation selector, or a zero-width joiner sequence.

A Unicode code point is the numbered value assigned by Unicode, such as U+0061 for a. A grapheme cluster is a user-perceived text element. The cluster can be one code point, but it can also be a sequence. UTF-8 bytes are encoded storage or transport units; their number depends on encoding rather than on what is visibly drawn.

Example Visible symbols UTF-16 code units Code points UTF-8 bytes
A 1 1 1 1
é (precomposed) 1 1 1 2
U+1F600 (grinning face) 1 2 1 4
U+1F44D U+1F3FD (modified thumbs-up) 1 4 2 8

The second row carries an extra warning. The visible é can also be written as e followed by a combining acute accent. That alternative has two code points and three UTF-8 bytes while still usually displaying as one grapheme cluster. A rule that needs stable comparison may also need a normalization policy; counting alone does not supply one.

Use an explicit counter

This small utility makes the four choices visible. It uses for...of for code points, Intl.Segmenter for grapheme clusters when available, and TextEncoder for UTF-8 length.

function counts(text, locale = "en") {
  const codeUnits = text.length;
  const codePoints = Array.from(text).length;
  const bytesUtf8 = new TextEncoder().encode(text).length;

  const graphemes = typeof Intl.Segmenter === "function"
    ? Array.from(
        new Intl.Segmenter(locale, { granularity: "grapheme" }).segment(text)
      ).length
    : null;

  return { codeUnits, codePoints, graphemes, bytesUtf8 };
}

console.log(counts("A\u{1F600}\u{1F44D}\u{1F3FD}"));
// { codeUnits: 7, codePoints: 4, graphemes: 3, bytesUtf8: 13 }
Enter fullscreen mode Exit fullscreen mode

Array.from(text) and [...text] iterate by code point, so they avoid splitting a surrogate pair into two values. They do not solve the grapheme problem: the modifier in the U+1F44D U+1F3FD sequence remains a second code point. Intl.Segmenter asks the platform to apply Unicode text-segmentation rules, which is much closer to the count a person expects in an input box.

Treat the fallback carefully. Returning null is more honest than quietly substituting code-point length and labelling it a grapheme count. If a server enforces a visible-symbol limit, use a compatible segmentation library or define the accepted client platforms; client-side advice alone cannot enforce a server rule.

“With spaces” is a policy, not another encoding

The difference between characters with spaces and characters without spaces is usually a filtering rule applied before counting. On the counter page, the regular character total includes spaces and line breaks, while the no-spaces total excludes whitespace. That means a paragraph break affects one total and not the other.

Do not implement “without spaces” as text.replaceAll(" ", "") unless a single ASCII space is truly the stated policy. Tabs, newlines, non-breaking spaces, and other Unicode whitespace may be present after a paste. For a broad whitespace rule, use a Unicode-aware regular expression:

function countWithoutWhitespace(text) {
  return Array.from(text.replace(/\s/gu, "")).length;
}
Enter fullscreen mode Exit fullscreen mode

This example deliberately returns code-point length. If the product promise is “visible characters without whitespace,” segment the filtered string into graphemes instead. Also decide whether line endings are normalized before validation. \r\n is two code units but represents one line break in many environments; a copied Windows text block can otherwise surprise both users and tests.

Match the counter to the boundary

The right question is: what will reject, store, bill for, or display this text? A database byte limit calls for encoded-byte validation using the exact database encoding. A JavaScript API documented in UTF-16 units may legitimately use .length. A field described to ordinary readers as “up to 50 characters” should normally consider grapheme clusters, then explain exceptions if the receiving platform uses another rule.

Avoid claiming that a generic counter predicts every external service. Some platforms use weighted limits, normalize text, treat URLs specially, or apply their own version of Unicode segmentation. The counter page offers presets for X posts, meta descriptions, Instagram captions, and essays, but a final paste into the destination remains the authoritative check.

A useful UI exposes the chosen measure in the label itself: “50 visible characters,” “256 UTF-8 bytes,” or “280 platform units.” It should show the remaining amount from the same function used to block submission. A meter driven by graphemes alongside a server rule driven by bytes creates the exact off-by-several error it was meant to prevent.

Test strings belong in the suite

ASCII-only tests make almost every length function look correct. Add cases for a supplementary-plane character, combining marks, a skin-tone sequence, a family emoji joined with zero-width joiners, tabs, line endings, and non-breaking spaces. Store the expected unit in the test name, not merely a bare integer.

Requirement A defensible measurement Test that catches the wrong choice
Browser field limited by UTF-16 units text.length A supplementary-plane emoji
Reader-facing symbol counter Grapheme segmentation A modified symbol or a joined emoji sequence
API or database byte quota TextEncoder().encode(text).length Korean text or emoji
Copy limit excluding whitespace Stated filter plus stated unit Tab and newline input

This vocabulary may feel fussy until a customer cannot submit a short-looking name, or a byte-limited request fails only for Korean or emoji text. Naming the unit turns a mysterious edge case into an inspectable contract.

FAQ

Is string.length a character count?

It is a UTF-16 code-unit count. It matches a simple character count for many basic strings, but not for all Unicode text.

Does spreading a string count what a person sees?

Not always. [...text] counts code points, while one visible grapheme can contain several code points.

Why do UTF-8 byte counts matter?

They matter where an encoded payload has a byte ceiling, such as a protocol or storage boundary. They are not automatically a reader-facing character count.

Should spaces be removed before validation?

Only when the stated rule excludes whitespace. Preserve the original text for display unless the product explicitly transforms it.

Compare words, character totals, whitespace-free totals, and UTF-8 bytes with our free counter

Top comments (0)