DEV Community

Cover image for How to Remove Special Characters in JavaScript Without Breaking Unicode Text
sam khan
sam khan

Posted on

How to Remove Special Characters in JavaScript Without Breaking Unicode Text

Cleaning text sounds simple until your input contains accented letters, punctuation, emoji, line breaks and characters from different languages.

A common approach is to use a regular expression that removes everything except A–Z, a–z and 0–9. That works for a narrow set of English-only inputs, but it can also delete valid letters from names and words such as café, München and São Paulo.

In this tutorial, we’ll build a small JavaScript function that removes special characters while preserving Unicode letters, numbers and spaces.

The problem with a basic regular expression

You may have seen code like this:

const cleaned = text.replace(/[^a-zA-Z0-9 ]/g, "");
Enter fullscreen mode Exit fullscreen mode

It removes punctuation and other characters, but it also removes letters outside the specified English alphabet.

For example:

const text = "Café ☕ costs €5!";

console.log(text.replace(/[^a-zA-Z0-9 ]/g, ""));
// "Caf  costs 5"
Enter fullscreen mode Exit fullscreen mode

The é disappears even though it is a letter. That may be unacceptable when processing names, international content or multilingual text.

A Unicode-aware solution

Modern JavaScript supports Unicode property escapes in regular expressions. We can use \p{L} for letters and \p{N} for numbers.

function removeSpecialCharacters(text) {
  return text.replace(/[^\p{L}\p{N} ]/gu, "");
}

const input = "Café ☕ costs €5!";
console.log(removeSpecialCharacters(input));
// "Café  costs 5"
Enter fullscreen mode Exit fullscreen mode

Here’s what the regular expression does:

  • \p{L} matches Unicode letters.
  • \p{N} matches Unicode numbers.
  • The literal space allows ordinary spaces.
  • [^...] matches characters outside the allowed set.
  • The u flag enables Unicode-aware matching.
  • The g flag processes all matches, not just the first one.

This preserves accented letters while removing the coffee emoji, currency symbol and punctuation.

Clean up extra spaces too

Removing symbols can leave multiple spaces behind. If you want a tidy single-line result, normalize the spacing afterward.

function cleanText(text) {
  return text
    .replace(/[^\p{L}\p{N} ]/gu, "")
    .replace(/ +/g, " ")
    .trim();
}

console.log(cleanText("Café ☕ costs €5!"));
// "Café costs 5"
Enter fullscreen mode Exit fullscreen mode

The second replacement collapses consecutive ordinary spaces into one, and trim() removes spaces at the beginning and end.

Be careful with this version if your input contains multiple lines: the first replacement removes line breaks. If preserving paragraphs matters, use a different pattern.

Preserve line breaks when needed

For multiline text, you can explicitly allow \n and \r:

function cleanMultilineText(text) {
  return text
    .replace(/[^\p{L}\p{N} \r\n]/gu, "")
    .replace(/ +/g, " ")
    .trim();
}

const input = "First line! 🚀\nSecond line: Café.";

console.log(cleanMultilineText(input));
// First line
// Second line Café
Enter fullscreen mode Exit fullscreen mode

This keeps line breaks while removing punctuation and emoji.

What about apostrophes and hyphens?

Not every punctuation mark is unwanted. Removing the apostrophe from don't produces dont, and removing a hyphen from well-known produces wellknown.

If your use case requires those characters, add them to the allowed set:

function cleanReadableText(text) {
  return text
    .replace(/[^\p{L}\p{N} '\-\r\n]/gu, "")
    .replace(/ +/g, " ")
    .trim();
}

console.log(cleanReadableText("A well-known café — don't miss it!"));
// "A well-known café  don't miss it"
Enter fullscreen mode Exit fullscreen mode

The right cleaning rule depends on what the text will be used for. Search queries, display text, usernames and data imports often need different rules.

When to use a browser-based tool instead

If you only need to clean a few passages, writing and running JavaScript may be unnecessary. A browser-based tool can be more convenient for a one-off task.

TextNivo has a Remove Special Characters tool for stripping special characters while keeping letters, numbers and spaces.

Try the tool: []

For repeatable processing inside an application, keep the logic in code and add tests for the kinds of text your users actually enter.

Final thoughts

A regular expression can remove unwanted characters in one line, but the definition of “unwanted” matters.

For international text, Unicode property escapes are generally more appropriate than an English-only character range. Decide whether you need to preserve spaces, line breaks, apostrophes or hyphens before applying the cleaning rule.

Most importantly, test your function with realistic input—not just a simple English sentence.

Top comments (0)