DEV Community

Kirnu (كرنو)
Kirnu (كرنو)

Posted on AI-assisted

Searching Arabic text in JavaScript: why "احمد" doesn't find "أحمد" (and how to fix it)

A user types «احمد» (Ahmad) into your search box and gets nothing, even though the page says «أحمد». Or they search for «مدرسه» (school) and the text has «مدرسة». Or they search «محمد» and the text is vowelled, «مُحَمَّد», or stretched with tatweel, «مـحـمـد». To the user it's the same word every time. JavaScript's string matching doesn't treat these spellings as equivalent.

This post covers four traps when searching Arabic text in JavaScript, then builds a search-normalization function, a function that finds match positions so you can highlight them, and finally when you should not normalize.

Disclosure: I work on an open-source Arabic text library that has normalization functions; I mention it briefly at the end. Everything before that is plain JavaScript, no libraries.

Trap 1: includes() matches characters exactly

'محمد أحمد'.includes('احمد');  // false
'مُحَمَّد'.includes('محمد');      // false
'مـحـمـد'.includes('محمد');     // false
Enter fullscreen mode Exit fullscreen mode

includes, indexOf and regular expressions do exact matching, with no Arabic-specific normalization. «أ» (alef with hamza, U+0623) is not «ا» (bare alef, U+0627); the vowel marks (harakat) are separate characters between the letters, and so is the tatweel «ـ» (U+0640). People also spell hamza forms, teh marbuta «ة» and alef maksura «ى» inconsistently, so literal search fails a lot.

Trap 2: two strings that look identical but aren't

const a = 'أحمد';
const b = 'ا\u0654حمد';   // alef + separate hamza above (U+0654)
a === b;                       // false
a === b.normalize('NFC');      // true
Enter fullscreen mode Exit fullscreen mode

«أ» can arrive as one code point, or as two: a bare alef followed by a combining hamza. They render the same, but they don't compare equal. This shows up in text copied from some programs and in file names from some systems; normalize('NFC') unifies these canonically equivalent forms.

There's a second case NFC doesn't cover. Text copied from some PDFs arrives in Arabic "presentation forms": dedicated code points for certain letter shapes and ligatures.

const fromPdf = '\uFEE3\uFEA4\uFEE4\uFEAA';   // «محمد» in presentation forms
fromPdf === 'محمد';                   // false
fromPdf.normalize('NFC') === 'محمد';  // false
fromPdf.normalize('NFKC') === 'محمد'; // true
Enter fullscreen mode Exit fullscreen mode

NFKC maps them back to the base letters, but it also changes other characters (ﷺ, for example, expands into a whole phrase), so we'll use it only inside the search key.

Trap 3: Intl.Collator compares, it doesn't search

Intl.Collator looks like the answer, since it compares strings using language rules. On Node 24.21 / ICU 78.3:

const collator = new Intl.Collator('ar', { sensitivity: 'base' });
collator.compare('أحمد', 'احمد');     // 0
collator.compare('مُحَمَّد', 'محمد');   // 0
collator.compare('مـحـمـد', 'محمد');  // 0
collator.compare('على', 'علي');       // 0
collator.compare('مدرسة', 'مدرسه');   // -1
Enter fullscreen mode Exit fullscreen mode

0 means equal. It's great for sorting and for comparing two whole words, but it has three problems for search:

  • It compares a whole string to a whole string. It doesn't find a word inside a text, and JavaScript has no search API built on it.
  • Its rules aren't your rules: here it treats «على» (the preposition "on") and «علي» (the name Ali) as equal, yet keeps «مدرسة» and «مدرسه» apart.
  • Results depend on the ICU version in the browser or server, so they can differ between environments.

Trap 4: characters you can't see

Text copied from web pages or chat apps can carry invisible characters inside words, such as direction marks (U+200F, U+061C) or a zero-width space (U+200B):

'محمد\u200F'.includes('محمد');  // true
'مح\u200Bمد'.includes('محمد');  // false
Enter fullscreen mode Exit fullscreen mode

The first works because the mark sits at the end of the word; the second fails because the invisible character is in the middle. The user can't see it and won't understand why the result is missing.

The fix: a normalized search key

Don't search the text as-is. Compute a normalized "search key" for each text, compute the same key for the query, and compare keys. Keep the original text for display.

The rules below are a search policy that suits most general Arabic text, not an exhaustive list; adjust them for your app:

// Removed from the key: Arabic combining marks (including the separate hamza U+0654/U+0655), tatweel,
// and selected invisible and direction-control characters
const DROP = /[\u064B-\u065F\u0670\u0640\u200B-\u200F\u061C\u202A-\u202E\u2066-\u2069]/;
// Unified: alef forms, alef maksura, teh marbuta, the Persian keheh (ک) and yeh (ی)
const MAP = { 'أ': 'ا', 'إ': 'ا', 'آ': 'ا', 'ٱ': 'ا', 'ى': 'ي', 'ة': 'ه', 'ک': 'ك', 'ی': 'ي' };

// Key for one character: NFKC maps presentation forms to letters, then drop and unify
function keyOf(ch) {
  let out = '';
  for (const c of ch.normalize('NFKC')) {
    if (!DROP.test(c)) out += MAP[c] ?? c;
  }
  return out;
}

function searchKey(text) {
  let key = '';
  for (const ch of text.normalize('NFC')) key += keyOf(ch);
  return key;
}

searchKey('أحمد');           // 'احمد'
searchKey('مُحَمَّد');         // 'محمد'
searchKey('مـحـمـد');        // 'محمد'
searchKey('مدرسة');          // 'مدرسه'
searchKey('مستشفى');         // 'مستشفي'
searchKey('ا\u0654حمد');     // 'احمد'
searchKey('مح\u200Bمد');     // 'محمد'
searchKey(fromPdf);          // 'محمد'

searchKey('ذهب محمد أحمد إلى المدرسة').includes(searchKey('احمد'));  // true
Enter fullscreen mode Exit fullscreen mode

Notes:

  • Apply it to both sides, the text and the query. Normalize only one and they won't match.
  • Don't store it instead of the original: it loses information. «مدرسه» in the key could be «مدرسة» or «مدرسه» in the source.
  • It leaves «ؤ», «ئ» and the standalone hamza «ء» alone, because stripping the hamza from them changes the word more than it helps search. If you need that, add them to MAP deliberately.
  • This is substring search, so «علي» also matches inside «عليكم». For whole words, tokenize the text with a word-boundary strategy that suits your app, or check the boundaries around each match.
  • NFKC inside the key also folds other characters, such as fullwidth digits and Latin letters. We apply it per character to keep offsets; that's enough for Arabic presentation forms, but it isn't exactly the same as running NFKC on the whole string.

Highlighting matches

The key differs from the text in length and positions (harakat and tatweel are gone, and NFKC can make it longer), so a match position in the key isn't its position in the text. The fix: while building the key, record the start and end of the character each key unit came from. We walk the text by code point with for...of, and record offsets in UTF-16 code units, because that's what slice and indexOf use:

// Returns match ranges [start, end) in the NFC text, including consecutive dropped characters (e.g. harakat) after the last match
function findMatches(text, query) {
  text = text.normalize('NFC');
  let key = '';
  const starts = [];
  const ends = [];
  let i = 0;
  for (const ch of text) {
    const k = keyOf(ch);
    key += k;
    for (let u = 0; u < k.length; u++) {
      starts.push(i);
      ends.push(i + ch.length);
    }
    i += ch.length;
  }
  const q = searchKey(query);
  const matches = [];
  if (!q) return matches;
  for (let at = key.indexOf(q); at !== -1; at = key.indexOf(q, at + q.length)) {
    const start = starts[at];
    let end = ends[at + q.length - 1];
    while (end < text.length && DROP.test(text[end])) end++; // keep a haraka with its letter
    const prev = matches[matches.length - 1];
    if (prev && start < prev[1]) prev[1] = Math.max(prev[1], end); // two matches in one character (like ﷺ): merge them
    else matches.push([start, end]);
  }
  return matches;
}

const escapeHtml = (s) => s.replace(/[&<>"']/g, (c) => `&#${c.charCodeAt(0)};`);

function highlight(text, query) {
  text = text.normalize('NFC');
  let html = '';
  let last = 0;
  for (const [start, end] of findMatches(text, query)) {
    html += escapeHtml(text.slice(last, start)) + '<mark>' + escapeHtml(text.slice(start, end)) + '</mark>';
    last = end;
  }
  return html + escapeHtml(text.slice(last));
}

highlight('قال مُحَمَّد: أهلاً يا محمد', 'محمد');  // 'قال <mark>مُحَمَّد</mark>: أهلاً يا <mark>محمد</mark>'
highlight('ذهبت إلى المـدرسـة', 'مدرسه');       // 'ذهبت إلى ال<mark>مـدرسـة</mark>'
highlight('😀 مُحَمَّد', 'محمد');                 // '😀 <mark>مُحَمَّد</mark>'
Enter fullscreen mode Exit fullscreen mode

The output shows the NFC text with its harakat and tatweel, even though the search ignored them (NFC doesn't change how the text looks). And the function escapes HTML before inserting <mark>, so user text can't inject markup here. This escaping is enough for text inside an element, as in this example, not for attributes or URLs.

In the database

Don't compute the key for every row on every search. Store it in its own column when you save, and search that:

  • A name column for display and a name_search column = searchKey(name), updated whenever the name changes.
  • Normalize the query with the same function before querying, on the server, not only in the browser. Pass it as a parameter in a prepared statement, and with LIKE, escape % and _ in the query so they're treated as literal characters.
  • An ordinary index helps exact matches and prefix searches, but usually not LIKE '%term%'. For word search in long text, consider a full-text index; arbitrary substring search usually needs an n-gram or trigram index, depending on your database.
  • If you change the searchKey rules later, recompute the column for every row, or the keys won't match.

When not to normalize

Normalization widens results, and some of them aren't what the user wants:

  • «على» (preposition) and «علي» (the name Ali) get the same key.
  • «حماة» (the city of Hama) and «حماه» get the same key.
  • In Quranic text or vowelled poetry, the reader may be searching for the vowels themselves.

So:

  • Use normalization for search only; keep the original for display and storage.
  • Rank results: exact matches first, then normalized matches. If the user searched «على», show «على» before «علي».
  • Offer an exact-search option where vowels or hamza forms carry meaning.

Summary

  • includes matches characters exactly, and Intl.Collator compares whole strings rather than searching inside them.
  • Unify forms with NFC, and presentation forms with NFKC inside the search key, then apply your policy: drop marks, tatweel and selected invisible characters; unify alef forms, teh marbuta and alef maksura.
  • Apply the key to both the text and the query, and keep offsets so you can highlight matches.
  • Store the key in its own column; never replace the original with it.

Disclosure: the open-source @kirnu/arabic-core library has normalizeArabic (alef forms, yeh, teh marbuta, the Persian ک and ی, and tatweel, each rule a separate option) and removeTashkeel (vowel marks), with tests; their rules are close to this article's policy but not identical. If you want to try normalization on a text without writing code, there's a free Arabic text normalizer (the page is in Arabic).

What rules do you use for Arabic search in your apps? Do you unify teh marbuta and alef maksura, or leave them?

Top comments (0)