DEV Community

Kirnu (كرنو)
Kirnu (كرنو)

Posted on AI-assisted

Arabic digits in web forms - why a valid phone number gets rejected, and how to fix it in JavaScript

A user in Riyadh types their mobile number with Arabic digits: ٠٥٠١٢٣٤٥٦٧. Your form rejects it, even though it's a perfectly valid number. Or worse, it accepts it and stores it as typed, so your database now holds two "different" numbers for the same person: 0501234567 and ٠٥٠١٢٣٤٥٦٧. Arabic-Indic digits represent the same decimal values as 0–9, but they are different Unicode characters.

This post covers where the problem comes from, four traps in JavaScript, and then how to normalize, validate and store one canonical form.

Disclosure: I maintain an open-source Arabic text library that has a digit-conversion function; it's mentioned briefly at the end. Every example before that is plain JavaScript.

Three sets of digits

  • Western digits 0–9: U+0030 to U+0039.
  • Arabic-Indic digits ٠–٩: U+0660 to U+0669. They may be entered by Arabic keyboard layouts or Arabic locale settings.
  • Extended Arabic-Indic (Persian/Urdu) digits ۰–۹: U+06F0 to U+06F9. Most look like the Arabic ones, but some are shaped differently, such as ۴, ۵ and ۶ (and ۷ in Urdu), and all ten are separate code points.

Number formatting also has its own marks: the Arabic decimal separator ٫ (U+066B), the Arabic thousands separator ٬ (U+066C) and the Arabic percent sign ٪ (U+066A).

Trap 1: \d doesn't match Arabic digits

/^\d+$/.test('٠٥٠١٢٣٤٥٦٧');  // false
/^\d+$/u.test('٠٥٠١٢٣٤٥٦٧'); // false
Enter fullscreen mode Exit fullscreen mode

In JavaScript, \d means 0–9 only, even with the u flag. Any validation built on \d rejects a number typed with Arabic digits.

Trap 2: Number() and parseInt() return NaN

Number('١٢٣');      // NaN
parseInt('١٢٣');    // NaN
parseFloat('١٢٫٥'); // NaN
Number('12٣');      // NaN
Enter fullscreen mode Exit fullscreen mode

JavaScript doesn't convert Arabic-Indic digits to numbers, and a single Arabic digit inside a Western number is enough to make the whole conversion fail.

Trap 3: \p{Nd} accepts more than you want

\p{Nd} (any Unicode decimal digit) looks like the fix:

/^\p{Nd}+$/u.test('٠٥٠١٢٣٤٥٦٧'); // true
/^\p{Nd}+$/u.test('৫৫৫');        // true (Bengali digits)
Enter fullscreen mode Exit fullscreen mode

It does accept Arabic digits, but also the digits of many other scripts: 770 characters in 77 digit sets in Node 24. And it converts nothing: Number() still returns NaN afterwards. Validation alone isn't enough: normalize first, then validate.

Trap 4: invisible direction marks

Copy a number from an Arabic page, or from Intl output, and invisible bidi control characters can come along. On Node 24.21 / ICU 78.3:

const date = new Intl.DateTimeFormat('ar-SA', { timeZone: 'UTC' }).format(new Date(Date.UTC(2024, 2, 11)));
date;                    // '١١‏/٣‏/٢٠٢٤'
date.includes('‏'); // true: a RIGHT-TO-LEFT MARK (RLM) after the day and the month

const percent = new Intl.NumberFormat('ar-EG', { style: 'percent' }).format(0.25);
percent.includes('؜'); // true: an ARABIC LETTER MARK (ALM) after the percent sign
Enter fullscreen mode Exit fullscreen mode

These can break a strict ^...$ check that expects digits only, even though the text looks fine on screen. In numeric fields you can strip them before validating, or reject them explicitly, depending on your input policy. Don't strip them from general Arabic text: there they control the order of mixed Arabic/Latin words.

The fix: normalize, validate, then store one form

Normalization and validation are different steps: normalization changes the representation; validation decides whether the resulting value is allowed in your application.

Step one is a function that only normalizes: it converts Arabic-Indic and Persian digits to 0–9 and strips bidi marks (for numeric fields only):

// For numeric fields: strips bidi marks and converts ٠–٩ and ۰–۹ to 0–9; touches nothing else.
function normalizeDigits(input) {
  return input
    .replace(/[‎‏؜‪-‮⁦-⁩]/g, '') // bidi marks
    .replace(/[٠-٩]/g, (d) => String(d.charCodeAt(0) - 0x0660))     // ٠–٩ → 0–9
    .replace(/[۰-۹]/g, (d) => String(d.charCodeAt(0) - 0x06f0));    // ۰–۹ → 0–9
}

normalizeDigits('٠٥٠١٢٣٤٥٦٧'); // '0501234567'
normalizeDigits('٠50١٢٣٤٥٦٧'); // '0501234567' (mixed digits)
normalizeDigits('۰۵۰');        // '050'
Enter fullscreen mode Exit fullscreen mode

Don't drop separators without validating, or ١٬٢ silently becomes 12. For amounts and decimals, check the format first:

// An amount like ١٬٢٥٠٫٥ or 1250.5: thousands separators only in valid positions; null otherwise.
function parseAmount(raw) {
  const s = normalizeDigits(raw).replace(/٫/g, '.').replace(/٬/g, ',');
  if (!/^(?:\d{1,3}(?:,\d{3})+|\d+)(?:\.\d+)?$/.test(s)) return null;
  return Number(s.replace(/,/g, ''));
}

parseAmount('١٬٢٥٠٫٥'); // 1250.5
parseAmount('١٢٫٥');    // 12.5
parseAmount('١٬٢');     // null
Enter fullscreen mode Exit fullscreen mode

For phone numbers, normalizing digits isn't enough: 0501234567, +966501234567 and 00966501234567 are one number in three formats. Convert to one international form (E.164) and store that:

// Saudi mobile in local or international form → '+9665XXXXXXXX', or null if the format doesn't match.
// Spaces and hyphens are allowed because they're only formatting.
function toSaudiE164(raw) {
  const s = normalizeDigits(raw).replace(/[\s-]/g, '');
  if (/^05\d{8}$/.test(s)) return '+966' + s.slice(1);
  if (/^(?:\+|00)9665\d{8}$/.test(s)) return '+966' + s.replace(/^(?:\+|00)966/, '');
  return null;
}

toSaudiE164('٠٥٠١٢٣٤٥٦٧');        // '+966501234567'
toSaudiE164('+٩٦٦ ٥٠ ١٢٣ ٤٥٦٧');   // '+966501234567'
toSaudiE164('00966501234567');    // '+966501234567'
toSaudiE164('0501234567‏');  // '+966501234567'
toSaudiE164('٠٥٠١٢٣٤٥');          // null
Enter fullscreen mode Exit fullscreen mode

This is a simplified example that checks the format only, not whether the number is assigned or active.

Three practical rules:

  • Store the canonical form, not what the user typed: Western digits for numbers, E.164 for phone numbers, so the same value never exists in two forms and searches don't miss it.
  • Normalize and validate on the server too. Browser validation can be bypassed, and requests may come from other clients.
  • Don't rely on inputmode alone. inputmode="numeric" asks for a number keypad; it doesn't guarantee which digits you receive.

Displaying numbers: pick the numbering system explicitly

Display matters too. Don't rely on the default when you want Arabic or Western digits on screen. On Node 24.21 / ICU 78.3, new Intl.NumberFormat(locale).format(1234567.89) gives:

  • ar: 1,234,567.89
  • ar-SA and ar-EG: ١٬٢٣٤٬٥٦٧٫٨٩
  • ar-AE: 1,234,567.89
  • ar-MA: 1.234.567,89 (dot for thousands, comma for decimals)

Defaults differ by country and can change between ICU versions, so put the numbering system in the locale:

new Intl.NumberFormat('ar-SA-u-nu-latn').format(1234567.89); // '1,234,567.89'
new Intl.NumberFormat('ar-u-nu-arab').format(1234567.89);    // '١٬٢٣٤٬٥٦٧٫٨٩'
Enter fullscreen mode Exit fullscreen mode

Use this for display only: never parse formatted text back into a number; keep the original numeric value.

Summary

  • \d and Number() don't understand Arabic-Indic digits; \p{Nd} accepts too much and converts nothing.
  • Normalize digits first, then validate the format, then store one canonical form (E.164 for phones), in the browser and on the server.
  • Don't drop separators without validating, and strip bidi marks only from numeric fields.
  • When displaying, set -nu-latn or -nu-arab explicitly.

Disclosure: my open-source library @kirnu/arabic-core has a tested toWesternDigits function that converts Arabic-Indic and Persian digits (and separators between digits). For a non-technical explanation of Arabic vs Western digits, there's a short guide on Kirnu (Arabic).

Have you hit this in your apps? Where do you normalize: in the browser, on the server, or in the database?

Top comments (1)

Collapse
 
mohith_kumar_05846f3211f3 profile image
Mohith kumar •

Great catch, and an easy one to miss if nobody on the team types with Eastern Arabic numerals. Normalising digits before validation should be standard, along with accepting spaces and a leading plus. Building chatform.in taught me the general rule: validation that's stricter than the data it protects just loses real users.