DEV Community

Ujjwal Dubey
Ujjwal Dubey

Posted on

A Mixed-Language Test Set for WhatsApp Assistants in Gulf Businesses

Disclosure: I run NxFlowAI, an automation agency serving UAE businesses remotely from Mumbai. This post is a vendor-neutral testing pattern.

Customers in the UAE often write in more than one language in the same chat: English with Arabic words, Arabic in Latin letters, Hindi or Urdu phrases, or a voice note in between. If you build or buy a WhatsApp assistant here, test those messages before launch. This is the test set I would ask any AI automation agency building for UAE customers to run, ours included.

1. Build the set from your own inbox

Export a few hundred real messages, remove names and numbers, and tag each one:

- id: m041
  text: "<anonymised real message>"
  languages: [en, ar-latin]
  intent: booking
  expected: route_to_person   # or auto_reply / draft_for_approval
  notes: "time mentioned in words, not digits"
Enter fullscreen mode Exit fullscreen mode

2. Cover the awkward cases on purpose

  • Same intent, three phrasings: English, Arabic script, Arabic in Latin letters.
  • Language switch mid-conversation.
  • Numbers written in Arabic-Indic digits and in Western digits.
  • Area names spelled several ways.
  • A voice note with no text.
  • Short replies ("ok", "tmrw", a single emoji) that only make sense with context.

3. Score routing, not just reply quality

def score(case, result):
    if case["expected"] == "route_to_person":
        return result.routed_to_person          # never auto-reply here
    if result.language_confidence < THRESHOLD:
        return result.routed_to_person          # unsure = a person
    return result.intent == case["intent"]
Enter fullscreen mode Exit fullscreen mode

The most important rule: when the system is unsure about language or intent, it hands over. A wrong confident reply in the wrong language costs more than a short wait.

4. Reply in the customer's language, or say who will

If the assistant cannot reply well in a language, the honest message is a short acknowledgement and a handoff to a team member who can.

5. Re-run after every change

Prompt edits, model changes and new templates can all shift results. Keep the set in the repo and run it in CI.

We put a set like this together during a 72-hour audit before any build. For the bigger question of whether you need custom AI at all, see our custom AI vs off-the-shelf chatbots write-up.

Top comments (0)