DEV Community

Noor Ul Huda
Noor Ul Huda

Posted on

Orderly: An AI WhatsApp Order-Taker That Refuses to Guess

Orderly: An AI WhatsApp Order-Taker That Refuses to Guess

Small shops often receive customer orders through WhatsApp like:

β€œ2 kg rice, 1 litre oil, and 1 kg sugar.”

But real messages can be much messier β€” Telugu + English, corrections, unclear quantities, or voice notes.

The difficult part isn't just understanding the message.

The real problem is knowing whether the AI understood the complete order correctly.

That's why I built Orderly.

What I Built

Orderly is an AI-assisted WhatsApp order-taker for small shops.

It converts messy text and voice orders into structured orders and checks them before confirmation.

The core principle is:

AI extracts. Code verifies. Catalog decides. Human confirms when evidence is insufficient.

If the order is clear and the evidence is sufficient, Orderly can mark it CONFIRMED.

If something is unclear, missing, incomplete, or unsafe to verify, it returns NEEDS_REVIEW instead of guessing.

Demo

πŸŽ₯ Watch the Orderly Demo Video

The demo shows Orderly processing customer orders, validating extracted items, and handling cases where confirmation should not be automatic.

Video Transcript

Hi, this is Orderly, an AI-powered WhatsApp order-taker for small shops.

It handles messy text and voice orders, including Telugu-English messages.

The key principle is:

AI extracts. Code verifies. Catalog decides. Human confirms when evidence is insufficient.

Orderly checks the extracted items against the catalog and validates the customer's request.

If the evidence is sufficient, the order can be CONFIRMED.

If anything is unclear or incomplete, Orderly returns NEEDS_REVIEW instead of guessing.

The goal is simple:

Never falsely confirm an incomplete order.

Try Orderly, explore the code, and share your feedback or ideas for improving safe AI-powered ordering for small shops.

Code

πŸ”— GitHub Repository

Orderly is built as a small Python application with:

  • Web interface
  • CLI
  • Deterministic validation
  • Catalog checks
  • Automated safety tests
  • Fake extractors for deterministic adversarial testing

How I Built It

The pipeline is:

Customer message β†’ transcription β†’ AI extraction β†’ deterministic validation β†’ catalog verification β†’ confirmation/review

For voice orders, speech is transcribed before extraction.

The AI produces structured order information, but the application does not treat the AI response as proof that the extraction is complete or correct.

The deterministic code checks the extracted information against the customer's message and the shop catalog.

This separation is intentional.

An LLM can misunderstand a message, hallucinate an item, or simply miss something.

So the system is designed so that:

AI proposes. Code verifies. Evidence determines whether confirmation is allowed.

Safety First

One of the biggest risks I focused on was customer-item omission.

For example, if a customer says:

β€œ2 kg rice and 5 litres sunflower oil”

but the AI extracts only:

β€œ2 kg rice”

a normal pipeline might accidentally confirm the incomplete order.

Orderly is designed to treat this as a NEEDS_REVIEW situation rather than assuming that the extracted list is complete.

I specifically tested adversarial cases involving:

  • Multiple products with one item omitted
  • Three products with one or more items omitted
  • Middle-item omission
  • Final-item omission
  • The same product with different sizes
  • Corrections
  • Mixed Telugu + English orders
  • β€œAnd also” constructions
  • Long messages containing unrelated text

The goal is not to maximize the number of CONFIRMED orders.

The goal is to avoid falsely confirming an incomplete order.

Why Open Innovation Matters

This project uses open-source AI because useful AI systems should not always require a closed cloud service.

Open models and open-source tools make it possible for developers to experiment, inspect the pipeline, run models locally, and build systems for specific real-world problems.

For a small-shop use case, local and open tooling can also provide more control over how customer data is processed.

The important lesson for me is that open AI becomes more useful when it is combined with traditional deterministic software rather than being trusted blindly.

My Agent Session

I used AI-assisted development to inspect the project, improve the implementation, create adversarial safety tests, investigate edge cases, and verify the application.

The development process focused heavily on finding cases where an AI extraction could appear correct while still being incomplete.

That led to a stronger design principle:

Never confuse a successful AI response with proof that the customer's complete request was understood.

Testing

The automated test suite contains 156 tests: 155 passed, 0 failed, and 1 expected failure (xfail).

Orderly automated test suite showing 156 tests collected, with 155 tests passed, 0 failed, and 1 expected failure (xfail). The terminal output demonstrates the project's deterministic safety, validation, and adversarial testing results.

The expected failure is intentional. It documents a known limitation involving an unknown product such as bellam when the extractor completely omits that product and the customer message does not provide enough recognizable evidence for the completeness checker to detect the omission.

The limitation is documented rather than hidden.

The tests are primarily deterministic safety and validation tests. They use controlled/fake extraction behavior to test whether the application fails safely when extraction is incomplete.

Real-Gemma validation was not performed in the current environment, so these test results should not be interpreted as evidence of real-model extraction quality.

Current Limitations

Orderly is a prototype, not a production-ready autonomous ordering system.

Completeness is difficult to guarantee perfectly for:

  • Completely unknown products
  • Ambiguous language
  • Speech-recognition errors
  • Unusual phrasing
  • Product names appearing in unrelated text
  • Corrections that are difficult to resolve
  • Messages where the customer's intent cannot be reliably identified

A particularly important limitation is that a completeness checker cannot mathematically prove that an LLM extracted every possible customer intent from arbitrary natural language.

Therefore, the safest behavior is to request human confirmation whenever sufficient evidence is unavailable.

Because of this documented limitation, Orderly should not be considered production-safe for unattended autonomous ordering.

Final Thoughts

Orderly started as an AI order-taking idea, but the more important lesson became verification.

AI is good at understanding messy human language.

Code is better at enforcing deterministic rules.

Putting the two together creates a safer system than relying on either one alone.

AI extracts. Code verifies. Catalog decides. Human confirms when evidence is insufficient.

The most important design decision was not making the AI say CONFIRMED.

It was making sure that when the system cannot establish that the customer's complete request is represented, it refuses to guess.

Try Orderly, explore the code, and share your feedback or ideas for making AI-powered ordering safer for small shops.


devchallenge #weekendchallenge #hf26challenge #opensource #ai #python

Top comments (0)