DEV Community

Cover image for How to test an LLM chatbot or a RAG app in 2026: a tester's checklist
AutomationDataCamp
AutomationDataCamp

Posted on Originally published at automationdatacamp.com

How to test an LLM chatbot or a RAG app in 2026: a tester's checklist

More and more teams ship a chatbot or a “chat with your documents” feature, and testers are asked to sign it off. The usual tools still apply, but the method changes: the same question can get two different answers, the answer depends on documents the model retrieved, and the user can try to talk the model into misbehaving. Here is how we approach it, with the public references behind each check.

TL;DR

  • Don’t assert exact text: assert properties (required facts, refusal, format, sources) and run each case several times.
  • Test retrieval and generation separately: a wrong answer often comes from the wrong documents, not from the model.
  • Use the OWASP Top 10 for LLM Applications (2025) as a security checklist: prompt injection, sensitive information disclosure, system prompt leakage, excessive agency…
  • Check the AI Act transparency rule: since 2 August 2026, users must be told they are interacting with an AI system.
  • Tools exist: promptfoo (open-source evals and red teaming), Ragas (RAG metrics), and Playwright for the UI.

Why do classic assertions break on a chatbot?

Because the output is not deterministic. Ask the same question twice and you may get two correct answers worded differently, or one correct and one wrong. An assertion like expect(answer).toBe("Your order ships in 3 days") fails on a good answer and tells you nothing about a bad one.

What works better is to assert properties of the answer:

  • it contains the required facts (the delay, the price, the right product name);
  • it does not contain forbidden content (another customer’s data, an internal URL, a competitor’s price);
  • it refuses what is out of scope, instead of inventing;
  • it respects the expected format (JSON schema, length, language);
  • it cites its sources when the product promises it.

Then run each case several times and track a pass rate, not a single green or red. A case that passes 7 times out of 10 is a finding, not a pass.

How do you test a RAG application?

A RAG (retrieval-augmented generation) app first retrieves documents, then generates an answer from them. Test the two steps separately, otherwise you cannot tell where a wrong answer comes from.

  • Retrieval: for a set of questions with known answers, check that the right passages come back. Ragas calls this context recall: how many of the relevant documents were successfully retrieved.
  • Generation: check that the answer sticks to what was retrieved. Ragas calls this faithfulness: how factually consistent the response is with the retrieved context, on a 0 to 1 scale.

Build a small reference set first: 30 to 50 real questions, with the expected answer and the document it should come from. It becomes your regression suite every time someone changes the prompt, the model or the chunking.

Which security tests come from the OWASP LLM Top 10?

The OWASP Top 10 for LLM Applications (2025) is the most used public checklist. Several entries translate directly into test cases:

  • LLM01 Prompt Injection: user input that alters the model’s behaviour. Test it directly (“ignore your instructions and…”) and indirectly, by putting hidden instructions in a document or web page the app will read.
  • LLM02 Sensitive Information Disclosure: try to obtain another user’s data, keys or internal information.
  • LLM05 Improper Output Handling: check what happens when the model’s output is rendered or passed on, for example HTML or script in an answer displayed in the page.
  • LLM06 Excessive Agency: if the bot can call tools (send an email, cancel an order), check it cannot do more than the user is allowed to.
  • LLM07 System Prompt Leakage: ask for the instructions, in several languages and phrasings.
  • LLM09 Misinformation: questions with no answer in the documents; the right behaviour is to say so.
  • LLM10 Unbounded Consumption: very long inputs, rapid repeated requests; check limits and costs.

What does the AI Act add for a chatbot?

A test case that is easy to forget. According to the European Commission, the transparency obligations of Article 50 apply since 2 August 2026: among other things, people must be informed that they are interacting with an AI system. Check that the disclosure is there, visible, and in the user’s language. (The high-risk rules are a separate matter: since the AI Omnibus they apply from 2 December 2027.)

Which tools help?

  • promptfoo: an open-source CLI and library for evaluating and red-teaming LLM apps, usable in CI (a GitHub Action exists). Good for running your reference set and attack prompts on every change.
  • Ragas: RAG metrics such as context precision, context recall, faithfulness and response relevancy.
  • Playwright: for the chat interface itself (streaming, errors, disclosure banner, accessibility), as for any web UI.

Tools score; they do not decide. A human still reviews the failing cases and the thresholds.

A starter checklist

  1. A reference set of 30 to 50 real questions with expected facts and source documents.
  2. Each case run several times, with a target pass rate.
  3. Retrieval checked on its own (right passages returned).
  4. Answers checked against the retrieved context (no invented facts).
  5. Out-of-scope and unanswerable questions: the bot says it does not know.
  6. Direct and indirect prompt injection attempts.
  7. System prompt and other users’ data cannot be extracted.
  8. Tool calls limited to what the user is allowed to do.
  9. Model output safely rendered in the page.
  10. AI disclosure visible to the user.

Over to you

How do you test your chatbot today? Pass rates, golden datasets, red-teaming prompts, or still mostly by hand? I'd love to compare approaches in the comments.

Sources: OWASP — Top 10 for LLM Applications (2025) · OWASP — LLM01:2025 Prompt Injection · Ragas — Available metrics · promptfoo — Introduction · European Commission — Navigating the AI Act


Originally published on AutomationDataCamp. Written with AI assistance; definitions and dates were checked against the sources listed above on 5 October 2026.

AutomationDataCamp is an online software testing academy (Playwright, API testing, CI, AI-assisted testing, ISTQB prep). More on our site.

Top comments (0)