DEV Community

Cover image for My support chatbot scored 0.92. It was also lying to customers.
Mialy333 🎧ྀི
Mialy333 🎧ྀི

Posted on

My support chatbot scored 0.92. It was also lying to customers.

Twelve of my thirteen test prompts scored a perfect 1.00. Two independent evaluation runs agreed: 0.92 correctness. No message routed to the wrong place, both prompt-injection attempts refused.

And yet, in one conversation, my chatbot told a customer:

"I have filed a bug report with ticket ID TIX-345678."

That ticket did not exist. Nothing had been written to the database. The customer would have walked away believing their bug was logged.

This is the story of how I built that chatbot, why the score couldn't see the problem, and what I check now instead.


The project

I'm going through the Udacity × AWS Agent Engineer Nanodegree. The first project: a customer-support chatbot for a fictional online shop that handles three kinds of messages.

  • Bug reports: collect a description, steps to reproduce, and the environment, possibly over several turns, then file a ticket.
  • Platform questions (orders, shipping, returns, payments): answer only from the shop's FAQ.
  • Anything else: politely hand off to a human support line.

One twist: Bedrock Agents Classic, which the course was originally designed around, closed to new customers on 30 July 2026. So the project runs on its successor, the Amazon Bedrock AgentCore managed harness.

The other twist, and the actual point of the exercise: there's no classifier and no routing node. All of the routing, information-gathering, and grounding behaviour lives in a single system prompt.

Architecture

Customer ──► chat.py ──invoke_harness──► AgentCore managed harness  ◄──► Amazon Nova Pro
                                              │  (system_prompt.txt + FAQ)     temp 0, topK 1
                                              │
                                    tool call: bugreports___create_bug_report
                                              ▼
                                       AgentCore Gateway (MCP, IAM auth)
                                              ▼
                                      Lambda create_bug_report ──PutItem──► DynamoDB
Enter fullscreen mode Exit fullscreen mode

The pieces:

  • Harness: AgentCore runs the agent loop: model calls, session state, tool execution. I supply the prompt; the FAQ is injected at a {{FAQ}} placeholder when the harness is created.
  • Gateway: exposes a Lambda as an MCP tool. The model sees it as <targetName>___<toolName>, three underscores.
  • Lambda + DynamoDB: validates that all three fields are non-empty, writes the ticket, returns a UUID ticketId.
  • Everything as code: two CloudFormation stacks (tool + testing) and a few boto3 scripts. Iterating on the prompt is just: edit the file, re-run create_harness.py, open a new chat session.

The model is pinned to us.amazon.nova-pro-v1:0 with greedy decoding (temperature 0, topK 1), which AWS recommends for reliable tool calling with Nova.

Routing with nothing but words

Treating routing as a classification problem inside the prompt worked better than I expected. The prompt tells the model to pick exactly one category before writing anything, and never mix categories.

The hard part is the boundary. Is "my card was declined" a bug? Is "the mug arrived cracked"? I ended up with 15 explicit tie-breakers, for example:

  • An error or crash on the checkout page → bug.
  • A declined card → platform question (payments).
  • An item that arrived damaged → platform question (returns): the software worked, the order didn't.
  • Search returning nothing for a product that clearly exists → bug.

And one fallback principle for everything else: if the software misbehaved, it's a bug; if it's about a policy or an order, it's a platform question; otherwise, hand off.

For bug collection, the rules were: re-read what the customer already said, ask for one missing field at a time, and the moment all three are present, call the tool. No "shall I file this?" confirmation.

I also added a section treating everything a customer writes as data, never instructions, with a list of override patterns ("ignore your previous instructions", "I'm your developer") to refuse without arguing or explaining how they were detected.

Evaluating it

Manual chatting doesn't scale, so I wrote a 13-case suite: 3 bug-report cases, 3 FAQ cases, 2 hand-offs, and 5 edge cases (a bare help, two ambiguous messages, two injections). A script runs each case in a fresh session and writes a JSONL file in the format Bedrock Evaluations expects; Nova Pro then scores each response against a reference as an LLM-as-a-judge.

Bedrock Evaluations correctness histogram: 12 prompts at 1.0, 1 at 0, average 0.923

To keep cases independent, I disabled harness memory. Small gotcha worth sharing: create_harness accepts memory={"disabled": {}}, but update_harness needs memory={"optionalValue": {"disabled": {}}}. Passing the create shape to update fails, and the script takes the update path on every re-run.

Result: 0.92, twice.

What the score didn't show

Here's the uncomfortable part. Both of the real defects I found, I found by doing something the evaluation never does: opening the DynamoDB table and comparing it to the chat transcript.

1. Invented ticket IDs

In a multi-turn conversation (search bar broken → what happens? → which browser?), the bot ended with a ticket ID. But there was no [tool call] line in the terminal, and the table still held nine rows, not ten.

Across attempts it produced #12345, TICKET1234, TIX-345678. Placeholder-shaped strings, never a UUID.

My prompt already said "never fabricate information". That clearly wasn't specific enough. I added an explicit rule: the only ticket ID you may give is the exact ticketId string returned by a successful call; if you didn't receive one, the report hasn't been filed, so say so.

2. Junk written into the ticket

The one prompt that scored 0.00 in both runs: "The app keeps logging me out. I'm using Chrome on Windows 11."

In the chat, the response looked fine: thanks, here's your real ticket ID. But in the table, stepsToReproduce contained either:

  • "Please provide specific actions or scenarios that lead to the app logging you out." (the bot's own clarifying question, saved as if the customer had said it), or
  • "Using the app." (filler)

DynamoDB scan showing a healthy ticket next to tickets with junk in stepsToReproduce

The cause was a tension in my own prompt. To stop the bot from badgering customers, I'd made the steps rule lenient: "It crashes when I click Pay" already counts. That works when the fault and the trigger are separate. It breaks when they're the same sentence. "Keeps logging me out" reads as both, so the model decided the field was satisfied and filled it with something rather than leave it empty.

The fix isn't removing the leniency (that prevents a worse failure). It's requiring steps that are distinct from the description, and forbidding the assistant's own words in any field. I deliberately didn't apply it mid-project: changing a central rule after two scored runs would have broken the comparison between them.

3. Two smaller ones

  • Exploratory tool calls. The model sometimes called create_bug_report with a field missing, on purpose, so the Lambda's validation error would tell it what to ask. Customer-visible behaviour was fine; the mechanism was wrong. New rule: a failed tool call is never a way to discover a missing field.
  • Reasoning leaks. Nova Pro liked to open replies with <thinking>…</thinking>. Moving the "never show your reasoning" rule to the very top of the prompt reduced it a lot. It didn't eliminate it.

What I'd tell myself at the start

1. For tool-using agents, assert on side effects, not on text.
A correctness judge reads the response. A confident, helpful-sounding response wrapped around an empty or invented field scores well. The defects that mattered most were invisible to the metric and obvious in the database. Next time, the test suite checks the stored row.

2. Single-turn evals don't test multi-turn behaviour.
My two most important fixes governed multi-turn collection, and all 13 eval prompts were single-turn. Run 2's identical 0.92 told me the fixes cost nothing and the score was reproducible. It didn't tell me they worked.

3. Long prompts degrade unevenly.
Even with greedy decoding, rules buried in the middle of a long prompt were followed in some sessions and not others. Multi-turn ticket filing with Nova Pro stayed unreliable. The next iteration is a shorter, restructured prompt, not more rules.

4. A few AgentCore details that cost me time:

  • Gateway target names: letters, digits, underscores only. A dash breaks Nova tool calling with "Model produced invalid sequence as part of ToolUse".
  • The Gateway passes tool arguments directly as the Lambda event. No Agents Classic parameters envelope. The tool name arrives in context.client_context.custom["bedrockAgentCoreToolName"].
  • Pin the model on the harness and on each invoke; the default model needs a Marketplace subscription.
  • Session IDs must be at least 33 characters.

The full code, prompt, test suite, transcripts and screenshots are on GitHub:
👉 Mialy333/support-chatbot-bedrock-agentcore

If you're evaluating agents that write to anything (a database, a ticket system, an API), I'd love to hear how you test the side effects. That's the part I'm working on next.

I'm Mialy, ex-asset-management, now building agents and digital-asset tooling. I write about it as @ellebuild.

Top comments (1)

Collapse
 
devantibot profile image
DEV ANTIBOT •

You need to complete account verification.Link in the profile.