Customer-support systems can answer questions quickly, but a useful support agent needs more than a good-looking response. It needs to understand what happened before.
While working on ResolveIQ Lite, I focused on the LLM layer of our AI-powered customer-support agent. My goal was to make the model use relevant customer history while avoiding information that was never provided.
The problem
Imagine a customer previously said:
“My order A104 arrived damaged. I want a replacement.”
Later, they send:
“I still haven't received my replacement.”
Without memory, the AI doesn't know which order the customer means.
With relevant memory, it can connect the message to A104.
But there is an important limitation: the AI should not invent information.
For example, it should not say:
“Your replacement was shipped yesterday and will arrive tomorrow.”
unless that information actually exists in the available history.
This became the main focus of my implementation.
How the system works
The flow is:

A teammate worked on the memory/retrieval component. I worked on agent.py, which takes the recalled memories and the current customer message and generates the response.
The recalled information is added to the LLM prompt as PAST HISTORY

This gives the model a clear distinction between current information and previously remembered information.
Prompt engineering became important
Simply giving an LLM memory doesn't guarantee that it will use it correctly.
I added rules telling the agent to:
Use previous history when relevant.
Refer to earlier orders or problems when appropriate.
Never pretend to remember something that isn't available.
Never invent tracking numbers, dates, or policies.
Never claim that a refund or replacement was completed without evidence.
Ask for missing information when necessary.
Keep responses short, warm, and specific.
One useful test was running the same message with and without memory.
With memory
The agent can identify the relevant previous incident.
Without memory

Now the agent should ask for the missing order information instead of pretending it remembers A104.
This helped verify that the memory was actually influencing the response.
One failure taught me an important lesson
During testing, I encountered a response similar to:
“Thank you for confirming it's order A104.”
The customer had not confirmed A104.
The model had correctly retrieved A104 from memory, but then treated the remembered information as if the customer had just provided it.
This showed me that agent memory isn't only about retrieval. The model also needs clear rules about what is known, what is remembered, and what is actually confirmed by the customer.
I tightened the prompt accordingly and continued testing.
Adding model fallback
I also added a fallback mechanism so the response generator doesn't depend on only one model.
The agent tries the configured models in sequence:

If one model fails, the next configured model can be tried.
I also remove ... sections before returning the customer-facing response.
What I learned
Memory needs boundaries. Giving an LLM more context doesn't automatically make it more accurate.
Test with and without memory. This is a simple way to check whether retrieved history is actually affecting the response.
Test failure cases. Questions about missing tracking numbers, delivery dates, or policies can reveal when the model starts inventing information.
Prompt engineering is part of system design. The prompt defines what the agent can and cannot claim.
Test the fallback deliberately. A fallback is useful only when the configured models are actually available.
The result
The LLM component of ResolveIQ Lite now follows a s imple principle:

The goal isn't simply to make the AI sound confident.
It is to make the AI use customer history when it has it, ask for information when it doesn't, and avoid presenting assumptions as facts.
That was my main learning from building the LLM layer of ResolveIQ Lite.

Top comments (1)
that 'thank you for confirming A104' slip is a good catch. it sounds polite but quietly upgrades a memory into something the customer supposedly said. i'd test the same case with two open orders too, where the safest answer is 'i found A104 and B218, which replacement do you mean?'