Customer returns and refund requests are among the costliest operational friction points for modern e commerce platforms. Traditional solutions force businesses to choose between two undesirable extremes:
- Rigid rule engines that deny valid edge cases, frustrating loyal shoppers.
- Naive generative chatbots that can be easily manipulated through prompt injection, promising free money and unauthorized chargebacks.
To solve this, I designed and built an autonomous AI Customer Support Refund System. It balances instant automated decisions with deterministic policy checks, velocity safeguards, and a supervisor review loop.
Here is a full breakdown of how the platform is architected and built.
1. High Level Architecture and the Tracer Bullet Approach
Instead of building individual layers in isolation, we followed the Tracer Bullet build strategy. Every development milestone implemented a thin, working slice through the database, backend services, AI decision layer, and user interface.
+--------------------------------+
| React 18 + TypeScript SPA |
| (Tailwind, Lucide, TanStack) |
+---------------+----------------+
|
Session Cookies (HTTP Only, Lax)
|
v
+--------------------------------+
| FastAPI REST Layer |
| (Clean Architecture Services) |
+---------------+----------------+
|
+-----------------------+-----------------------+
| |
v v
+-----------------------+ +-----------------------+
| PostgreSQL 16 Engine | | LiteLLM Policy Core |
| (SQLAlchemy 2 Async) | | (Ollama, OpenAI, Gem) |
+-----------------------+ +-----------------------+
The Technology Stack
- Backend: Python 3.11 with FastAPI, Pydantic 2, SQLAlchemy 2 (asyncio with asyncpg), and LiteLLM.
- Database: PostgreSQL 16 with normalized relational models for customers, orders, order items, refund claims, LLM providers, and audit logs.
- Frontend: React 18, TypeScript, Vite, TanStack Query (React Query), and Tailwind CSS.
- Infrastructure: Fully containerized multi stage Docker Compose environment with health checks.
2. Deterministic Guardrails Meet Generative AI
One of the biggest pitfalls when integrating large language models into fintech or support operations is unpredictability. An LLM should not make decisions in an unconstrained void.
Our engine applies a layered decision pipeline:
[Customer Claim Submission]
|
v
+-------------------------------+
| 1. Fast Path Validation | -> Window checks (within 30 days)
| (Pure Code & Database) | -> Item condition and final sale tags
+---------------+---------------+ -> Velocity & sum limits (7 day caps)
|
v
+-------------------------------+
| 2. Risk Scoring & History | -> Lifetime return rate calculations
| (Customer Profile Metrics)| -> Fraud cluster anomaly detection
+---------------+---------------+
|
v
+-------------------------------+
| 3. Structured AI Reasoning | -> Evaluates nuanced testimony
| (LiteLLM Provider Engine) | -> Strict JSON output schema
+---------------+---------------+ -> Auto fallback (OpenAI to Ollama)
|
v
+-------------------------------+
| 4. Audit Trail & Disposition | -> Approved: Instant settlement
| (Immutable Event History) | -> Escalated: Queued for supervisor
+-------------------------------+ -> Denied: Policy citation logged
Structured Policy Output
The AI decision engine produces typed outputs enforcing structured Pydantic schemas:
class PolicyEvaluationResult(BaseModel):
decision: Literal["Approved", "Denied", "Escalated"]
confidence_score: float = Field(ge=0.0, le=1.0)
reasoning_summary: str
policy_citations: list[str]
detected_red_flags: list[str]
matched_rules: list[str]
If a prompt injection attack is attempted (such as ignoring return windows or commanding the agent to approve high value claims), the deterministic validation layer flags anomaly markers and immediately forces an Escalated state for human review.
3. Scoped Customer Portal and Real Time Claims History
Customers interact with a dedicated three step refund wizard and history dashboard:
- Order and Item Selection: Displays verified purchases fetched strictly from session cookies. Items already claimed are disabled to prevent duplicate abuse.
- Reason and Testimony Input: Captures the customer return justification with character boundary validation.
- Instant Evaluation: Invokes the engine and renders a transparent verdict card with confidence percentage and detailed policy reasoning.
+---------------------------------------------------------------+
| AutoRefund Portal Sarah Jenkins [Customer badge] |
+---------------------------------------------------------------+
| [ File a Refund ] [ My Claims (3) ] |
|---------------------------------------------------------------|
| Total Claims: 3 | Approved: 2 | Under Review: 1 |
|---------------------------------------------------------------|
| Search: [ Filter by item or claim number... ] [All] [Appr] |
|---------------------------------------------------------------|
| Claim ID Order Date Status Action |
| REF-9021 ORD-104 Sep 26 2026 Approved [Inspect] |
| REF-8842 ORD-089 Sep 20 2026 Escalated [Inspect] |
+---------------------------------------------------------------+
Slide Over Inspection Drawer
Clicking Inspect on any claim row opens a slide over drawer containing:
- Complete order and refund item breakdowns.
- The verbatim customer statement quote.
- The AI policy evaluation with natural language rationale and confidence meters.
- Supervisor override justification alerts when manual human intervention occurred.
4. Human in the Loop: The Supervisor Review Desk
Autonomous systems must always provide an escape hatch for human judgment.
When claims exceed velocity thresholds or involve high value items, the system places the request in an Escalated queue on the Admin Dashboard. Human supervisors can:
- Review customer historical risk scores and lifetime spend.
- Inspect the initial AI decision and model explanation.
- Apply a Supervisor Override to approve or deny the claim, accompanied by mandatory written justification notes.
- Every override event is immutably recorded in the audit logs with the supervisor identifier and timestamp.
5. Testing and Verification Rigor
A financial decision engine demands comprehensive automated testing. We built an automated test harness covering backend and frontend layers:
-
137 Backend Pytest Tests:
- Unauthenticated access rejection (HTTP 401).
- Session scoping and strict tenant isolation (confirming Customer A never sees Customer B data).
- Order ownership verification and duplicate claim blocks.
- LLM fallback mechanisms and structured JSON response parsing.
-
121 Frontend Vitest Tests:
- Segmented tab switching and URL query parameter preservation.
- Client side search filtering and status pill filtering.
- Slide over inspection drawer focus trap, Escape key handling, and accessibility.
- Real time cache invalidation upon claim submission.
All tests execute inside Docker containers, ensuring complete parity between local development and production environments.
6. What I Learned
- Never let an LLM run without guardrails: Generative models are fantastic at interpreting messy human explanations, but hard constraints (deadlines, return caps, duplicate detection) belong in deterministic code.
- Multi provider resilience is essential: Using LiteLLM allowed us to seamlessly route requests between cloud providers like OpenAI and Gemini, falling back to local Ollama instances during outages.
- Session scoped security beats client provided parameters: Never trust the frontend to pass a customer identifier. Deriving user identity strictly from secure, HTTP only cookies eliminates impersonation attacks at the root.
Check out the full repository and architecture specs on GitHub!
https://github.com/abbeymaniak/ai-customer-support-refund
Top comments (0)