DEV Community

How I Built Proactive Churn Alerts Using Hindsight Reflection

How I Built Proactive Churn Alerts Using Hindsight Reflection

How can an AI support copilot detect a customer's growing frustration before they explicitly ask for a manager?

Most customer support platforms react to escalation signals too late.

A supervisor is often alerted only after a customer:

  • Explicitly demands a manager
  • Posts a public complaint
  • Threatens a chargeback
  • Repeatedly contacts support

By that point, the customer's sentiment may have already deteriorated significantly.

To address this, I designed a support copilot that tracks customer frustration trajectories and surfaces proactive escalation alerts before the customer explicitly demands intervention.

By integrating Hindsight into a FastAPI + React architecture, the system combines historical conversation experiences with Hindsight's reflection engine (areflect) to generate higher-level opinions about customer effort and frustration risk.

This article explains how the system tracks multi-session effort, enforces risk thresholds, and surfaces proactive escalation signals to support representatives.


Architecture Overview






The system operates as a rep-facing support copilot.

It intercepts incoming support queries, retrieves historical customer context, and combines raw experiences with reflection opinions before sending the context to the LLM.

Architecture Diagram


┌───────────────────────────────────────────────────────────────────┐
│                        React + Vite UI                            │
│                                                                   │
│  Ticket Queue       Conversation Thread      Agent Memory Panel   │
│  Risk Badges        Escalation Banner        Customer Context     │
└──────────────────────────────┬────────────────────────────────────┘
                               │
                               │ HTTP / REST API
                               ▼
┌───────────────────────────────────────────────────────────────────┐
│                       FastAPI Backend                             │
│                                                                   │
│  Ticket Routes       Fallback Engine       Core Memory            │
│  /api/tickets        Local Cache           Pinned Facts            │
│                                                                   │
│                  Risk & Trajectory Engine                         │
└───────────────────────┬────────────────────────┬──────────────────┘
                        │                        │
                 async recall/reflect         inference
                        │                        │
                        ▼                        ▼
┌────────────────────────────────┐    ┌─────────────────────────────┐
│       Hindsight Cloud           │    │       Groq LPU Engine       │
│                                │    │                             │
│  • Retain Experiences          │    │  • gpt-oss-120b             │
│  • Scoped Tag Recall           │    │    Primary                   │
│  • Reflect & Form Opinions     │    │  • qwen3-32b                │
│                                │    │    Fallback                  │
└────────────────────────────────┘    └─────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

Core Components

Component Responsibility
FastAPI Backend API contracts, local caches, trajectory calculation, prompt assembly
Groq LPU Engine Primary and fallback LLM inference
Hindsight Memory Temporal experiences, tagged recall, and reflection opinions

The architecture uses arecall for historical retrieval and areflect to generate higher-level opinions about customer behavior.


The Core Problem: Customer Effort Happens Over Time






A customer might contact support three times about different minor problems.

Individually, none of those conversations may appear particularly serious.

But together, they represent increasing customer effort.

A stateless LLM looking at only the current ticket cannot easily see this pattern.

For example:

Session 1
Delivery issue
      ↓
Session 2
Replacement issue
      ↓
Session 3
Still waiting
      ↓
Customer effort increases
      ↓
Potential escalation risk
Enter fullscreen mode Exit fullscreen mode

To track this progression, the system combines:

Semantic memory recall + Hindsight reflection + risk invariants

Whenever an issue is marked resolved, an asynchronous reflection task analyzes the customer's historical threads and updates an opinion containing information such as:

  • Customer effort count
  • Sentiment trajectory
  • Repeat-contact pattern
  • Overall churn risk

Risk Tier Invariants




One of the most important design decisions was preventing the escalation system from becoming too sensitive.

The system uses three risk tiers:

Risk Tier Condition
Normal 1 contact turn + stable sentiment
Watch ≥ 2 contact turns OR declining sentiment
Escalate ≥ 3 contact turns AND declining sentiment

The critical rule is:

Escalation requires both repeated contact and declining sentiment.

This prevents a single negative message from immediately producing an escalation alert.


Code-Backed Implementation

1. Asynchronous Hindsight Reflection



When a support representative resolves a case, the system triggers Hindsight reflection asynchronously.

async def reflect(
    self,
    customer_id: str
) -> Optional[dict]:

    """
    Triggers Hindsight reflection after an issue
    is resolved.

    Updates opinions on customer effort trajectory
    and frustration risk.
    """

    client = self._get_client()

    if not client:
        return self._local_opinions.get(customer_id)

    try:
        res = await client.areflect(
            bank_id=self.bank_id,
            query=(
                f"Assess customer effort trajectory, "
                f"repeat contact patterns, and "
                f"frustration churn risk for {customer_id}"
            ),
            tags=[customer_id]
        )

        return {
            "customer_id": customer_id,
            "confidence": 0.94,
            "reflection_text": getattr(
                res,
                "text",
                ""
            ) or ""
        }

    except Exception as e:
        logger.error(
            f"Hindsight reflect exception: {e}"
        )
        return None
Enter fullscreen mode Exit fullscreen mode

Why asynchronous reflection?

Reflection can be computationally heavier than a normal recall operation.

Instead of making the customer-facing response wait, the system performs reflection when the ticket is resolved.

Ticket Resolution
       │
       ▼
Async Reflection
       │
       ▼
Hindsight Opinion
       │
       ▼
Updated Risk Context
Enter fullscreen mode Exit fullscreen mode

This keeps the live support interaction separate from the heavier memory-consolidation process.


2. Risk Profile Classifier

The risk engine evaluates:

  • Contact frequency
  • Distress indicators
  • Sentiment trend
def compute_risk_profile(
    self,
    customer_id: str,
    threads_count: int,
    messages_text: str = ""
) -> RiskProfile:

    """Derives Customer Effort Trajectory
    & RiskProfile with strict invariant assertions."""

    contact_count = max(1, threads_count)

    distress_keywords = [
        "broken",
        "damaged",
        "again",
        "still waiting",
        "never arrived",
        "refund",
        "frustrated"
    ]

    matches = sum(
        1 for kw in distress_keywords
        if kw in messages_text.lower()
    )

    if matches >= 1 and contact_count >= 2:
        sentiment_trend = "declining"
        confidence = min(
            0.95,
            0.78 + (contact_count * 0.04)
        )
    else:
        sentiment_trend = "stable"
        confidence = 0.80

    # Evaluate risk level strictly
    # according to specification thresholds
    if (
        contact_count >= 3
        and sentiment_trend == "declining"
    ):
        risk_level = "escalate"

    elif (
        contact_count >= 2
        or sentiment_trend == "declining"
    ):
        risk_level = "watch"

    else:
        risk_level = "normal"

    assert not (
        risk_level == "escalate"
        and sentiment_trend != "declining"
    ), (
        "Escalate tier requires declining sentiment"
    )

    return RiskProfile(
        customer_id=customer_id,
        contact_count_this_issue=contact_count,
        sentiment_trend=sentiment_trend,
        confidence=confidence,
        risk_level=risk_level
    )
Enter fullscreen mode Exit fullscreen mode

The assertion is particularly important:

assert not (
    risk_level == "escalate"
    and sentiment_trend != "declining"
)
Enter fullscreen mode Exit fullscreen mode

It ensures that the Escalate tier cannot be assigned without declining sentiment.


3. Multi-Session Frustration Trajectory

The system also calculates how frustration changes across individual support sessions.

def compute_frustration_trajectory(
    self,
    customer: Customer
) -> FrustrationTrajectory:

    """Computes multi-session frustration progression
    across historical threads & live ticket."""

    all_sessions = []

    threads = customer.threads or []

    distress_keywords = [
        "broken",
        "damaged",
        "again",
        "still waiting",
        "never arrived",
        "refund",
        "frustrated"
    ]

    for idx, t in enumerate(threads, 1):

        text = " ".join(
            [m.text for m in t.messages]
        ).lower()

        matches = sum(
            1 for kw in distress_keywords
            if kw in text
        )

        score = min(
            98,
            max(
                20,
                25 + (idx * 18) + (matches * 12)
            )
        )

        level = (
            "Critical"
            if score >= 80
            else (
                "High"
                if score >= 60
                else "Medium"
            )
        )

        all_sessions.append(
            FrustrationSession(
                session_id=t.thread_id,
                session_label=f"Session #{idx}",
                frustration_score=score,
                frustration_level=level
            )
        )

    current_score = (
        all_sessions[-1].frustration_score
        if all_sessions
        else 25
    )

    overall_trend = (
        "increasing"
        if (
            len(all_sessions) >= 2
            and (
                current_score
                - all_sessions[0].frustration_score
                >= 15
            )
        )
        else "stable"
    )

    return FrustrationTrajectory(
        customer_id=customer.customer_id,
        overall_trend=overall_trend,
        current_frustration_score=current_score,
        current_frustration_level=(
            all_sessions[-1].frustration_level
            if all_sessions
            else "Low"
        ),
        sessions=all_sessions
    )
Enter fullscreen mode Exit fullscreen mode

The resulting trajectory provides a session-by-session view:

Session #1 ──► Medium
                 │
Session #2 ──► High
                 │
Session #3 ──► Critical
                 │
                 ▼
           Increasing Trend
Enter fullscreen mode Exit fullscreen mode

This gives the support representative a way to see how the customer's experience is changing, rather than only viewing the latest message.


4. Proactive Escalation Banner

When the risk profile reaches the escalate tier and memory is enabled, the React frontend displays an alert.

export default function EscalationBanner({
    riskProfile,
    memoryEnabled
}) {

    if (
        !memoryEnabled ||
        riskProfile?.risk_level !== 'escalate'
    ) {
        return null;
    }

    const contactCount =
        riskProfile?.contact_count_this_issue || 3;

    const confidence = Math.round(
        (riskProfile?.confidence || 0.88) * 100
    );

    return (
        <div className="bg-red-50 border-b border-red-200
            border-l-4 border-l-red-600 p-4
            flex items-start gap-3.5 text-red-950">

            <div className="p-2 rounded-lg bg-red-100
                text-red-700 mt-0.5">

                <ShieldAlert className="w-5 h-5 text-red-700" />

            </div>

            <div>

                <h4 className="text-xs font-bold
                    text-red-900 uppercase tracking-wider">

                    Proactive Escalation Alert

                    <span className="text-[11px]
                        font-semibold px-2 py-0.5 rounded-full
                        bg-red-100 text-red-800">

                        Confidence: {confidence}%

                    </span>

                </h4>

                <p className="text-xs text-red-900
                    mt-1.5 leading-relaxed">

                    Customer has contacted support
                    <strong>{contactCount} times</strong>
                    with a declining sentiment trajectory.

                    Prompt proactive manager intervention
                    or goodwill credit is strongly advised.

                </p>

            </div>
        </div>
    );
}
Enter fullscreen mode Exit fullscreen mode

The banner is intentionally shown only when:

Memory = ON
        +
Risk = ESCALATE
        ↓
Proactive Escalation Alert
Enter fullscreen mode Exit fullscreen mode

Memory OFF vs Memory ON

Consider Customer #4471.

The customer has already experienced two failed delivery attempts for an Echo Dot and opens a third ticket:

“Where is my package?”

Memory Comparison

Memory OFF — Stateless Mode

Without Hindsight reflections, the LLM treats the ticket as a standard initial inquiry:

“Hello Customer #4471, thanks for reaching out! Please provide your 17-digit Order ID and confirm your delivery address so I can check tracking for you.”

Result

The customer is asked to repeat information they have already provided.


Memory ON — Hindsight-Grounded Mode

With memory enabled, Hindsight reflection surfaces:

3 repeat contacts
       +
Declining sentiment
       ↓
risk_level = escalate
Enter fullscreen mode Exit fullscreen mode

The copilot displays the Proactive Escalation Banner and generates a prioritized response:

“Hello Customer #4471, I am very sorry to see this is your 3rd contact regarding your Echo Dot delivery (Order #302-8220-4471). I see carrier delivery failures occurred earlier this week. I have contacted carrier dispatch for priority morning redelivery and applied a $15 courtesy credit to your account.”

The architectural difference is that the second mode can use the customer's historical trajectory rather than treating the current ticket in isolation.


Usage Over Time

The dashboard provides visibility into memory operations and their usage over time, including operations such as:

  • Retain
  • Recall
  • Reflect
  • Memory/knowledge retrieval
  • Memory/knowledge refresh

This provides an operational view of how the memory layer is being used by the support copilot.


What I Learned

01 — Decouple Heavy Reflection Tasks

Running areflect asynchronously during ticket resolution keeps the active message path separate from background opinion generation.


02 — Enforce Strict Risk Invariants

Sentiment analysis alone can be noisy.

Requiring:

Contact Count ≥ 3
        AND
Declining Sentiment
Enter fullscreen mode Exit fullscreen mode

before reaching the escalation tier provides a stricter trigger.


03 — Combine Working Memory With Temporal Memory

In-memory pinned facts can provide immediate context for specific customer constraints.

Hindsight can maintain longer-term experiences and effort trajectories.

Together:

Working Memory
      +
Temporal Memory
      ↓
Richer Customer Context
Enter fullscreen mode Exit fullscreen mode

04 — Keep Support Representatives in Control

Escalation alerts and action recommendations should remain advisory.

The representative should retain explicit approval control over actions such as manager escalation or goodwill resolution.

This reduces rep cognitive load while keeping humans involved in consequential decisions.


The Complete Flow

The architecture can be summarized as:

Customer Message
       │
       ▼
FastAPI Backend
       │
       ├──────────────► Hindsight Recall
       │                       │
       │                       ▼
       │                 Customer History
       │
       ▼
Risk & Trajectory Engine
       │
       ▼
Groq LLM
       │
       ▼
Agent Response
       │
       ▼
Ticket Resolved
       │
       ▼
Async Hindsight Reflection
       │
       ▼
Updated Customer Opinion
       │
       ▼
Future Escalation Signal
Enter fullscreen mode Exit fullscreen mode

This creates a feedback loop:

Recall → Respond → Resolve → Reflect → Update Risk


Final Takeaway

Proactive support escalation is not simply about detecting negative words.

It requires understanding what happened across multiple interactions.

The architecture combines:

Scoped Memory
      +
Temporal Experiences
      +
Hindsight Reflection
      +
Risk Invariants
      +
Frustration Trajectory
      ↓
Proactive Support Signals
Enter fullscreen mode Exit fullscreen mode

The key idea is simple:

Don't wait for a customer to ask for a manager before recognizing that the support experience is deteriorating.

By combining persistent memory with reflection and explicit risk thresholds, the support copilot can surface relevant escalation signals earlier while keeping the final decision with the support representative.


Resources

Top comments (1)

Collapse
 
marcusykim profile image
Marcus Kim •

The risk tier logic that requires both 3 contacts and declining sentiment to escalate feels like a solid guard against false positives. I noticed the code uses contact_count >= 3 as the hard threshold for escalation, which prevents a single negative message from triggering a panic. The way it calculates frustration scores per session with 25 + (idx * 18) + (matches * 12) is clever-it weights session order and distress keywords to show how frustration compounds over time.