DEV Community

Statewave
Statewave

Posted on Originally published at statewave.ai

Why our support health score drops 25 points on a quiet week

Why our support health score drops 25 points on a quiet week
I scored one support account once a day for eleven days and added nothing to it. No new tickets, no messages, no resolutions. It went from 80 to 55 and crossed a band boundary on the way down. Two of the eight signals in the score read the clock, and that is the whole explanation.

We build Statewave, an open-source memory runtime for AI support agents, so treat the framing as ours and the numbers as checkable. The scoring rules are plain arithmetic in the Apache-2.0 core, which is what makes this post possible: every number below can be re-derived from the published constants.

What this score measures, and what it leaves to your CSM tool

Worth separating before anything else, because the two answer different questions.

Customer success platforms score the whole relationship. Gainsight's Staircase AI health score combines sentiment, engagement, open items and response time, with weights set by a statistical model and lifecycle modifiers layered on top.

HubSpot's customer success workspace takes a different route: an admin builds a points-based score from event and property groups, with a point cap per group and optional decay.

Ours reads only the support record: open and resolved sessions, urgency in the conversation, idle tickets, response and resolution times. Narrower on purpose, because it matches the question the next agent has at escalation: how are this customer's tickets going right now? Predicting renewal is a different job and this number will not do it.

How the eight rules compute the number

Every customer starts at 100. Eight rules move the number, the result is clamped between 0 and 100, and a band is assigned: 70 and above is healthy, 40 to 69 is watch, under 40 is at_risk.

Signal Points Cap Fires when
Unresolved issues -15 per open session -45 A session's resolution status is open
Repeated issue -20 Once An open session shares at least 30% of its keywords with a resolved session
Escalations -10 per episode -20 A message contains an urgency marker such as outage, p0, sev1, or escalat
Idle open issue -15 Once An open session has had no activity for 7 days
Resolution breach -10 per session -20 A resolved session took more than 24 hours from the first customer message
Slow first response -5 Once Average first response across sessions is above 10 minutes
Recent resolution +10 Once A session was resolved in the last 7 days
High resolution rate +10 Once At least 80% of two or more sessions are resolved

No model call, no stored score, nothing cached in the scoring path. The same history at the same moment always returns the same number and the same factor list.

The factor list matters more than the number

Take a constructed mid-market account with three tickets. An SSO login loop fixed in two hours, three weeks ago. A duplicate-invoice export problem that took 36 hours to close last week. A warehouse sync to Snowflake failing since a schema change, untouched for seven days.

{
  "subject_id": "acct_ridgeline",
  "score": 65,
  "state": "watch",
  "factors": [
    { "signal": "unresolved_issues", "impact": -15, "detail": "1 open session(s)" },
    { "signal": "idle_open_issue", "impact": -15, "detail": "Session t-102 idle for 7 days" },
    { "signal": "recent_resolution", "impact": 10, "detail": "1 session(s) resolved in last 7 days" },
    { "signal": "sla_resolution_breaches", "impact": -10, "detail": "1 session(s) exceeded 24h resolution SLA" },
    { "signal": "slow_first_response", "impact": -5, "detail": "Avg first response 10.7 min (threshold: 10 min)" }
  ]
}
Enter fullscreen mode Exit fullscreen mode

A receiving agent does not have to trust 65. They can see that 30 of the 45 penalty points come from one stale ticket, and go straight to t-102.

Fixed rules beat learned weights for a number that lands in an escalation brief, for one reason. Nobody can argue with a learned weight at 2am. A -15 labeled Session t-102 idle for 7 days gets checked in ten seconds.

Why it drifts when nothing happens

Two signals read elapsed time. An idle penalty fires once an open ticket passes seven days without activity, and the recent-resolution bonus expires seven days after a ticket closes. Neither needs a new event to change the score.

Scored daily across eleven days with no new episodes and no new resolutions, that same account walks down on its own:

  • Day 0: 80, healthy
  • Day 3: the Snowflake ticket crosses seven idle days, 65
  • Day 7: last week's invoice fix ages out of the bonus window, 55 Urgency markers behave the opposite way. Nothing ages them out, because the escalation check counts matching episodes regardless of when they were written.

So an "outage" message from a ticket that closed months ago keeps costing points forever. Only the 20-point cap bounds how much.

Alerts fire only when something asks for the score

The subject.health_degraded and subject.health_improved webhooks compare the current band against the last band cached. That comparison runs only when the score is computed, which happens on GET /v1/subjects/{subject_id}/health and on POST /v1/handoff.

In the walkthrough above, the drop to watch on day 3 sends nothing at all until one of those calls runs. If you want band-change alerts on the day they happen, run a scheduled job that requests health for every subject with an open session.

One more detail worth knowing: a subject scored for the first time is compared against a default of healthy. A new subject whose first score lands in watch therefore fires health_degraded immediately.

What can actually reach at_risk

I enumerated every combination of signal values the rules allow, respecting dependencies such as a repeated issue requiring at least one open session, and scored each one against the published constants. Doing so produced 936 feasible combinations.

No single signal can do it. Largest penalty any one signal applies is 45, from three or more open sessions, which leaves a score of 55. Still watch. Every at_risk customer has at least two separate things going wrong.

Zero open tickets means a floor of 55. With no open sessions the unresolved, repeated-issue and idle signals cannot fire at all. Everything left adds up to 45 at most, and 66 of the 72 no-open-ticket combinations come out healthy.

The fastest route is three open tickets plus one full-strength signal. Three open sessions (-45) plus a repeated issue (-20), two urgency episodes (-20), or two resolution breaches (-20) lands on exactly 35. With only one open ticket, reaching at_risk takes at least three other signals firing together.

Read an at_risk alert as "this customer's queue needs a person today." Risks that live outside the queue, like a champion leaving or a renewal slipping, have to reach your team some other way.

Four ingest mistakes that move the score with no error

Each of these is a mismatch between how your ingest path records support data and the fields the rules actually read. None of them raise an error. I reproduced all four against the scoring rules before writing them down.

What you see Why it happens What to do
An outage ticket scores like a routine one The escalation and repeated-issue checks read payload.messages[].content. Text sent as payload.text is skipped. An open ticket reading "Full outage, P0" scored 85 as text and 75 as messages. Record support turns in the messages shape
A 48-hour resolution costs no points Timing checks recognize user, chat and support-chat as customer sources, and assistant, agent, system, tool as responders. The same history scored 100 with customer/staff and 95 with user/agent. Normalize source names at ingest
Billing questions count as escalations Urgency markers are substring matches, so "download" contains "down" Never alert on the escalations factor alone. The 20-point cap limits the damage
Scores run high right after a bulk import Idle and timing checks read created_at, which is ingest time. Backfilled tickets look answered in seconds and resolved today Backfill before go-live, or hold health alerts for seven days after an import

One failure we see more often than any of those is not in the table at all: tickets that never get closed. Open-session penalties stack toward 45 and the resolution bonuses never fire, so long-standing customers drift into watch for no reason other than bookkeeping.

What the handoff pack drops first

A score is one line in a larger brief. POST /v1/handoff assembles nine sections from stored memory, with no model call, then trims to a token budget: 4,000 by default, up to 16,000.

Trimming cuts from the end. Customer and health come first so they survive, and recent context from the customer's other sessions is the first thing lost. If your briefs are getting cut, raise max_tokens before you change anything else.

Three quirks worth knowing before you wire this into an escalation path:

  • The health line keeps the first three factors rather than the biggest three. Factors are added in the fixed order the rules run, so for the account above that means the open session, the idle ticket and the recent resolution. Its resolution breach and slow first response reach the next agent through the SLA section instead.
  • Active Issue is the first message in the session, cut to 200 characters. If the session opens with "hi", so does the brief. Record the problem statement, or a ticket-created event carrying the subject line, as the first episode.
  • What Has Been Tried keeps five entries, oldest first. On a long-running ticket the most recent attempts are the ones left out, so a specialist picking one up should pull the full session timeline for anything past step five. Because nothing is generated, the brief cannot invent a detail that was never recorded. Which cuts both ways. A model-written summary reads more smoothly, and it can also tell you something that never happened.

If you are building your own

Five things that transfer, whatever you build this on:

  1. Make every point traceable to a reason string. A score that goes into an escalation brief has to survive a skeptical human during a live incident.
  2. Decide deliberately which signals age. Two of ours do, and both surprised us once we watched them over a week: bonuses that expire, and urgency markers that never expire at all.
  3. Compute on demand and schedule the computation. If your alerts compare against a cached band, they fire when something asks, not when the state changes.
  4. Your ingest shape is what the score actually reads. Field names and payload shapes change the number silently, whatever you meant them to record. Normalize at ingest, and write a resolution when a ticket closes.
  5. Cap anything matched by substring. "download" contains "down", and a cap is what stops that from becoming an alert. Start with one account that has an open ticket. Call the health endpoint and check each factor against the ticket history. It takes about ten minutes and it tells you whether your ingest path records what the rules actually read.

Scoring rules, handoff assembly and the eval suite all live in the Apache-2.0 core runtime.

Read the constants there rather than taking our word for any of this.

Top comments (0)