title: A 40% Faster Fault-Detection Pipeline Without WebSockets — the Polling Trick We Shipped
published: false
description: Real-time monitoring for 1,000+ network nodes over plain REST — a change-flag endpoint, RTK Query cache discipline, and an operator map that shows the causal chain, not the noise.
tags: react, rtkquery, architecture
This is the dev.to version of a story I first published as a LinkedIn article — restructured for a code-first audience. It's a real production system: real-time monitoring for 1,000+ GSM service stations, built for an international partner. And one honest constraint at its core: we could not touch the backend protocol.
The ticket-queue blind spot
The worst part of a monitoring outage isn't the failure itself. It's not being able to tell the difference between two situations that look identical on a ticket queue:
- a node recovered by itself, and
- a node is still broken, but nobody noticed yet.
Our network engineers — a field team plus a back office — lived inside that ambiguity for years. The back office created tickets from Grafana data, and a ticket meant wait: someone opened it, someone processed it, and a certain amount of time always passed before the incident became visible to people who could act on it.
The constraint that reframed everything
I checked the obvious answers first: whether the Romanian backend team could support WebSockets or SSE.
They couldn't. The backend was a monolith, and a deep refactor to introduce a new transport protocol was off the table. Not "hard" — off the table. Different project, different budget, different risk.
If the client can't be pushed updates, then maybe it can be woken up cheaply.
The reframe: change the data granularity, not the protocol
My proposal split the information flow into two layers:
1. A cheap change-flag. The backend exposes a tiny flag/version field on an ordinary REST endpoint — a number that increments whenever anything relevant changes underneath. We designed this together: they implemented a small flags feature to my spec, I tracked it from the frontend.
2. A heavy data pull. The actual node states — statuses, incidents, telemetry the dashboard renders.
The frontend polls only the flag on a light interval. When the flag hasn't changed, the request costs nothing meaningful. When it has — RTK Query re-fetches the heavy data and its cache is immediately invalidated and repopulated. The dashboard behaves as if it were pushed to, while under the hood everything is still plain REST.
That's the whole trick:
naive polling: [timer] → pull full payload → render (every N seconds, always)
flag-driven polling: [timer] → pull 1 number ─┬─ unchanged → nothing
└─ changed → invalidate → pull full payload → render
The RTK Query side: cache and refetch discipline
Two things came out of the RTK Query migration itself:
- Caching. Identical requests were served from cache instead of re-hitting the API — removing redundant traffic the legacy thunk-based store generated on its own.
- Refetch discipline. Refetches happened on demand — flag changed, cache invalidated, window refocused — not on a dumb interval burning everyone's patience and the backend's CPU.
The polling setup in practice:
// poll the flag only — light interval, e.g. 5s
const flag = usePollChangeFlag(intervalMs = 5000)
// the heavy query is keyed on the flag value:
// flag moves → new cache key → automatic refetch + repopulate
const nodes = useGetNodesQuery({ version: flag.version })
The exact API shape was refactored a few times since — the pattern is the deliverable, not the signature. The heavy-fetch key follows the flag; refetch happens when and only when the flag moves.
Showing pain, not data: the map
Raw node counts are useless to a human operator. 1,000+ stations on one screen is noise.
The stations had real topology: parent–child hierarchy. A parent going down didn't just break itself — it degraded its children. So we drew the map differently:
- the map shows only the affected subtree, not the whole network;
- when a parent fails, its dependent children light up as potentially damaged;
- in a typical impact window the map displayed around 400 nodes — the ones potentially hurt through connectivity — instead of 1,000+.
An engineer opens the map and sees a causal chain, not a cloud of dots. And the back office, which had been hand-writing tickets out of Grafana, now pulled the incident structure — parent, children, order of failure — from a single source.
The numbers
We validated the improvement with two existing instruments: Grafana metrics and the ticket system itself (opening-to-resolution timestamps, across the before/after periods).
Result: incident detection time improved by ~40%.
Being precise about what the number means — never defend a fuzzy metric in an interview: the clock we shortened is "failure happens → failure is actionable" — previously dominated by ticket latency, now dominated by flag-poll latency. Nodes that self-recovered no longer polluted the queue and no longer hid the ones that didn't.
If you can't change the protocol, change the granularity
You don't always need WebSockets. You need the cheapest possible signal that something changed — a version number, a hash, an etag — and a client disciplined enough to pull the real payload only when that signal moves. Pair it with cache-first architecture (RTK Query, React Query — whatever your stack provides) so heavy data is fetched rarely and reused often.
And when you present it to a human operator: don't show them the network. Show them the pain — the affected subtree, the causal chain. A map of 400 hurting nodes is worth more than a map of 1,000 quiet ones.
Built with React, TypeScript, RTK Query — and with backend partners who were straightforward about what the system could and couldn't give us.
I write about React patterns and the engineering behind my AI tooling — building in public on Telegram, the LinkedIn version of this story is here.
How do you do "real-time" when the protocol won't budge — flag-polling, SSE workarounds, or plain polling? Curious what others landed on.
Top comments (0)