DEV Community

Vladimir Elchinov for Session Replay

Posted on

Your Status Page Is Written in the Present Tense. Bug Reports Are Not.

For eight days in late September the network between Tokyo and Singapore was congested because subsea cables had been cut. Cloudflare's status page carried it from 23 September and closed it on the 29th. The cause, in their own words: "multiple subsea cable outages have caused congestion between Tokyo and Singapore datacenters", dating back to 21 September at 02:20 UTC.

No deploy caused it. It threw no errors. And it was more than enough to make a site feel slow across a large part of the world.

What that produces on your side is the worst report you can receive: "the site was slow yesterday". No error, no stack trace, no status code to search for. One region, some users, intermittently. Nothing in your own metrics either, because your monitoring almost certainly runs somewhere the cables were fine.

The asymmetry nobody designs around

A status page is written in the present tense. The front page answers exactly one question: is this broken right now.

A bug report is written in the past tense. It asks a different one: was this broken then.

So by the time somebody mentions the site was slow on Tuesday, the incident has closed and the dashboard says all systems operational. That is true, and useless, and it quietly reads as confirmation that the user imagined it. You close the ticket as not reproducible, and you are both slightly wrong.

The history is an API, and you already depend on several of them

The useful part is that the incident archive is machine readable, on a path that is the same almost everywhere, because most of these pages are the same hosted product rather than something each company built.

I checked five this morning. Every one answered with JSON on /api/v2/incidents.json: www.cloudflarestatus.com, www.githubstatus.com, status.openai.com, www.redditstatus.com and status.datadoghq.com.

Each incident record carries name, impact, status, started_at, resolved_at, the affected components, and the full incident_updates history with a timestamp on every update. Which means "was anything on fire around 14:30 yesterday" stops being a question you ask in Slack and becomes one you answer:

require "json"
require "net/http"
require "time"

def incidents_around(host, at, window: 3600)
  body = Net::HTTP.get(URI("https://#{host}/api/v2/incidents.json"))

  JSON.parse(body)["incidents"].select do |incident|
    times = [incident["started_at"], incident["resolved_at"], incident["created_at"]]
            .compact.map { |t| Time.parse(t) }
    next false if times.empty?

    times.min - window <= at && at <= times.max + window
  end
end

incidents_around("www.cloudflarestatus.com", Time.parse("2026-09-25T09:00:00Z"))
  .each { |i| puts "#{i['name']} (#{i['impact']})" }
Enter fullscreen mode Exit fullscreen mode

One thing that will bite you

Do not assume started_at comes before resolved_at.

Incidents get filed retroactively, when somebody notices afterwards that something was broken and writes it up. The record then carries the time it was written as its start. On GitHub's status page as I write there is an incident titled "[Retroactive] Actions workflow run failures after deployment gate approvals" whose started_at is 11:34 UTC today and whose resolved_at is 02:00 UTC the same day, nine and a half hours earlier.

A natural-looking started_at <= t && t <= resolved_at filter silently excludes exactly the incidents somebody cared enough about to go back and document. Take the minimum and maximum of whatever timestamps the record carries instead, which is what the snippet above does.

The precondition is the whole problem

All of this runs on two inputs: roughly when, and roughly where.

Neither of them comes out of a free-text box. "It was slow yesterday" has neither, and the conversation to get them is the one everybody has had. Which page? I don't remember. About when? Afternoon? A week later the cables are repaired, the report is closed as not reproducible, and the only lesson anybody took is that reporting things here is a waste of time.

So the cheap half of this is a script you write once and keep. The expensive half is arriving at a report that carries a timestamp and a location at all, which is a capture problem rather than an investigation one.

Write the script anyway. It is twenty lines and it will settle an argument every few months. It just cannot do anything with a report that has nothing to join on.

Top comments (0)