Every directory of indie software I have used had the same flaw. A third of the links were dead, and nothing told you which third. Listings were added once, by someone excited about them, and never looked at again.
So I built one that checks. It Still Works is a catalogue of independent software: web tools, games, open-source projects. Every published listing is fetched on a six-hour schedule and the result is written down. Right now that is 672 listings and 40,817 recorded checks. Each listing page shows whether the app answered at the last check, when it last did, how fast, when its visible content last changed, and when its TLS certificate expires. An app that fails three checks in a row moves to a public graveyard, and the graveyard stays, because knowing what used to work is half the point.
This post is about what the record taught me, because the checking turned out to be the easy part.
What a check is
A check is an HTTP fetch with a 20-second budget. The verdict is live on a 2xx under five seconds with no parking-page text in the body, slow above that, error on anything else. We also hash the visible text so a redesign registers and a rotating advert does not, read the certificate's expiry, and note where the URL redirected to. Three consecutive non-working checks promote the listing to dead.
The sweep runs in three phases: check everything, decide what to record, then write. The split matters for the next section. Requests are paced per host, so forty games on one platform are fetched one at a time with a gap, not in a burst.
The schedule that did not run
The first version relied on an external cron. It never fired on the host I deployed to, and nobody noticed until the checks were eight days stale while the site still said "every six hours". That is the worst kind of bug for a product whose only asset is being right.
The fix was to move the scheduler into the app process. Each job takes a lease in a Postgres table, writes a heartbeat while it runs, and records how it ended. The next run is scheduled from when the last one finished, not from wall-clock, so a slow night does not pile runs on top of each other. A run that stops heartbeating is recovered as timed out. The public evidence page shows the last full sweep and how many of the scheduled sweeps in the window actually completed, so the claim is checkable from outside.
The night the record was wrong
A few weeks in, itch.io started answering our server with 403 on every page. Forty working games were recorded as failing, twice. One more sweep and they would all have been marked dead.
The checks were correct as observations. The server really was refused. They were wrong as facts about the games. The fix had two parts.
First, a rule in the decide phase. If a host has at least five listings and at least 80% of them failed in a way that looks like refusal, the sweep records nothing for that host that night. A real 404 or an unresolvable domain still counts. A 429 is never recorded for anyone.
Second, a repair that re-applied that rule to every recorded run, removed the checks the fixed sweep would have declined to write, along with the false "content changed" rows and queued emails they caused, and rebuilt each listing's current state from the checks that remained. It runs from an admin page as a dry run first.
For itch.io games specifically, we now check the game's own build files, which the platform serves from a host that does answer us, rather than the store page. That is the better signal anyway: a store page can stay up after the developer deletes the upload.
Honesty as an engineering constraint
The site has a short list of rules that read like policy but behave like tests. Never display a number we cannot defend as measured. Label detected facts as detected. Never say "working" about software we never loaded.
That last one changed the titles. For an installed CLI or desktop tool, the only thing our server can reach is the project's website, so its page now asks "Is X still available?" and says "Site checked", while a web app's page asks "Is X still working?". A repository's page asks "still maintained?" and answers from commit dates GitHub reports, with the date of the snapshot it rests on.
Two bugs taught me that "honest" is not a wording problem. The sitemap's last-modified query compared a string with a Date, which is always false in JavaScript, so recorded changes never moved it and 601 of 672 URLs told crawlers nothing had changed since a screenshot run in September. And the sweep compared each redirect with the stored URL instead of the previous observation, so every listing whose URL permanently redirects wrote an identical "redirect" row every six hours: 3,370 rows across 60 listings before I noticed. Both made the site say things that were not true, and no copy review would have caught either.
What I would ask you
If you make something, you can claim your listing and put the status badge in your README. It updates on every check and links to the verification record.
What I want to know from people who read this far: what would you need to see on a listing page before you trusted it, and what would you expect to find there that is missing?
Top comments (1)
tr.ee/dev-to