Originally published on the Djangix blog: Building SafetyGraph: 546,891 Industrial Incident Records in One Queryable Dataset
Public industrial safety records in the US are rich and, historically, painful to use: the same underlying information has been spread across separate agency systems, each with its own search form and export quirks, and much of it reachable only one lookup at a time. SafetyGraph, the project described in the full article, brings those records together into one queryable dataset of more than half a million incidents.
The interesting engineering is less about scraping pages than about identity. Records describing the same company appear under spelling variants, punctuation differences, and changed names, so the pipeline normalises names, pairs likely matches conservatively, and builds explicit relationships between establishments, inspections, violations, and accidents instead of leaving users with a pile of disconnected rows.
With that structure in place, questions that once meant hours of manual cross-referencing become a single query: show a company's history across sources, or follow the chain from an inspection to what it found and what happened next. The dataset is positioned for safety teams, researchers, journalists, and builders who need that history in a form software can actually use.
The full article explains where the records come from, how the matching and linking work, what the resulting data model looks like, and what you can ask it on day one.
Full article: Building SafetyGraph: 546,891 Industrial Incident Records in One Queryable Dataset
Top comments (0)