DEV Community

Cover image for Where Hospital Price Files Actually Live in 2026
Pennyforge
Pennyforge

Posted on

Where Hospital Price Files Actually Live in 2026

A dated census of US hospital machine-readable price files: formats, hosting, and file sizes — and the federal catalog that used to list them all, now gone.

Pennyforge (SendCheck) — 2026-10-02. CC BY 4.0.


If you build anything on top of US hospital prices — negotiated-rate datasets, cost estimators, agent pipelines — you already know the rule: every inpatient hospital must post a machine-readable file (MRF) of its charges, discounts, and negotiated rates, plus a plain-text pointer file in its web root. For years, CMS ran the Provider Data Catalog (PDC): a single machine-readable list of every hospital's MRF URL. It was the thing you pointed a crawler at.

It's gone. And the gap it leaves is worth mapping. This post is the map, with dates and sources.

1. The federal catalog is gone (8 dated evidence points)

As of 2026-10-02, the PDC no longer appears anywhere in CMS's own machine-readable surface, after the data.cms.gov platform rebuild published 2026-09-29:

  1. The new slug API returns 404 for the old PDC path (/data-api/v1/slug?path=/data/pdc/252m-zfp9 → "doesn't match anything in the system").
  2. The rebuilt DCAT-US catalog (data.json, 3,422 datasets + 159 series, published 09-29) contains exactly one price-titled dataset: enforcement activities.
  3. The new site's 648-URL sitemap has no PDC entry.
  4. The legacy SODA endpoint returns 410 Gone.
  5. The HPT program page now links only the enforcement dataset and the GitHub docs repo.
  6. The HPT FAQ (as of 2026-06-26) never mentions the catalog.
  7. The GitHub README carries no PDC URL.
  8. We scanned all 6,583 distribution files in the new DCAT catalog: exactly one is a price file — the July-2026 enforcement CSV.

The de-facto catalogs have moved to third parties: a public GitHub tracker covering all 5,419 CMS-listed hospitals (anthonyisnotadev/cms-hpt-tracker, per-row checked 2026-09-29, repo updated 10-02), vendor trackers (Turquoise's MRF tracker reports 77.3% of hospitals meet both requirements, refreshed 10-01), and an open monitoring project (mrf-watch). Fine — but dated, attributed, and not federal.

2. Format diversity: the answer to "what formats are hospitals allowed to post in?"

CMS's own HPT FAQ invites the question. Counting the 3,983 assessable hospitals (federal/DoD/IHP excluded) with an MRF URL in the tracker's 09-29 check:

Format (by listed URL) Count Share
.csv 2,016 50.6%
.json 706 17.7%
Dynamic handlers (.aspx 417, .ashx 101, extensionless 385) 903 22.7%
.zip 347 8.7%
Other (.xlsx 2, .php 5, .com 2, .txt 1, .xml 1) 11 0.3%

So only about 68% of hospitals serve a plain, statically-named file. Nearly a quarter are URLs that generate the file on demand — which is what a crawler gets.

Our live spot check (25 hospitals, seeded random sample, 2026-10-02, residential egress) shows the listed distribution undercounts CSV: four of the extensionless URLs we probed were Azure Blob Storage endpoints serving plain CSV with an application/octet-stream content type. Recounted by served content, CSV is closer to two-thirds of the population.

Template versions: 3.0.0 dominates — 3,334 of the versioned listings (83.7%) — with ~36 already on 4.0.0, ~200 still on 2.x, and a long tail of malformed values (a data-quality note for anyone joining on this).

3. Hosting: about one in five files comes from five platforms

The 3,977 distinct MRF URLs spread over 1,196 domains are far less independent than "5,000 hospitals each host a file":

  • Top 5 domains serve 21.8% of all listed MRFs (a Para Healthcare FS app, two ST Health Azure blob buckets, Hyve Healthcare's MRF host, Craneware's pricing API).
  • Azure Blob buckets alone carry 11.7% of the population.
  • Top 10 domains: 29.7%.

If you're building a price-data pipeline, your real integration surface is a handful of hospital-IT vendors — not 5,000 independent web roots. (Vendors' rate limits, auth quirks, and outages will shape your pipeline more than individual hospitals do.)

4. File sizes: one in five sampled files is over 100 MB

Of the 25 MRFs we fetched live, 5 (20%) exceeded 100 MB — 112–120 MB CSVs and a 106 MB JSON. The smallest was ~120 KB. If your pipeline does "download the file, parse, store", budget for >100 MB as the common case for large health systems, not the edge case.

Liveness, from the same 25: 17 returned 200 immediately; one 302 redirect; one 403 (Cloudflare blocking data-center IPs — a real trap for crawlers); 6 timed out from our egress; one hospital had removed its root pointer file (404). Pointer files that resolved returned the canonical location-name: / source-page-url: / mrf-url: body.

What this means for builders

  • Point at a dated, cited catalog, not "CMS". The federal list is absent as of 10-02 (evidence above); the third-party rebuilds are the working surface, and they drift.
  • Don't trust the extension. ~23% of URLs are dynamic; extensionless cloud URLs frequently serve CSV. Sniff the first bytes.
  • Treat five hospital-IT platforms as first-class upstreams. They carry a fifth of the population.
  • Budget >100 MB files as routine.

Method, sources, caveats

  • File list: anthonyisnotadev/cms-hpt-tracker (public GitHub; 5,419 hospitals; per-row checked_at 2026-09-29; repo updated 2026-10-02). Cited as a de-facto PDC rebuild; not an official CMS source.
  • Our verification: 25-hospital stratified random sample (seed 42) probed 2026-10-02 from a residential IP; results above are that spot check, not a full census. Full 5,419-row pass is next.
  • Incumbent context: Turquoise MRF tracker (77.3% compliance, 10-01 refresh) measures compliance; this post measures format/hosting/size diversity — the census it doesn't do.
  • Counts are as-of-dated; re-run the query before you build on them.

CC BY 4.0. AI-assisted drafting; all counts machine-verified. Questions → pennyforge.xyz.

Top comments (0)