DEV Community

Timothy Kelvin
Timothy Kelvin

Posted on

Replacing GovTrack's retired bill API with official bulk data

I maintain a small tool that looks up new U.S. federal bills by keyword. For its whole life it read from GovTrack's API (https://www.govtrack.us/api/v2/bill). Last week a cloud run failed with a 15-second timeout, requests from my own machine got empty replies, and another HTTP client got a 403. A day later the same URL answered again.

Before adding another retry and moving on, I checked what I was actually depending on.

The API had been retired for years

  • In December 2016 GovTrack announced that its API and bulk data would shut down the following summer, and pointed developers to other sources.
  • The API documentation page, /developers/api, returns a 404. The Wayback Machine shows it as a 404 since at least February 2023.
  • robots.txt has Disallow: /api for every user agent, with a 30-second crawl delay.

The endpoint still answers, but nothing about it is supported. An intermittent failure there isn't something to retry around, so I looked for a replacement.

Two official options

The Congress.gov API. It's official and well documented, but it needs an api.data.gov key. The shared DEMO_KEY came back with X-Ratelimit-Limit: 10, which is nothing for a tool running from shared cloud IPs. The /bill list endpoint also has no keyword parameter, and list items don't include the sponsor, so every bill would need its own request.

GovInfo's BILLSTATUS bulk data. The Government Publishing Office publishes one XML file per bill at https://www.govinfo.gov/bulkdata/BILLSTATUS/{congress}/{type}/, built from Congress.gov data. Each file has the titles, sponsor, cosponsors, committees and the full action history. No key, no account, the files are marked public domain, and govinfo.gov's own robots.txt lists the bulk data sitemaps.

I went with the bulk data. It has no search either, but there's no request budget to protect.

Finding recent bills without a search endpoint

Each Congress and bill type has its own sitemap:

https://www.govinfo.gov/sitemap/bulkdata/BILLSTATUS/119hr/sitemap.xml
Enter fullscreen mode Exit fullscreen mode

It lists every file (10,608 House bills in the 119th Congress when I checked). From where I am it downloaded in under 3 seconds, while the JSON directory listing for the same folder took 9 to 13.

async function listBillNumbers(congress, type) {
  const url = `https://www.govinfo.gov/sitemap/bulkdata/BILLSTATUS/${congress}${type}/sitemap.xml`;
  const xml = await (await fetch(url)).text();
  const re = new RegExp(`BILLSTATUS-${congress}${type}(\\d+)\\.xml`, 'g');
  const numbers = new Set([...xml.matchAll(re)].map((m) => Number(m[1])));
  return [...numbers].sort((a, b) => b - a);
}
Enter fullscreen mode Exit fullscreen mode

Bill numbers are handed out in the order bills are introduced. So for "bills from the last 90 days" I walk each type (H.R., S., H.Res. and so on) from the highest number down and stop once three bills in a row are older than the window. Three rather than one, as a margin in case a number was introduced out of order.

Status isn't in the file

The XML has no "current status" field. It has a list of actions, newest first, with free-text descriptions:

<item>
  <actionDate>2026-09-24</actionDate>
  <text>Referred to the House Committee on Oversight and Government Reform.</text>
  <type>IntroReferral</type>
  <actionCode>H11100</actionCode>
</item>
Enter fullscreen mode Exit fullscreen mode

GovTrack gets its data through the open source unitedstates/congress project, and statuses like "Passed House (Senate next)", "Failed Cloture" or "Pocket Vetoed" come from that project's parsing of these action texts. The rules live in parse_bill_action and new_status_after_vote in bill_info.py, about 400 lines of Python and regular expressions. I ported them to JavaScript rather than inventing my own, so existing status filters would keep meaning the same thing.

The core is a small state machine. After a vote, the new status depends on whether it passed, which chamber voted, and whether that's the chamber the bill started in. Simplified:

function statusAfterVote(voteType, passed, chamber, billType, suspension) {
  if (voteType === 'vote') { // vote in the originating chamber
    if (passed) {
      if (billType === 'hres' || billType === 'sres') return 'PASSED:SIMPLERES';
      return chamber === 'h' ? 'PASS_OVER:HOUSE' : 'PASS_OVER:SENATE';
    }
    if (suspension) return 'PROV_KILL:SUSPENSIONFAILED';
    return chamber === 'h' ? 'FAIL:ORIGINATING:HOUSE' : 'FAIL:ORIGINATING:SENATE';
  }
  // ...second-chamber, ping-pong, cloture, veto override and conference votes
}
Enter fullscreen mode Exit fullscreen mode

Two details mattered:

  • The same vote often appears twice: once from the House or Senate floor system and once from the Library of Congress (sourceSystem code 9), with slightly different text. The original code drops the Library of Congress copy when its text ends with the text of the action just before it. Skip that and a single vote can count twice.
  • Actions are replayed oldest first, and each one sees the status the earlier ones produced.

Checking the port against the old source

While GovTrack still answered, I pulled bills from it across as many statuses as I could and compared field by field: status, which chamber acts next, status date, introduced date, whether the bill is still alive, and the sponsor's party and state. That was 203 bills over 18 different statuses, including failed veto overrides from earlier Congresses. All 203 matched.

That check is why I'm comfortable with the switch. A port of that many regular expressions is easy to get subtly wrong.

What got worse: keyword search

GovTrack's search looked at more than titles. For "artificial intelligence" over 90 days it returned 123 bills. Matching titles in the bulk data finds 57, all of which GovTrack also returned. I checked eight of the missing ones: the phrase appears nowhere in their BILLSTATUS files, and none had a CRS summary or legislative subject terms yet, including bills introduced in July.

Two changes won part of it back:

  • OR between alternatives. artificial intelligence OR AI finds 68, and 66 of those are in GovTrack's set. The other two mention AI in their titles but weren't in GovTrack's results.
  • Singular and plural both match, so "data centers" also finds the "Data Center Fair Share Act".

Real full-text search would mean downloading bill text as well, at least one more file per bill. I've left that out for now.

Keeping it cheap

A 90-day keyword search reads about 2,000 XML files, and most of them get thrown away. A full parse took about 2.5 ms per file on my machine, so each file goes through two stages:

const xml = await res.text();
// the bill's own introducedDate comes before any nested sections
const introduced = xml.match(/<introducedDate>([^<]+)<\/introducedDate>/)?.[1];
if (introduced < cutoff) return 'too-old';

// every word of the keyword has to be somewhere in the raw file
const lower = xml.toLowerCase();
if (!tokens.every((t) => lower.includes(t))) return 'no-match';

const bill = parser.parse(xml).billStatus.bill; // full parse only for candidates
Enter fullscreen mode Exit fullscreen mode

The raw-text check can let extra files through but never drops a real match, because a title match means every word is in the file somewhere. The exact decision still runs on the parsed titles afterwards.

In a cloud container in us-east-1, 2,030 files took about seven seconds at 10 requests at a time. GovInfo sits behind Cloudflare, and plain requests at that rate went through without any challenge.

Takeaways

  • An endpoint that answers is not the same as a supported API. Check the docs page and robots.txt before building on one, and again when it starts failing.
  • When an API has no search anyway, bulk files plus a sitemap can be simpler than an API with a tiny rate limit.
  • If you replace a data source, run old and new side by side on real data before switching. 203 matching bills told me more than rereading the regular expressions ever would.

Top comments (0)