If you have ever pointed an AI agent at "public data", you know the hard part is not calling the API. It is finding the right official one. Which agency publishes company records in Denmark? Does Poland's VAT register have a real endpoint or only a web form? Is there an API behind a statistics office, or just PDFs? Agents burn their first steps guessing, and they often end up at a paid reseller instead of the government source.
So I built a catalog for exactly that step: otwarteAPI.pl, an open index of official public and government APIs that an agent can read directly.
Disclosure up front: it is my own project. It is free, there is no login and no tracking, and it does not proxy or resell anyone's data. Every entry points straight at the official source.
What is in it
Right now it lists 92 APIs across 40 jurisdictions: 38 countries and territories, plus EU institutions and international bodies. Company registers, VAT and tax lookups, statistics offices, central banks and open data portals. A few examples: Poland's VAT white list, KRS and CEIDG, the UK's Companies House, Denmark's CVR, Norway's Brønnøysund register, France's Sirene, SEC EDGAR, FRED, Eurostat, the ECB data warehouse, GLEIF, the World Bank, and the national data portals of Singapore, India and Australia.
69 of them need no key at all. Each entry names the publisher (a ministry, a registrar, a statistics office, a central bank), gives the base URL, says whether you need a key, and lists a few example questions it can answer. Everything is in English and Polish, and 38 entries have a test button on their page that calls the official API straight from your browser.
The part that matters for agents
The whole catalog is one machine-readable file:
curl https://otwarteapi.pl/.well-known/ai-catalog.json
It uses the ai-catalog.json format from the Agentic Resource Discovery (ARD) spec, which a working group with people from Microsoft, Google and Hugging Face is developing. The same file drives the search on the site, so the page a human reads and the data an agent reads can't drift apart. There is also an llms.txt, and JSON-LD on every page for crawlers that don't speak ARD yet.
Every entry has a stable identifier (a URN), the URL, a description, tags and a jurisdiction code (an ISO country code, EU for Union institutions, INT for international bodies), so an agent can filter by country or topic before it makes a single call.
If your client speaks MCP, there is also a remote MCP server at https://otwarteapi.pl/mcp with one tool that searches the catalog. It is listed in the official MCP registry as io.github.bartosz-kuc/otwarteapi.
How you would use it
Say your agent has to check a company somewhere in Europe. Instead of hardcoding one country's endpoint, it reads the catalog, filters by jurisdiction and the company register tag, and gets the official base URL for that country. One lookup, then a direct call to the government source. No scraping, no guessing, no middleman.
It is open
The catalog data is mirrored to a public repo, bartosz-kuc/otwarteapi-catalog. Issues and pull requests are welcome there: a missing official API from your country, a base URL that moved, a key requirement that changed. I review the changes and fold them back into the site.
If you build agents that touch public data, I would really like to hear which registers or countries you wish were in there.
Top comments (2)
The single source for the human page and machine catalog is useful. For agent consumers, I'd add a last-checked date and a link to the publisher's docs for each entry's key requirement, separate from whether the endpoint currently responds. "Official source" and "still usable without a key" can change independently.
A useful regression fixture would keep the same URN while an API moves its base URL or starts requiring a key. Can a client see that change in the catalog rather than treating a failed call as "no company found"? That distinction seems especially important for register lookups.
Does each entry say how a key is obtained, and not only whether one is needed? For an agent, "free key by email" and "needs a registered company" are completely different blockers, and that's the field I'd filter on first. Mirroring the data to a public repo is the right move, government endpoints change without much notice.