Before a migration, a content audit, or a crawl budget review you want the list of pages a site itself says it has. Crawling gets you there slowly and noisily. The sitemap gets you there in seconds, if you handle the format properly.
What a sitemap reader has to handle
-
Discovery:
robots.txtmay list severalSitemap:lines; if not, try/sitemap.xml,/sitemap_index.xml,/sitemap-index.xml,/wp-sitemap.xmland/sitemap.txt. - Indexes: large sites publish a sitemap index pointing at dozens of child sitemaps, sometimes nested.
-
Compression:
.xml.gzis common on big sites. -
Extensions:
xhtml:link rel="alternate" hreflang="..."for language versions,image:imageentries,lastmod,changefreq,priority. - Size: a single file can hold 50,000 URLs; an ecommerce site can have millions across files, so you need a cap and streaming parsing or you run out of memory.
The two-minute version
Sitemap URL Extractor does discovery, follows indexes recursively, decompresses gzip, parses the extensions, and returns one record per page URL.
- Paste website URLs or sitemap URLs into Website or sitemap URLs.
- Set Max URLs per site (default 5,000) and optionally Include / Exclude URL patterns such as
**/blog/**or*.pdf. - Click Start; export CSV, Excel or JSON, or feed the URL list into another Actor.
Switch on List sitemap files only to get one row per sitemap file (URL count, depth, lastmod) instead, which is the quickest way to map a large site's sitemap structure.
From code
curl -X POST "https://api.apify.com/v2/acts/josh99smith~sitemap-url-extractor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "urls": ["https://www.apify.com", "https://blog.apify.com/sitemap.xml"], "maxUrlsPerSite": 5000, "excludePatterns": ["*.pdf"] }'
What you get back
{
"site": "https://www.apify.com",
"sitemapUrl": "https://apify.com/sitemap.xml",
"url": "https://apify.com/store",
"lastmod": "2026-09-17T06:12:41.000Z",
"changefreq": "daily",
"priority": 0.8,
"alternates": [{ "hreflang": "en", "href": "https://apify.com/store" }],
"imageCount": 0
}
A site with no sitemap, or one hidden behind bot protection, comes back as a single free { "success": false, "errorType": "not-found" | "blocked" | "dns" } record.
Cost and limits
$0.0002 per extracted URL: a 5,000-URL site costs $1. Failed sites and unloadable sitemap files are free, there is no start fee, and both the per-site cap and the run's cost cap protect you from a surprise on a million-URL site. The Actor reads sitemap files only and never fetches the pages themselves.
Use it from an AI agent
Add https://mcp.apify.com?tools=josh99smith/sitemap-url-extractor to your MCP client and ask "list every URL under /blog/ on blog.apify.com".
Disclosure: I built this Actor. Source: github.com/josh99smith/sitemap-url-extractor.
Top comments (0)