DEV Community

Cover image for Get every URL of a website from its sitemaps, without crawling it
Joshua Smith
Joshua Smith

Posted on

Get every URL of a website from its sitemaps, without crawling it

Before a migration, a content audit, or a crawl budget review you want the list of pages a site itself says it has. Crawling gets you there slowly and noisily. The sitemap gets you there in seconds, if you handle the format properly.

What a sitemap reader has to handle

  • Discovery: robots.txt may list several Sitemap: lines; if not, try /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml and /sitemap.txt.
  • Indexes: large sites publish a sitemap index pointing at dozens of child sitemaps, sometimes nested.
  • Compression: .xml.gz is common on big sites.
  • Extensions: xhtml:link rel="alternate" hreflang="..." for language versions, image:image entries, lastmod, changefreq, priority.
  • Size: a single file can hold 50,000 URLs; an ecommerce site can have millions across files, so you need a cap and streaming parsing or you run out of memory.

The two-minute version

Sitemap URL Extractor does discovery, follows indexes recursively, decompresses gzip, parses the extensions, and returns one record per page URL.

  1. Paste website URLs or sitemap URLs into Website or sitemap URLs.
  2. Set Max URLs per site (default 5,000) and optionally Include / Exclude URL patterns such as **/blog/** or *.pdf.
  3. Click Start; export CSV, Excel or JSON, or feed the URL list into another Actor.

Switch on List sitemap files only to get one row per sitemap file (URL count, depth, lastmod) instead, which is the quickest way to map a large site's sitemap structure.

From code

curl -X POST "https://api.apify.com/v2/acts/josh99smith~sitemap-url-extractor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://www.apify.com", "https://blog.apify.com/sitemap.xml"], "maxUrlsPerSite": 5000, "excludePatterns": ["*.pdf"] }'
Enter fullscreen mode Exit fullscreen mode

What you get back

{
    "site": "https://www.apify.com",
    "sitemapUrl": "https://apify.com/sitemap.xml",
    "url": "https://apify.com/store",
    "lastmod": "2026-09-17T06:12:41.000Z",
    "changefreq": "daily",
    "priority": 0.8,
    "alternates": [{ "hreflang": "en", "href": "https://apify.com/store" }],
    "imageCount": 0
}
Enter fullscreen mode Exit fullscreen mode

A site with no sitemap, or one hidden behind bot protection, comes back as a single free { "success": false, "errorType": "not-found" | "blocked" | "dns" } record.

Cost and limits

$0.0002 per extracted URL: a 5,000-URL site costs $1. Failed sites and unloadable sitemap files are free, there is no start fee, and both the per-site cap and the run's cost cap protect you from a surprise on a million-URL site. The Actor reads sitemap files only and never fetches the pages themselves.

Use it from an AI agent

Add https://mcp.apify.com?tools=josh99smith/sitemap-url-extractor to your MCP client and ask "list every URL under /blog/ on blog.apify.com".

Disclosure: I built this Actor. Source: github.com/josh99smith/sitemap-url-extractor.

Top comments (0)