DEV Community

Cover image for How to Scrape Spotify Artist, Album and Playlist Data Without a Paid API
Tim Zinin
Tim Zinin

Posted on Originally published at apify.com

How to Scrape Spotify Artist, Album and Playlist Data Without a Paid API

The problem

Music analytics, A&R research, and playlist campaign tracking all start with the same humble need: read the public card for a Spotify artist, album, or playlist — the name, the monthly-listener figure, the release year, the track count, the item and save counts. Doing that by hand means opening every page in a browser and re-typing numbers into a spreadsheet, which does not scale past a handful of targets.

The programmatic alternatives are poor fits for a small watchlist. The Spotify Web API requires an app, OAuth tokens, and approval, and it meters you into quotas you do not need for a dozen pages. A naive scraper has a worse problem, and it is counterintuitive: Spotify's card pages only render their structured data server-side for a plain, non-browser request. Send a real browser User-Agent and you get HTTP 200 with an empty client-side shell — about 156 KB of "Spotify – Web Player" with zero data in it. A confident-looking response containing nothing.

What the actor does

The Spotify Artist, Album & Playlist Scraper takes full Spotify URLs you already have — https://open.spotify.com/artist/<id>, /album/<id>, or /playlist/<id> — and reads the server-rendered application/ld+json block each page ships, with one plain non-browser GET per target. No login, no API key, no Web Playback SDK, no browser runtime.

From the README:

  • The non-browser rule is the product. The actor always sends its own declared, non-browser User-Agent — never a browser signature — because that is the only request shape Spotify answers with data. Every row carries a requestUserAgent field echoing the literal header sent, so you can verify the discipline on your own runs; an acceptance test fails the build if that header ever becomes a browser string.
  • Type comes from the URL path, never guessed: artist rows return name and Spotify's displayed monthly-listener figure; album rows return name, release type, release year, and track count; playlist rows return name, item count, and save count.
  • One flat 15-field row per target, regardless of outcome: input, found, type, spotifyId, canonicalUrl, name, releaseType, releaseYear, trackCount, monthlyListeners, itemCount, savesCount, error, requestUserAgent, checkedAt.
  • Honest negatives, free of charge. A nonexistent object comes back as Spotify's own genuine HTTP 404 (found:false, error:"http 404"), not a silent-empty page. A non-Spotify host is refused before any request is made (requestUserAgent:null proves nothing was sent). An HTTP 200 with no usable ld+json block is reported, never guessed at. Only found:true rows are billed.
  • Disclosed source instability. Playlist itemCount/savesCount can legitimately come back null on the same URL across repeated fetches — the README measured three of five back-to-back fetches carrying the structured counts and two not. The actor reports the null honestly instead of inventing a value.
  • Up to 50 targets per run, concurrency 1–15 (default 5), plus a one-time KVS OUTPUT run summary with requested/delivered/paid/free/failed counts and replay safety.

Limits are stated plainly: no search by name, no track lists, no audio, no full discography, no pagination of playlist contents — the card-level summary only. monthlyListeners and savesCount are Spotify's own rounded display text ("100.9M"), not exact internal numbers.

Example: input and output

Input is one required field:

{
  "targets": ["https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02"],
  "maxConcurrency": 5
}
Enter fullscreen mode Exit fullscreen mode

The README's real happy-path row for that artist URL, verbatim:

{
  "input": "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
  "found": true,
  "type": "artist",
  "spotifyId": "06HL4z0CvFAxyc27GXpf02",
  "name": "Taylor Swift",
  "monthlyListeners": "100.9M",
  "error": null,
  "requestUserAgent": "Mozilla/5.0 (compatible; spotify-artist-intel/0.1; +https://apify.com/zinin/spotify-artist-intel)",
  "checkedAt": "2026-08-17T20:32:57.777Z"
}
Enter fullscreen mode Exit fullscreen mode

An album row in the same run returned name: "After Hours", releaseType: "album", releaseYear: 2020, trackCount: 14; a playlist row returned name: "Today's Top Hits", itemCount: 50, savesCount: "33.9M" (with a repeat fetch showing both counts as honest nulls).

Pricing and the free limit

Pay-per-event: $0.005 per actor start plus $0.002 per delivered card (result-found); the README computes 100 delivered cards at about $0.205 including the start fee. A run of 50 URLs where 10 are stale links is billed for the 40 it actually delivered, never for the 10 honest 404s. Apify's free plan includes $5 of usage credits per month, which covers about 24 full 100-card runs — roughly 2,500 delivered cards — before you pay anything.

Try it

Run the prefilled artist URL with zero edits to see a real row, then feed your own watchlist: Spotify Artist, Album & Playlist Scraper.

For AI agents and MCP

The actor takes JSON in and returns structured JSON via the Apify API, and the README documents an agent/MCP pattern plus the Apify MCP server setup. An agent that already holds a Spotify URL passes it into targets, branches on found before trusting any field, uses requestUserAgent to distinguish a locally blocked host from a genuine source 404, and treats a playlist row with null counts as "not available on this fetch" — a disclosed source behavior to re-run, not a defect to retry blindly.

Top comments (0)