Building a media monitoring database or an automated outreach engine requires fresh podcast data. The problem is that open podcast indexes are filled with abandoned projects. Millions of feeds exist across public directories where the creator published three episodes in 2021 and never posted again.
If you ingest raw search results without filtering up front, your downstream pipeline wastes processing cycles cleaning dead feeds, checking unreachable RSS URLs, and evaluating abandoned shows. Querying directories separately also forces you to maintain distinct wrappers for the iTunes Search API and Listen Notes REST API.
The Apple Podcasts + Listen Notes Scraper combines both data sources into a single schema. It includes built-in date filtering to eliminate inactive shows before they reach your storage.
The Dual-Directory Pipeline Problem
Apple Podcasts and Listen Notes serve different structural purposes in data ingestion:
- Apple Podcasts exposes public iTunes search and country-specific RSS chart endpoints without requiring an API key. It updates its top charts roughly every 1 to 2 hours across 47 country storefronts and 19 standard genres.
-
Listen Notes maintains an editorial database with structured social links, granular search types (
podcast,episode,curated), and reach metrics such aslistenScoreandlistenScoreGlobalRank. It requires an API key, providing 250 free calls per month on its base tier.
Each platform uses different identifiers and pagination behaviors. Apple identifies shows via numeric collectionId values, while Listen Notes uses 32-character hexadecimal hashes. Normalizing these records in custom scripts requires mapping varying date formats, categorizing genres, and parsing nested metadata.
{
"platform": "applePodcasts",
"recordType": "podcast",
"podcastId": "1200361736",
"title": "The Daily",
"publisher": "The New York Times",
"url": "https://podcasts.apple.com/us/podcast/the-daily/id1200361736?uo=4",
"primaryGenre": "Daily News",
"artworkUrl": "https://is1-ssl.mzstatic.com/image/.../600x600bb.jpg",
"summary": "This is what the news should sound like...",
"scrapedAt": "2025-01-01T00:00:00+00:00"
}
Dropping Dead Feeds with minLastEpisodeDays
When you query directories for broad topics like "technology" or "marketing", a substantial portion of the returned records have not published an episode in months or years.
The actor supports an input parameter called minLastEpisodeDays. Setting this field to an integer such as 30 evaluates the show's most recent release timestamp and immediately drops any podcast whose latest episode is older than 30 days.
This approach saves database space and prevents post-processing workers from querying dead RSS feeds found in the feedUrl field.
For Apple Podcasts queries, the actor handles public iTunes API constraints internally. The iTunes Search API throttles around 20 requests per minute per IP address. The actor sleeps approximately 0.6 seconds between calls and honors upstream Retry-After headers whenever it receives HTTP 429 status codes.
Step-by-Step Walkthrough
You can execute a search across active podcasts using either the Apify Console or the Python API client.
1. Set Up the Run Configuration
Construct an input configuration specifying the target directory, search term, and activity threshold.
{
"platform": "applePodcasts",
"mode": "searchPodcasts",
"searchQuery": "data engineering",
"country": "us",
"minLastEpisodeDays": 30,
"maxItems": 50
}
2. Trigger the Actor via Python
Run the Actor through the apify-client package and pull clean records straight into your data store:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"platform": "applePodcasts",
"mode": "searchPodcasts",
"searchQuery": "data engineering",
"country": "us",
"minLastEpisodeDays": 30,
"maxItems": 50,
}
run = client.actor("crawlerbros/applepodcasts-listennotes-scraper").call(run_input=run_input)
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(f"{item['title']} - {item.get('feedUrl')}")
3. Switch to Listen Notes for Deep Metadata
If you need metrics like listenScore, update the platform parameter, provide your key in listenNotesApiKey, and change the mode accordingly:
{
"platform": "listenNotes",
"mode": "bestPodcasts",
"listenNotesApiKey": "YOUR_LISTEN_NOTES_KEY",
"listenNotesGenreId": 127,
"listenNotesRegion": "us",
"maxItems": 25
}
The returned records will include social media links, update frequency estimates, and language flags alongside standard show descriptions.
Pricing and Event Charges
The actor runs on a pay-per-event pricing model. In addition to standard Apify platform usage (which is billed separately according to your plan's rates), runs incur charges based on two specific event types:
-
Actor Start (
apify-actor-start): $0.005 per GB of memory allocated to the run, charged once when the run begins. -
Result (
apify-default-dataset-item): Charged for each emitted record saved to the default dataset.
The per-result event base price is $0.005 on the FREE tier. The price decreases for users on higher Apify discount tiers:
- BRONZE: $0.00433 per result
- SILVER: $0.00367 per result
- GOLD: $0.003 per result
- PLATINUM: $0.003 per result
- DIAMOND: $0.003 per result
Filtering out stale records using minLastEpisodeDays directly limits the emitted records, which helps control result event charges.
What This Workflow Does Not Do
This setup does not capture Spotify-exclusive podcasts, as private catalog shows that do not expose public RSS feeds are omitted entirely from both Apple Podcasts and Listen Notes directories.
For pipelines targeting open RSS-distributed podcasts, standardizing ingestion around these two indexes provides broad global coverage while filtering out dead feeds directly at the edge.
Runs in this article used Apple Podcasts + Listen Notes Scraper. Its README is the reference for input fields and output structure; this post is only one path through them.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-04. Check the Actor page for the current rates.
Top comments (0)