Building a unified catalog of online courses requires writing and maintaining separate web scrapers for every major platform. Coursera, edX, Udemy, and YouTube each use different page layouts, frontend frameworks, and anti-bot checks. When platforms alter their DOM or deployment pipelines, custom parsers break.
Class Central functions as a course meta-search engine, aggregating listings from over 100 course providers into standardized pages. Instead of writing platform-specific scrapers, you can extract course metadata, ratings, syllabi, and pricing across multiple providers using a single interface with the Class Central Scraper.
Standardizing cross-platform course schemas
Scraping course data directly from underlying providers creates schema fragmentation. A course workload might be defined as total hours on one platform, weeks of effort on another, or video length on a third. Class Central normalizes these disparate inputs into standard fields.
When pulling course listings, the Actor yields records containing consistent fields across platforms:
{
"courseId": "18207",
"title": "Intro to Programming Nanodegree",
"slug": "udacity-intro-to-programming-nanodegree--nd000-18207",
"courseUrl": "https://www.classcentral.com/course/udacity-intro-to-programming-nanodegree--nd000-18207",
"format": "credential",
"level": "beginner",
"language": "English",
"provider": "Udacity",
"rating": 4.5,
"numRatings": 120,
"isFree": false,
"hasCertificate": true,
"isUniversityCourse": false,
"pricingLabel": "Paid Course",
"effort": "109 hours",
"rank": 1,
"sourceMode": "search",
"recordType": "courseListing",
"scrapedAt": "2026-03-30T10:00:00.000Z"
}
If a field is missing from the source provider (such as an academic institution tag for a platform-native course), the Actor omits the key from the output JSON rather than populating it with null values or empty strings.
Target query routing with specific execution modes
The Actor supports six execution modes set via the mode parameter: search, bySubject, byProvider, byUniversity, collection, and courseDetail. Choosing the right mode prevents unnecessary page navigation and minimizes execution overhead.
Filtering by subject or provider taxonomy
To monitor course additions within a specific domain, use bySubject alongside predefined taxonomy slugs like cs, data-science, ai, or devops.
{
"mode": "bySubject",
"subjectSlug": "data-science",
"level": "intermediate",
"minRating": 4.0,
"sortBy": "highestRated",
"maxItems": 100
}
If you need to catalog courses from a specific provider or academic institution, set mode to byProvider or byUniversity using slugs such as coursera, mit, or stanford.
Fetching granular course details and user reviews
Listing modes omit deeper metadata like syllabi outlines and individual review texts. To capture these fields, pass target URLs to courseUrls with mode set to courseDetail.
{
"mode": "courseDetail",
"courseUrls": [
"https://www.classcentral.com/course/udacity-intro-to-programming-nanodegree--nd000-18207"
]
}
The detailed record appends structural attributes:
-
syllabus: Ordered list of syllabus topic titles -
subjectTags: Topic breadcrumbs associated with the course -
reviews: Up to 20 user reviews containingauthor,date,text, andrating -
institutionandinstitutionUrl: Academic affiliations (e.g., University of Michigan hosting on Coursera)
How to set up and run the Actor via API
You can trigger execution using the Apify Python SDK or standard HTTP requests.
Step 1: Install the client SDK
pip install apify-client
Step 2: Configure and trigger the run
The snippet below initializes the client, configures input filters to search for free, certificate-bearing Python courses, and fetches the resulting dataset items.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"mode": "search",
"searchQuery": "python",
"level": "beginner",
"freeOnly": True,
"certificateOnly": True,
"minRating": 4.5,
"maxItems": 50
}
run = client.actor("crawlerbros/class-central-scraper").call(run_input=run_input)
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(f"{item.get('title')} - {item.get('provider')} ({item.get('rating')} stars)")
Cost mechanics for event-based billing
This Actor uses a flat pay-per-event pricing model alongside platform usage. Platform usage is billed separately at the rates of your Apify plan.
The fixed event charges on the Apify Store are:
- Actor Start: $0.005 per GB of memory allocated to the run, charged once when the Actor starts running.
- Result Output: $0.005 per event emitted to the default dataset.
Emitting result items benefits from discount-tier pricing based on user tier:
- FREE: $0.005 per event
- BRONZE: $0.00433 per event
- SILVER: $0.00367 per event
- GOLD: $0.003 per event
- PLATINUM: $0.003 per event
- DIAMOND: $0.003 per event
For example, collecting 100 course results on a default run configuration with 1 GB of memory allocated incurs one Actor Start event ($0.005) and 100 result events ($0.50 at the standard FREE tier), plus underlying platform usage.
Technical constraints and anti-bot behavior
Direct HTTP requests to Class Central URLs using standard Python libraries like requests or httpx frequently yield 403 Forbidden errors due to browser fingerprinting checks. The Actor runs headless network operations that bypass these checks automatically, allowing output URLs to be logged cleanly.
However, courseDetail mode does not support full historical pagination for reviews. The output reviews array only captures the sample of reviews displayed directly on the primary course page (up to 20 items). If your workflow requires extracting thousands of legacy reviews for a single course, this Actor will not fulfill that specific requirement.
Class Central Scraper is what these steps drive. The README covers the inputs this article skipped, including the ones that change how much a run costs.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-04. Check the Actor page for the current rates.
Top comments (0)