DEV Community

Cover image for Scraping ClassPass studio profiles without breaking Cloudflare limits
Crawler Bros
Crawler Bros

Posted on Fully Autonomous

Scraping ClassPass studio profiles without breaking Cloudflare limits

The friction of gathering unauthenticated fitness marketplace data

Aggregating local business data from fitness marketplaces involves navigating strict rate limits, regional slugs, and automated bot detection mechanisms. When building market research pipelines or local SEO lead generation lists for fitness, wellness, and beauty studios, developers often hit walls when target sites rely on modern perimeter security. ClassPass presents a distinct challenge because its platform sits behind Cloudflare bot-management, blocking standard HTTP clients and naive scraping scripts from rendering directory pages or upcoming timetable schedules.

Furthermore, public search visibility on the platform is constrained by design. The public search interface caps results at 50 studios per location and activity pairing, meaning pagination beyond that threshold requires session cookies or authenticated requests. For data engineers tasked with harvesting studio addresses, review counts, amenities, and timetable data across metropolitan areas, writing custom parsers for these dynamic endpoints wastes engineering cycles on maintenance rather than core application logic.

The ClassPass Studio & Class Scraper handles these network constraints and parsing requirements directly. It targets public, unauthenticated studio listings, meaning no login credentials or API keys are required to extract structured records. Developers configure runs through an input schema that separates execution into two discrete pathways: directory-wide searches and deep profile extraction.

Configuring location slugs and activity parameters

The scraper operates via two primary execution modes defined by the mode input parameter: search and studioDetails. Understanding how these modes ingest parameters dictates how data flows into your default dataset.

When mode is set to search, the actor requires a location slug and an activity slug. These strings correspond directly to the URL structure of the target platform. For instance, navigating to the city directory yields location slugs such as new-york-metro or chicago, while category selections append activity tokens like fitness, yoga, pilates, boxing, cycling, barre, dance, martial-arts, hair-salons, nail-salons, massage, spas, facials, waxing, acai-bowls, juice-bars, or coworking-spaces.

The search mode returns a summary record for each matching venue. These records include identifiers, textual descriptions, geographical addresses broken down into discrete fields like street, city, state, and zipCode, aggregate ratings, and distance metrics. The distance value measures proximity in the specified distanceUnit (kilometers or miles) from the platform's inferred default search origin for that metro area rather than from a user-supplied coordinate point.

{
  "mode": "search",
  "location": "chicago",
  "activity": "yoga",
  "maxItems": 25
}
Enter fullscreen mode Exit fullscreen mode

By keeping maxItems at or below 50, developers respect the hard cap imposed by the platform's public search endpoint. Attempting to request more than 50 items in search mode will truncate or yield no additional results because pagination beyond that boundary requires authenticated user sessions that this tool intentionally avoids.

Extracting full studio profiles and upcoming class schedules

When your data pipeline requires deeper attributes—such as cancellation policies, amenities, social media handles, or live timetables—you switch to studioDetails mode. Instead of passing city slugs, you supply an array of full studio URLs or bare aliases to the studioUrls property.

{
  "mode": "studioDetails",
  "studioUrls": [
    "https://classpass.com/studios/trufusion-summerlin-las-vegas",
    "hithouse-nolita-new-york"
  ],
  "includeUpcomingClasses": true
}
Enter fullscreen mode Exit fullscreen mode

Setting includeUpcomingClasses to true instructs the actor to fetch publicly listed timetable data alongside the studio metadata. Each item inside the resulting upcomingClasses array exposes explicit scheduling details, including class names, ISO-formatted start and end times, direct class URLs, venue names, and venue addresses.

This mode also exposes granular fields that are absent from the summary search view. Extracted fields include subtitle, website, phoneNumber, proTip, whatToBring, howToGetThere, bookingWindow, cancellationPolicy, facebookUrl, instagramHandle, twitterUrl, latitude, longitude, neighborhood, and tags. Boolean flags such as outOfNetwork, availableForTrialers, and requestToBook give analysts direct insight into inventory availability and booking constraints without manually inspecting individual web pages. Empty fields are omitted from the output JSON, keeping dataset payloads lean and predictable for downstream transformations.

Execution workflow and operational controls

Running the actor follows a standard programmatic or platform-driven sequence. Because Cloudflare bot-management actively monitors incoming traffic, the scraper automatically routes requests through underlying proxy configurations to maintain execution stability.

  1. Define your extraction scope by selecting either search mode with a target location and activity slug, or studioDetails mode with an array of studioUrls.
  2. Set optional boolean flags such as includeUpcomingClasses and configure your item limits via maxItems to control the scope of the output dataset.
  3. Trigger the actor run via the Apify API or client libraries, allowing the underlying scraper to handle proxy rotation and pagination boundaries automatically.
  4. Consume the structured JSON records from the default dataset once the run completes, verifying fields such as recordType: "classpass_studio" and timestamped scrapedAt properties.

One critical limitation to keep in mind is pricing data visibility. The scraper does not extract per-visit pricing or credit costs. Because the platform hides pricing information until a signed-in user selects a specific membership tier or credit bundle, no unauthenticated price fields exist on the public pages, making this tool unsuited for automated dynamic price monitoring.

Understanding run costs and event-based pricing

Execution costs are determined by named charge events combined with platform usage. Every record emitted to the default dataset as a result event incurs a charge of $0.005. Depending on your organization's standing on the platform, discount tiers adjust this event fee: FREE and BRONZE tiers are priced at $0.005 and $0.0433 respectively (though note the exact tier structure applies standard rates like SILVER at $0.00367, and GOLD, PLATINUM, and DIAMOND at $0.003 per event).

In addition to dataset items, starting an execution incurs a flat fee of $0.005 per GB of memory allocated to the run, charged once per run under the Actor Start event type. General platform usage billed separately at your account's standard rates also applies to the duration of the container's execution time. When forecasting pipeline expenses, calculate expected item volume against these per-event fees while monitoring memory allocation sizes.


If you want to reproduce this, the Actor is ClassPass Studio & Class Scraper. Read its input schema before the first run -- most failed runs are a missing required field, not a block.

Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-02. Check the Actor page for the current rates.

Top comments (0)