DEV Community

Cover image for Why the Eventbrite Events Scraper Drops Records Past the 500 Item Cap
Crawler Bros
Crawler Bros

Posted on

Why the Eventbrite Events Scraper Drops Records Past the 500 Item Cap

Extracting event data from Eventbrite presents unique architectural hurdles that go far beyond basic HTTP requests. While scraping simplifies data extraction by bypassing login requirements, developers frequently run into empty datasets, silent payload truncations, or unexpected integration failures. These issues rarely stem from code bugs. Instead, they are the direct result of conflicts between the Actor's input schema and the underlying Apify platform mechanics. To build reliable data pipelines, backend engineers must understand how input parameters interact, how the platform handles long-running synchronous requests, and how the event-based pricing model scales.

Why does an empty startUrls array return zero results from the Eventbrite Events Scraper?

An empty startUrls array explicitly overrides the default location and category filters, causing the Actor to return zero results. The scraper treats startUrls as the primary entry point, so any defined array (even an empty one) prevents the system from falling back to alternative inputs. To resolve this, you must entirely omit the startUrls field from your API payload when querying by location or category.

The input schema of the Eventbrite Events Scraper allows you to target events using direct browse URLs or localized search slugs. When you trigger the run from an automated API call, your application might generate an input dictionary dynamically. If your internal logic initializes optional list parameters with an empty list [] instead of omitting the key, the Apify platform will pass this empty array directly to the Actor container.

When the scraper boots, its entry-point controller checks for the presence of the startUrls parameter. If the key exists, even as an empty array, the engine assumes you have provided an explicit list of URLs to process. It bypasses the fallback parameters location and category entirely. Because the target list contains zero items, the scraper exits immediately, producing a successful run but writing zero items to your dataset.

To prevent this silent failure mode, your integration code must defensively strip out empty startUrls fields before dispatching the payload to the Apify API. If your upstream user interface submits a blank text field or an empty array, sanitize it programmatically.

import apify_client

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

def build_safe_input(raw_input: dict) -> dict:
    # If startUrls is present but empty, remove it to allow location/category to work
    if "startUrls" in raw_input and not raw_input["startUrls"]:
        cleaned_input = raw_input.copy()
        del cleaned_input["startUrls"]
        return cleaned_input
    return raw_input

# This raw input would fail silently and return zero results
raw_user_input = {
    "startUrls": [],
    "location": "ny--new-york",
    "category": "music",
    "maxItems": 50
}

safe_input = build_safe_input(raw_user_input)

# Execute the run with the cleaned input parameters
run = client.actor("crawlerbros/eventbrite-scraper").call(run_input=safe_input)
print(f"Started run {run['id']} with sanitized input parameters.")
Enter fullscreen mode Exit fullscreen mode

What causes the Eventbrite Events Scraper to truncate event results at five hundred items?

The scraper enforces an internal hard limit of 500 items per run, clamping any higher maxItems value without throwing an error. This pagination safety cap prevents excessive resource consumption and potential anti-bot triggers. To collect more than 500 events from a broad target region, you must execute multiple separate runs segmented by specific category slugs rather than relying on a single large query.

The default value for maxItems is 50, and the absolute maximum is 500. The scraper processes results by walking through Eventbrite's pagination (approximately 20 events per page) using the ?page=N parameter until it hits your maxItems limit or runs out of results. If you specify 1000 items, the platform will execute the run successfully, but the output dataset will never contain more than 500 items.

If your workflow expects a larger volume, you must detect when this cap is reached and split your queries. For instance, rather than querying the entire city of New York without filters, you can trigger individual runs for categories like music, business, and food-and-drink. This avoids hitting the single-run limit and maximizes your total output across multiple datasets.

import apify_client

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

requested_max = 800
target_location = "ca--san-francisco"

run_input = {
    "location": target_location,
    "maxItems": requested_max
}

run = client.actor("crawlerbros/eventbrite-scraper").call(run_input=run_input)
dataset_id = run["defaultDatasetId"]
items = list(client.dataset(dataset_id).iterate_items())

print(f"Requested: {requested_max} | Retrieved: {len(items)}")

# Detect if the internal hard cap truncated your dataset
if len(items) == 500 and requested_max > 500:
    print("WARNING: Dataset was capped at the 500-item limit.")
    print("To retrieve more data, segment your queries using specific category slugs.")
Enter fullscreen mode Exit fullscreen mode

Why do synchronous API requests to the Eventbrite Events Scraper fail past three hundred seconds?

Synchronous HTTP requests on the Apify platform are subject to a strict 300-second execution cap, resulting in an HTTP 408 timeout error if a run takes longer than five minutes. While the scraper container continues to run in the background on the platform, your calling application loses its connection and cannot capture the dataset directly. For larger queries approaching the 500-item cap, you must initiate the Actor asynchronously and poll the run status or use webhooks.

If you are scraping a high volume of events (such as the maximum limit of 500), network latency, pagination loops, and anti-bot mitigation can easily push the execution time beyond 5 minutes. Initiating this run via a synchronous .call() block or a direct POST request to the sync endpoint will cause your client application to drop the connection with a timeout error.

To avoid this, you must run the Actor asynchronously by initiating the run without blocking, then polling the run status or leveraging webhooks. This asynchronous execution model guarantees that your client application remains stable regardless of how long the scrapers take to gather the paginated data.

import apify_client
import time

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

# Trigger an asynchronous run by calling start() instead of call()
run_summary = client.actor("crawlerbros/eventbrite-scraper").start(
    run_input={"location": "united-states", "maxItems": 500}
)

run_id = run_summary["id"]
print(f"Asynchronous run {run_id} started. Polling status safely...")

while True:
    status_details = client.run(run_id).get()
    current_status = status_details["status"]

    if current_status in ["SUCCEEDED", "FAILED", "ABORTED"]:
        print(f"Run completed with final status: {current_status}")
        break

    print(f"Current status is {current_status}. Waiting 15 seconds...")
    time.sleep(15)
Enter fullscreen mode Exit fullscreen mode

Handling typed default values instead of nulls in downstream databases

To parse the scraped event data safely, your parser must inspect fields for empty strings, zero values, and empty arrays instead of checking for nulls. The Actor schema guarantees that all 29 output fields are present in every record, replacing missing values with explicit typed defaults.

For example, when an event is virtual (isOnline is true), fields like venueName, venueAddress, and venuePostalCode will return as empty strings "", while the geocoordinates latitude and longitude will return as 0. If you downstream this directly into a database that expects a null or uses a strict coordinate validator, it may trigger database integrity constraint errors.

Below is a Python function designed to clean up these typed default values and restore native None values where appropriate before pushing data into database systems.

def sanitize_eventbrite_record(record: dict) -> dict:
    cleaned = record.copy()

    # Clean string fields that use empty strings as typed defaults
    string_fields = ["venueName", "venueAddress", "venueCity", "venueRegion", "venueCountry", "venuePostalCode"]
    for field in string_fields:
        if cleaned.get(field) == "":
            cleaned[field] = None

    # Clean coordinate fields where 0 indicates a missing venue
    if cleaned.get("latitude") == 0 and cleaned.get("longitude") == 0:
        cleaned["latitude"] = None
        cleaned["longitude"] = None

    # Clean array fields
    array_fields = ["tags", "gallery"]
    for field in array_fields:
        if cleaned.get(field) == []:
            cleaned[field] = None

    return cleaned

# Example output record from the scraper containing typed defaults
scraped_item = {
    "id": "1029384756",
    "name": "Virtual Tech Meetup",
    "isOnline": True,
    "venueName": "",
    "latitude": 0,
    "longitude": 0,
    "tags": []
}

sanitized_item = sanitize_eventbrite_record(scraped_item)
print("Sanitized Output:", sanitized_item)
Enter fullscreen mode Exit fullscreen mode

Avoiding configuration failures caused by Console UI prefill values

The prefill configurations specified in the Apify Console input schema are UI helpers and are never injected into direct API requests or task invocations. If your API call relies on those visual placeholders, the Actor will run with its default schema values or trigger validation errors.

When developing integrations, it is easy to assume that because a field is filled out in the Console UI, it will automatically apply when calling client.actor("crawlerbros/eventbrite-scraper").call(). This is false. Only fields marked as default in the schema are fallback-compatible. If you do not provide an explicit input payload in your API calls, the Actor will use its internal default values (e.g., maxItems defaults to 50), completely ignoring whatever was configured in the UI prefill fields.

Always construct and pass an explicit input payload dictionary when triggering the scraper from code.

{
    "startUrls": ["https://www.eventbrite.com/d/ny--new-york/all-events/"],
    "location": "ny--new-york",
    "category": "music",
    "maxItems": 100
}
Enter fullscreen mode Exit fullscreen mode

Protecting scraped Eventbrite datasets from platform storage expiration

Unresolved storage endpoints and unnamed datasets automatically expire on the platform, and free tier plans cap storage retention at the 10 most recent runs. If your integration relies on pulling historical data from old run IDs, those datasets will disappear after a maximum of 4 months.

When you execute the scraper, Apify generates an unnamed dataset. To prevent your historical event data from being deleted, you must assign an explicit name to your storage entities. Named storages are always exempt from deletion schedules, guaranteeing your extracted files remain accessible over time for analytical workflows.

The following script shows how to rename your dataset programmatically immediately after a successful scraping run.

import apify_client

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

# Trigger the scraper run
run = client.actor("crawlerbros/eventbrite-scraper").call(
    run_input={"location": "online", "category": "business", "maxItems": 20}
)

dataset_id = run["defaultDatasetId"]

# Apply a name to the dataset to exempt it from the auto-deletion policy
permanent_name = f"saved-eventbrite-online-{run['id']}"
client.dataset(dataset_id).update(name=permanent_name)

print(f"Dataset has been renamed to '{permanent_name}' and is safe from expiration.")
Enter fullscreen mode Exit fullscreen mode

Orchestrating parallel queries safely with isolated request queues

When scaling your scraper architecture to monitor multiple metropolitan areas or distinct event categories, you might be tempted to queue all locations into a single automated run. However, the Apify platform introduces a strict constraint: a request queue can only be processed by one Actor or task run at a time. Trying to execute parallel fan-out operations or running multiple concurrent instances of the scraper against a single, shared request queue will result in synchronization failures and skipped URLs.

To scrape multiple locations concurrently, you must isolate each query within its own distinct run execution. This approach prevents resource conflicts and allows you to sidestep the 500-item hard limit by segmenting your scraper tasks across multiple locations.

The following Python integration demonstrates how to orchestrate isolated parallel runs for different metropolitan areas safely.

import apify_client
from concurrent.futures import ThreadPoolExecutor

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

# List of target locations to process in parallel runs
target_locations = ["ny--new-york", "ca--san-francisco", "tx--austin"]

def run_scraper_for_location(location_slug: str):
    print(f"Initiating run for: {location_slug}")
    try:
        # Each run is completely isolated with its own unique input parameter payload
        run = client.actor("crawlerbros/eventbrite-scraper").call(
            run_input={
                "location": location_slug,
                "category": "business",
                "maxItems": 100
            }
        )
        print(f"Successfully completed run {run['id']} for {location_slug}.")
        return run["defaultDatasetId"]
    except Exception as e:
        print(f"Run failed for {location_slug}: {str(e)}")
        return None

# Execute runs concurrently using Python's ThreadPoolExecutor
with ThreadPoolExecutor(max_workers=3) as executor:
    dataset_ids = list(executor.map(run_scraper_for_location, target_locations))

print("All location scraping tasks have completed. Dataset IDs generated:", dataset_ids)
Enter fullscreen mode Exit fullscreen mode

Operational limitations and failure modes of the scraper

Every data extraction system is bound by specific rules, and this scraper is no exception. Understanding these boundaries ensures you do not design processes that violate core design assumptions.

First, the scraper relies on extracting the embedded window.__SERVER_DATA__ JSON directly from public Eventbrite browse pages. Because it does not require proxies, cookies, or user logins, it works directly from datacenter IPs via Chrome 131 TLS impersonation. However, this means that if Eventbrite changes the layout or name of this embedded JSON script block, the scraper engine will fail to locate the event data, returning empty datasets until the scraper is updated by its developer.

Second, the output is strictly restricted to a flat schema containing exactly 29 fields. If Eventbrite introduces new custom metadata fields or regional telemetry on their event detail pages, those fields cannot be retrieved through the standard scraper interface. The extracted dataset will only populate the standard 29 preconfigured fields.

Third, because the request queue is tied directly to the execution context of a single run, you cannot configure dynamic feed inputs that depend on sharing state in real-time across multiple instances. Any segmentation of your search (by location or category slugs) must be managed on your backend by spinning up isolated runs.

Finally, there is no direct query parameter to filter events by specific start and end dates or by ticket pricing levels within the input schema. The scraper retrieves all available live events from the browse results up to your configured maxItems limit. Any refined sorting, date range windowing, or price filtration must occur downstream inside your database after the raw items are successfully captured.

Cost structure and calculation model for running the scraper

Understanding the pricing model of the Eventbrite Events Scraper is critical for budgeting your data pipelines. This Actor uses the PAY_PER_EVENT pricing model. This means you are charged a set flat rate for every specific system event executed during a run, independently of any subscription plan rates. On top of these event fees, you also pay for the standard Apify platform usage that the run consumes, which is billed separately at your Apify plan's rates.

The event-based pricing structure contains two primary components:

  • "Actor Start" (apify-actor-start): Billed once per run execution when the container initializes. This is charged at a flat price of $0.005 per GB of memory allocated to the run. For instance, allocating more memory to the container increases this baseline flat charge per run.
  • "result" (apify-default-dataset-item): Charged for every single result written to the default dataset. This base rate scales directly with the number of events your run emits. The price per result event is determined by your assigned Apify user tier:
    • FREE: $0.002 per event
    • BRONZE: $0.00167 per event
    • SILVER: $0.00133 per event
    • GOLD: $0.001 per event
    • PLATINUM: $0.001 per event
    • DIAMOND: $0.001 per event

Your total event-based cost scales dynamically with the size of your query. When fetching data, your maxItems value acts as the primary cost multiplier, as it determines how many result items are scraped and written to your dataset. Additionally, allocating more container memory will increase your "Actor Start" event fee. Platform usage costs (representing processing resources, storage operations, and network bandwidth) are charged dynamically and are displayed separately on your Apify billing dashboard.

Checked against the Actor's input schema and Apify docs on 2026-10-05.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

Top comments (0)