As data engineers, we often encounter tools that promise straightforward data extraction. The Yelp Scraper (available on the Apify Store: https://apify.com/crawlerbros/yelp-scraper) is one such tool, designed to pull business data, reviews, and ratings from Yelp. While its input and output schemas seem clear, the implicit failure modes, the ways it can return incomplete or misleading data without an explicit error, are often where the real work begins. This article will focus on those less-obvious constraints and how to defensively code against them.
When does runTimeoutSecs lead to silently incomplete data?
The runTimeoutSecs parameter (default 1800 seconds or 30 minutes) acts as a wall-clock cap on the run. When this limit is reached, the Actor stops new fetches, pushes everything it has collected so far, and exits cleanly. This means your dataset will contain whatever was collected up to that point, which might be a partial result set if the scraper was in the middle of processing a large searchLimit or many directUrls. Crucially, this is not an error state from the Actor's perspective.
This mechanism is particularly relevant when you're using the synchronous run API. The synchronous run endpoint hard-caps at 300 seconds and returns an HTTP 408 (Request Timeout) past that. This means if your client makes a synchronous request and the Actor run on the platform exceeds 300 seconds, the client's connection will terminate with an HTTP 408, and it will not receive any further data for that specific request. However, the Actor itself will continue to run on the Apify platform until its runTimeoutSecs is met. For runs expected to last longer than 5 minutes, you must use the asynchronous run endpoint and poll for its status or configure a webhook.
Consider this scenario:
import apify_client
from datetime import datetime, timedelta
import time
client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
actor_id = "crawlerbros/yelp-scraper"
run_input = {
"searchTerms": ["coffee shop"],
"locations": ["London, UK"],
"searchLimit": 100, # A high limit to challenge the timeout
"runTimeoutSecs": 400 # Just over the 300s synchronous cap
}
try:
print(f"Starting Actor {actor_id} with runTimeoutSecs: {run_input['runTimeoutSecs']}")
# This call would typically be client.actor(actor_id).call(run_input=run_input)
# The Apify Python client's .call() method might internally poll for longer than 300 seconds
# if runTimeoutSecs allows, handling the asynchronous nature.
# However, if using a raw HTTP client for a synchronous call, a 300s hard cap is real
# and would return HTTP 408 to the client after 300 seconds.
run = client.actor(actor_id).call(run_input=run_input)
# If the run finishes within the synchronous client's timeout, you get results.
# If it exceeds 300s, an HTTP client would typically throw a timeout error,
# but the run would still be happening on Apify until its runTimeoutSecs is met.
print(f"Run completed with status: {run['status']}")
print(f"Access dataset items via: {client.dataset(run['defaultDatasetId']).url}")
except Exception as e:
print(f"An error occurred during the run call, possibly a client-side timeout: {e}")
# In a real synchronous HTTP client timeout, you'd get an error,
# but the run would still be happening on Apify. You'd need to
# find its dataset ID to retrieve any collected data.
print("Always use asynchronous execution for runs expected to exceed 300 seconds.")
print("Checked against the Actor's input schema and Apify docs on 2026-10-06")
In this case, runTimeoutSecs of 400 seconds combined with a synchronous API call might lead to a TimeoutError on the client side if not handled by the client library, because the synchronous run endpoint would return HTTP 408 after 300 seconds. The Actor on the platform would continue for its full 400 seconds, pushing data to the default dataset, but your client would never receive a direct response containing the dataset ID for that specific request. Always design your long-running data pipelines to use asynchronous execution and robust status checking.
How do I handle missing review data from reviewLimit?
The reviewLimit input parameter controls the maximum number of reviews to collect per business. However, reviews are extracted only from the reviews visible on the first page load of a business listing. Even if you set reviewLimit up to 50, only reviews rendered on the page will be captured. If a business has fewer reviews visible on the first page than your reviewLimit, or if reviewLimit is set to 0, the output for reviews will contain fewer items than expected or be an empty array.
This behavior implies that even if you set reviewLimit: 50, you might only get a handful of reviews, or even fewer if the business itself has only a few. The reviews array in the output will reflect what was actually collected. To detect this defensively, compare the length of the reviews array for each business against your reviewLimit. If len(business['reviews']) < review_limit (and review_limit > 0), it might indicate either a natural scarcity of reviews for that business or that the scraper couldn't load more than the first page.
Here's a Python snippet to check for this:
import apify_client
# Initialize Apify client
client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
# Example run configuration
actor_id = "crawlerbros/yelp-scraper"
run_input = {
"searchTerms": ["pizza"],
"locations": ["New York, NY"],
"reviewLimit": 10, # Requesting up to 10 reviews
"searchLimit": 1
}
# Start the Actor and wait for it to finish
run = client.actor(actor_id).call(run_input=run_input)
# Fetch results from the default dataset
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if "reviews" in item and run_input["reviewLimit"] > 0:
actual_reviews = len(item["reviews"])
expected_reviews = run_input["reviewLimit"]
if actual_reviews < expected_reviews:
print(f"Business '{item['name']}' returned {actual_reviews} reviews, less than the requested {expected_reviews}.")
# You might want to log this, trigger an alert, or re-process
elif actual_reviews == 0 and run_input["reviewLimit"] > 0:
print(f"Business '{item['name']}' returned no reviews despite reviewLimit > 0.")
elif "reviews" not in item:
print(f"Business '{item['name']}' has no 'reviews' field, potentially indicating an issue.")
print("Checked against the Actor's input schema and Apify docs on 2026-10-06")
Why does Yelp Scraper return fewer results than my searchLimit?
The Yelp Scraper prioritizes collecting results until the searchLimit is reached, the results on Yelp's side run out, or the runTimeoutSecs is hit. If your run terminates early, it's often due to the runTimeoutSecs being too low for the requested searchLimit and number of searchTerms and locations combinations. Yelp also uses Cloudflare, and persistent blocking can lead to fewer results if too many proxy sessions fail.
When setting searchLimit, it's crucial to understand how the scraper works. Yelp search pages use infinite scroll, and the Actor simulates this to collect results. This process takes time. The documentation states that "each scroll round adds a few seconds." If you set a high searchLimit (e.g., 50 per search term/location combination) but keep the default runTimeoutSecs of 1800 seconds (30 minutes) for a complex query, the run might hit its timeout before all results can be fetched. The Actor will "stop launching new fetches at the runTimeoutSecs mark, pushes everything it has collected so far, and exits cleanly." This results in a partial dataset without an explicit failure.
Consider this input, designed to fetch many results:
{
"searchTerms": ["restaurant", "cafe"],
"locations": ["New York, NY", "Los Angeles, CA", "Chicago, IL"],
"searchLimit": 50,
"reviewLimit": 5,
"runTimeoutSecs": 1800
}
This configuration attempts 6 distinct searches (2 searchTerms * 3 locations). If each search aims for 50 results, that's potentially 300 business records. Given that each search requires pagination and multiple scroll events, the 30-minute default timeout can easily be exceeded, especially under network latency or minor Cloudflare challenges. To avoid this, you need to either increase runTimeoutSecs significantly or reduce your searchLimit for exploratory runs. For large jobs, using the asynchronous run endpoint and polling or webhooks is also critical, as the synchronous run endpoint has a hard cap of 300 seconds.
What happens if the proxy fails, and how do I detect it?
The proxy input field is explicitly marked as "REQUIRED" because "Yelp uses Cloudflare protection. Residential proxy is mandatory; datacenter IPs are reliably blocked." If the proxy configuration is incorrect or the residential proxies themselves are exhausted or blocked by Cloudflare, the run will likely return a yelp_blocked record in the dataset and then terminate cleanly. This is a partial success state from the Actor's perspective, but a data failure from a user's.
A yelp_blocked record has type, reason, message, and scrapedAt fields. It's crucial to check for the presence of these records in your output dataset, especially if the expected number of business results is significantly lower than anticipated. While the Actor exits "successfully" with this record, your downstream processes should interpret it as a trigger for re-running the job, possibly with different proxy settings or a delay.
Here's an example of how to check for yelp_blocked records in your output:
import apify_client
client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
actor_id = "crawlerbros/yelp-scraper"
run_input = {
"searchTerms": ["invalidsearchterm"], # A term likely to return nothing, or cause blocks
"locations": ["Nowhere, US"],
"proxy": {
"useApifyProxy": True,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}
# To get a real dataset ID, you would run the Actor first:
run = client.actor(actor_id).call(run_input=run_input)
mock_dataset_id = run["defaultDatasetId"]
blocked_records = []
for item in client.dataset(mock_dataset_id).iterate_items():
if item.get("type") == "yelp_blocked":
blocked_records.append(item)
if blocked_records:
print(f"Detected {len(blocked_records)} 'yelp_blocked' records in the output:")
for record in blocked_records:
print(f" Reason: {record.get('reason')}, Message: {record.get('message')}")
print("Consider re-running the Actor with fresh residential proxy sessions or different input.")
else:
print("No 'yelp_blocked' records detected. Run likely completed without explicit blocks.")
print("Checked against the Actor's input schema and Apify docs on 2026-10-06")
It's also worth noting that residential proxy sessions typically persist for around 30 minutes. For long-running scrapes, especially those hitting a high searchLimit across many locations or direct URLs, proxy sessions might expire and need to be refreshed by the underlying Apify Proxy service. This is usually handled transparently, but severe, sustained blocking can indicate an issue where the proxy pool itself is struggling.
How can silent data loss from storage limitations be avoided?
Apify manages storage, and unnamed storages, like the default dataset created by Yelp Scraper runs, have retention policies. On the free plan, "only the 10 most recent runs are retained, for 4 months." If you are running the Yelp Scraper frequently on a free plan without explicitly naming your datasets, older data will be automatically deleted. This can lead to silent data loss if you aren't actively consuming and archiving your results.
To prevent this, ensure that any critical datasets generated by your Yelp Scraper runs are explicitly named. Named storages are always exempt from deletion. For example, if you schedule daily runs, you could dynamically name your dataset based on the run date.
Here's how you might create a named dataset before running the Actor:
import apify_client
from datetime import datetime
client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
# Create a named dataset
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
dataset_name = f"yelp-scraper-data-{timestamp}"
named_dataset = client.datasets().get_or_create(name=dataset_name)
actor_id = "crawlerbros/yelp-scraper"
run_input = {
"searchTerms": ["bakery"],
"locations": ["Paris, France"],
"searchLimit": 10,
"reviewLimit": 2
}
# Start the Actor and specify the named dataset for output
run = client.actor(actor_id).call(
run_input=run_input,
dataset_id=named_dataset["id"] # Direct the output to our named dataset
)
print(f"Run finished. Data saved to named dataset: {dataset_name} (ID: {named_dataset['id']})")
print(f"You can access it at: {client.dataset(named_dataset['id']).url}")
print("Checked against the Actor's input schema and Apify docs on 2026-10-06")
Using a named dataset ensures that your historical data remains accessible, regardless of your Apify plan or the number of subsequent runs.
What are the cost implications of Yelp Scraper's input parameters?
The Yelp Scraper follows a PAY_PER_EVENT pricing model. This means you are charged for specific named events that occur during its execution, in addition to general Apify platform usage. Understanding how your input parameters influence these events is key to managing costs. Platform usage for the run (memory and run time) is billed separately at your Apify plan rates.
The primary charge event for this Actor is "result" (technical name: apify-default-dataset-item). This event is emitted for "Single result in the default dataset." Therefore, the total number of business records (items) pushed to the dataset directly multiplies this cost.
- The "result" event costs $0.01 per event for FREE tier users.
- BRONZE tier users pay $0.00833 per event.
- SILVER tier users pay $0.00667 per event.
- GOLD, PLATINUM, and DIAMOND tier users pay $0.005 per event.
There is also an "Actor Start" event (technical name: apify-actor-start) which costs $0.05 per GB of memory allocated to the run, charged once per run. This means even a run that returns zero business results will still incur this initial charge.
-
searchLimit: This parameter directly impacts the number of "result" events. If you setsearchLimit: 10for a single search term and location, and 10 businesses are found, you will incur 10 "result" events. If you have multiplesearchTermsandlocations(e.g., 2searchTerms* 3locations* 10searchLimit= 60 potential results), the "result" events will scale proportionally, leading to higher costs. -
directUrls: Similarly, each business successfully scraped via adirectUrlsentry will count as one "result" event. -
reviewLimit: WhilereviewLimitaffects the content of each business result (by adding reviews to the business object), it does not generate separate "result" events. Each business record, regardless of how many reviews it contains, is still a single dataset item.
These event charges do not change based on run time or memory, but remember that platform usage for the run (memory and run time) is billed separately at your Apify plan rates.
How to manage scheduled Yelp Scraper runs?
When creating new schedules for the Yelp Scraper on the Apify platform, they are DISABLED by default. You must explicitly enable them after creation. Additionally, an Actor "must have run at least once before it can be scheduled." These are important considerations for setting up automated data extraction pipelines.
Consider a scenario where you want to schedule daily scrapes:
- Run the Actor manually once: Before you can schedule the
yelp-scraperActor, you must execute it at least once. This initializes its internal state on the platform for scheduling purposes. - Create a schedule: Define your cron expression and other schedule parameters. Remember that new schedules are
DISABLEDby default. - Enable the schedule: After creation, navigate to the schedule in the Apify Console and explicitly enable it.
This process ensures that your automated tasks will actually execute. For example, to set up a run for every day at 3 AM UTC:
{
"name": "Daily Yelp Scrape",
"actorId": "crawlerbros/yelp-scraper",
"cronExpression": "0 3 * * *",
"isEnabled": true,
"input": {
"searchTerms": ["restaurants"],
"locations": ["London, UK"],
"searchLimit": 10
}
}
This JSON represents the configuration for a schedule, which you would typically set up through the Apify Console or API. Note that isEnabled: true must be set after creation, as schedules default to disabled.
What are the key limitations and caveats for data engineers?
Beyond the specific failure modes, it's essential to understand the inherent limitations of web scraping, especially with a target like Yelp that employs active anti-bot measures.
- Cloudflare Protection: Yelp uses Cloudflare, which means continuous vigilance against bot detection. While the Actor mandates residential proxies to counter this, occasional blocks are still possible. If a run does get blocked across all available residential proxy sessions, it will exit successfully but with a
yelp_blockedrecord in the dataset, allowing you to detect this programmatically. - Dynamic HTML Changes: Yelp, like any dynamic website, can change its HTML structure without notice. The README mentions that "Yelp uses dynamic class names that can change," and "some fields may occasionally be empty if Yelp changes their HTML structure." This means even a perfectly configured run might return
nullor empty strings for fields likephone,description, or specific review fields if the scraper's parsing logic temporarily breaks. Defensively, your downstream pipelines should always account for the possibility of missing data for any non-essential field. - Review Extraction Limit: The scraper explicitly states that reviews are extracted only from "the reviews visible on the business page (first page of reviews)." Even if Yelp displays a total
reviewCountof thousands, you will only receive the reviews visible on that initial page, capped by yourreviewLimit. There is no current mechanism to paginate through all reviews for a single business using this Actor. - Search Result Reliability: "Search results depend on Yelp's ranking algorithm and may vary." This means repeated identical searches might yield slightly different orderings or even different sets of businesses over time. For auditing or consistency-critical tasks, this variability needs to be considered.
- Actor Task Consumption: While multiple runs can add to a request queue, a single request queue can only be processed by one Actor or task run at a time. This impacts design if you considered fanning out a single queue across multiple parallel Yelp Scraper runs.
- Input Prefill vs Default: The
prefillvalues in the input schema are only visible in the Console UI. When making API calls or using existing Actor tasks, only thedefaultvalues are applied if a field is omitted. Always pass an explicit input dictionary via API to ensure your desired configuration.
These limitations mean that robust data pipelines built around the Yelp Scraper require more than just calling an API. They need monitoring for yelp_blocked records, checks for partial or missing fields in the output, and an understanding of storage retention policies and proxy behavior.
Example Input
Search for businesses
{
"searchTerms": ["pizza", "sushi"],
"locations": ["New York, NY", "San Francisco, CA"],
"searchLimit": 5,
"reviewLimit": 3
}
Scrape specific businesses
{
"directUrls": [
"https://www.yelp.com/biz/prince-street-pizza-new-york-2",
"https://www.yelp.com/biz/joes-pizza-new-york"
],
"reviewLimit": 10
}
Checked against the Actor's input schema and Apify docs on 2026-10-06
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)