Understanding YouTube Comment Scraper's Event-Driven Cost Model
Working with Apify Actors like the YouTube Comment Scraper requires data engineers to understand its underlying charge events. This Actor operates on a pay-per-event model, where specific, named actions determine your bill. Understanding these events and how input parameters influence their frequency is key to predictable and efficient data extraction.
What are the primary charge events for the YouTube Comment Scraper?
The YouTube Comment Scraper charges for two distinct event types. An "Actor Start" event costs $0.05 per GB of memory allocated to the run, with a minimum of one event. The "result" event costs $0.005 per single result stored in the default dataset, and is often far more significant. Platform usage for the run (memory and run time) is billed separately at your Apify plan's rates.
The "result" event directly correlates with the number of top-level comments the Actor successfully scrapes and pushes to its default dataset. Each top-level comment counts as one "result" event, regardless of whether it includes replies. The cost for these "result" events is subject to Apify's discount tiers: FREE users pay $0.005, BRONZE $0.00433, SILVER $0.00367, and GOLD, PLATINUM, and DIAMOND tiers all pay $0.003. Users on higher discount tiers will see a lower per-result cost.
How does maxComments influence the number of billable events?
The maxComments input field determines the upper limit of top-level comments the Actor attempts to scrape for each videoUrl provided. Each top-level comment pushed to the dataset (including placeholder marker rows for filtered-out comments that still meet other criteria) constitutes a "result" event, meaning a higher maxComments value directly increases the potential count of these billable events.
For example, scraping 1,000 top-level comments across multiple videos, where maxComments allows for this, would generate 1,000 "result" events. Replies, importantly, are nested within the top-level comment's output object and do not count as separate "result" events themselves.
Consider a scenario where you're processing multiple videos. If you provide an array of videoUrls and set maxComments: 250, the Actor will aim to scrape up to 250 top-level comments for each video. If all videos yield their target number of comments, and all those comments are pushed to the dataset, your total "result" events will be len(videoUrls) * maxComments.
Here's an example input targeting multiple videos with a specific maxComments limit:
{
"videoUrls": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://www.youtube.com/watch?v=kYtJ0k2YfRk"
],
"maxComments": 250,
"includeReplies": false,
"sortBy": "newest"
}
In this case, for each of the two video URLs, the Actor will attempt to fetch up to 250 top-level comments. If successful and all comments are pushed to the dataset, this would lead to approximately 500 "result" events for these comments, plus the "Actor Start" charge.
Does includeReplies add separate billable events?
No, the includeReplies boolean input field does not add separate "result" events to your bill. The Apify YouTube Comment Scraper is designed such that each top-level comment is a single item in the default dataset. When includeReplies is set to true (the default), any fetched replies for a top-level comment are nested within that comment's replies array in the same dataset item. This means a top-level comment with 0 replies and a top-level comment with 50 replies both count as exactly one "result" event.
While includeReplies doesn't directly increase the "result" event count, enabling it does consume more platform resources (CPU, memory, network traffic) because the Actor performs additional requests and processing to fetch and structure those replies. This increased resource consumption will be reflected in the separate Apify platform usage charges, but not in the specific event prices for "result" items.
The output schema clearly illustrates this structure:
{
"commentId": "UgxB...",
"text": "Great video!",
"authorName": "John Doe",
"replyDepth": 0,
"videoId": "dQw4w9WgXcQ",
"replies": [
{
"commentId": "UgxB....AbCdEfGhIj",
"text": "I agree!",
"authorName": "Jane Smith",
"replyDepth": 1,
"parentCommentId": "UgxB...",
"commentUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ&lc=UgxB....AbCdEfGhIj"
}
]
}
In this example, the entire structure, including the nested replies array, constitutes a single "result" event. This design is highly beneficial for cost management if your primary concern is per-item billing for top-level comments.
How do maxRepliesPerComment, commentKeywordFilter, and minLikeCount affect cost?
These input fields are designed to refine the data pushed to the dataset, primarily affecting platform usage and the granularity of collected data, rather than adding separate "result" events.
maxRepliesPerComment: This integer field (default5) limits the number of replies fetched for each top-level comment whenincludeRepliesistrue. LikeincludeReplies, changing this value does not alter the "result" event count. Each top-level comment is still one "result." However, reducingmaxRepliesPerCommentwill reduce the platform usage associated with fetching replies, potentially lowering your platform usage costs. Setting it to0effectively skips all replies, even ifincludeRepliesistrue.commentKeywordFilter: This string input allows you to filter comments and their replies based on a case-insensitive keyword or phrase. If a top-level comment does not match the filter, its replies are skipped entirely, saving platform usage. This is an efficient optimization because the Actor applies this filter before attempting to fetch replies. If a reply is fetched, it's also filtered independently. ThecommentKeywordFilteronly reduces the number of comments pushed to the dataset; comments filtered out and not pushed do not incur a "result" event charge, though they may still contribute to platform usage by being initially fetched.minLikeCount: Similar to the keyword filter,minLikeCount(default0) filters comments and replies based on their like count. A top-level comment below this threshold will not have its replies fetched, reducing platform usage. Like the keyword filter, this parameter is applied after fetching from YouTube but before saving to the dataset. It helps you focus on more engaging comments without incurring "result" event charges for filtered-out items.
These filtering parameters offer powerful ways to optimize your data collection to only include relevant comments, saving on downstream processing and storage, and potentially reducing platform usage, without directly impacting the per-result event cost.
Consider a Python script using the Apify Client to run the Actor with specific filters:
from apify_client import ApifyClient
apify_client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"maxComments": 500,
"includeReplies": True,
"maxRepliesPerComment": 3,
"commentKeywordFilter": "great video",
"minLikeCount": 10,
"sortBy": "top"
}
run = apify_client.actor("crawlerbros/youtube-comment-scraper").call(run_input=run_input)
# Fetch results from the default dataset
for item in apify_client.dataset(run["defaultDatasetId"]).iterate_items():
print(item.get("text"))
In this example, commentKeywordFilter and minLikeCount will reduce the number of comments ultimately stored. Only comments that pass these filters and are pushed to the dataset will generate "result" events. If a top-level comment is filtered out, its replies are not fetched, saving platform usage.
Understanding Failure Modes and Their Cost Impact
The Actor's robust error handling ensures that even failed attempts to scrape a video will generate a single marker row in the dataset, containing error details rather than silently producing nothing. This is a critical design choice for visibility but also means that you're charged $0.005 (or your tier's price) for each failed video processed, even if no comments are extracted. Importantly, maxComments does not apply to these marker rows; each failed video URL always results in exactly one marker row.
Consider the various failure scenarios:
- Invalid URL or video ID: If
videoUrlscontains an unparseable string or a non-existent video ID, the Actor still processes it. It will push a marker row withsuccess: falseand anerrormessage. This counts as one "result" event for billing purposes. - Comments disabled: For a valid video where comments are disabled, a marker row with
success: falseandcommentsDisabled: trueis pushed. This counts as one "result" event. - Bot detection/network error: If YouTube blocks the scraper or a network issue occurs, a marker row with
success: falseand anerrormessage is created. This counts as one "result" event.
This behavior means that to accurately estimate your event costs, you must consider that each failed video URL results in one "result" event (a marker row), irrespective of the maxComments setting. An efficient strategy involves pre-validating video URLs where possible to minimize these "failed result" charges.
The output schema for an error entry would look something like this:
{
"success": false,
"error": "Video with ID 'INVALIDID' does not exist or is private/deleted.",
"inputUrl": "https://www.youtube.com/watch?v=INVALIDID",
"scrapedAt": "2026-02-11T12:00:00.000000+00:00"
}
This single error object still represents a billable "result" event.
Avoiding Duplicate Charges When Paginating Many Videos
For large-scale scraping of comments across numerous videos, a common pattern involves using Apify's request queues. However, a crucial platform limitation to remember is that a request queue can only be processed by one Actor or task run at a time. If you try to fan out multiple instances of the YouTube Comment Scraper to consume from a single shared queue, it will not work as expected, leading to inefficiencies or errors.
Instead, for processing a very large list of videos, you might consider orchestrating multiple independent runs, each with its own specific videoUrls array, or using a "parent" Actor to manage and fan out jobs to individual "child" runs of the YouTube Comment Scraper. Each videoUrls entry in an input will be processed sequentially within a single run, and the Actor is designed to avoid duplicate comments within a single video's processing by tracking comment IDs. It also skips duplicate video URLs within its own videoUrls input array.
If you are using external orchestration (e.g., via a Python script), ensure that your logic handles unique video URLs across runs to avoid redundantly processing and paying for the same video comments multiple times.
Consider using a simple loop to call the Actor for chunks of videos, ensuring each run is distinct:
from apify_client import ApifyClient
apify_client = ApifyClient("YOUR_APIFY_TOKEN")
actor_client = apify_client.actor("crawlerbros/youtube-comment-scraper")
all_video_urls = [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://www.youtube.com/watch?v=kYtJ0k2YfRk",
# ... many more URLs
]
# Process videos in chunks of 100
chunk_size = 100
for i in range(0, len(all_video_urls), chunk_size):
chunk_urls = all_video_urls[i:i + chunk_size]
run_input = {
"videoUrls": chunk_urls,
"maxComments": 100,
"includeReplies": False
}
print(f"Starting run for {len(chunk_urls)} videos...")
run = actor_client.call(run_input=run_input)
print(f"Run {run['id']} finished, default dataset: {run['defaultDatasetId']}")
# You would typically process results here or set up a webhook
This approach creates distinct runs, each independently charged for its "Actor Start" event and the "result" events it generates.
Limitations and Caveats for Cost Optimization
While understanding the event-driven pricing model helps, there are practical limitations that can indirectly affect your overall cost:
- Platform Usage: Always remember that event charges are in addition to platform usage. Heavy use of
includeRepliesand highmaxRepliesPerCommentvalues, even if they don't add "result" events, will increase the resources (CPU, memory, network) consumed by your Actor run. This will result in higher platform usage costs, which are billed separately at your Apify plan's rates. It's a trade-off: more comprehensive data per item versus lower resource consumption. - YouTube's Dynamic Loading and API Limits: The Actor README notes that YouTube loads comments dynamically and the scraper may not always see all comments on very large videos. While
maxCommentssets an upper bound on what you request, YouTube's actual response can be fewer. You will only be charged "result" events for the comments successfully pushed to the dataset, not for comments that YouTube's API simply didn't return. However, the platform usage for the attempt to fetch remains. - Synchronous Run Cap: For interactive development or smaller tasks, the synchronous run endpoint hard-caps at 300 seconds (5 minutes) and returns an HTTP 408 if exceeded. For longer-running tasks, you must POST to
/v2/acts/<actor>/runsand poll for status or use a webhook. This doesn't directly impact the per-event cost model but is a critical operational constraint for controlling run duration and thus platform usage. A run hitting this cap will still incur event charges for results pushed up to that point, plus the platform usage. - Data Retention: Unnamed storages (like the default dataset for an Actor run) expire. On the Apify Free plan, only the 10 most recent runs are retained for 4 months. If you need to retain your scraped comments long-term and avoid re-scraping (and thus re-paying event charges), ensure you either name your datasets or export the data promptly. Re-running the Actor for already-scraped data unnecessarily repeats event charges.
Checked against the Actor's input schema and Apify docs on 2026-10-05.
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)