Why do synchronous API requests to the YouTube Video Downloader fail?
Synchronous API requests fail with a 408 timeout error when the processing duration exceeds the platform-enforced 300-second limit. Because downloading multiple high-resolution videos or long playlists often requires more than five minutes of execution time, a synchronous call will terminate on the client side even if the actor continues to work in the background.
When you trigger a run using the synchronous endpoint, the Apify platform waits for the process to complete before sending the HTTP response. For small tasks like fetching metadata for a single video, this works well. However, when you pass arrays of videoUrls or playlistUrls that require significant data transfer, the time required to negotiate streams and write files to storage can easily cross the 300-second threshold. Once the connection is dropped, you lose visibility into the run's success, even though the actor remains active in your dashboard. Moving to an asynchronous pattern involves sending a POST request to the /v2/acts/crawlerbros~youtube-video-downloader/runs endpoint, which returns a run object immediately. You can then poll the status or use a webhook to handle the data once the work is finalized.
When should you use the extractMetadataOnly parameter?
Use the extractMetadataOnly parameter to retrieve rich video information without triggering expensive and time-consuming media file downloads. This setting is ideal for scenarios where you need to build datasets of titles, descriptions, view counts, and tags without consuming bandwidth or storage for the actual video files.
By setting extractMetadataOnly to true, you bypass the data-heavy segments of the actor, which directly reduces the memory footprint and processing time of the run. This approach is highly efficient for bulk analysis, such as monitoring metadata trends across hundreds of channels. Since this mode does not attempt to save large video objects, it also makes the run significantly more resilient to network fluctuations.
{
"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"extractMetadataOnly": true,
"maxVideosPerInput": 50
}
How do you implement stateful resume-download patterns?
Implement stateful resume patterns by tracking successful video_id outputs in your own database and filtering them from your input before triggering a new run. Because the Youtube Video Downloader processes each item independently, failures on individual items do not inherently stop the entire queue, but network interruptions during large batches can cause gaps.
Since unnamed storage on the platform expires after four months and the free plan only retains the ten most recent runs, you should always direct your actor output to a named dataset. By querying your named dataset after a partial failure, you can extract the video_id of every successfully processed item. Before initiating your next run, compare this list against your initial input URLs to ensure you only request missing content, which preserves your budget and avoids redundant processing. The API returns a download_status field for every video, which is your primary indicator for success.
import requests
# Example of identifying failed downloads from the dataset
dataset_id = "YOUR_NAMED_DATASET_ID"
url = f"https://api.apify.com/v2/datasets/{dataset_id}/items"
response = requests.get(url).json()
failed_videos = [item['video_url'] for item in response if item.get('download_status') != 'success']
print(f"Videos requiring retry: {failed_videos}")
What happens when a residential proxy session rotates?
Residential proxy sessions typically expire after approximately 30 minutes, which can abruptly terminate a file download if the stream is particularly large or slow. When the connection drops, the actor may report a failed status in the download_status field for that specific item, even if the metadata was successfully extracted beforehand.
This is a common occurrence with long-form video content where the total transfer time exceeds the lifespan of a single residential IP session. If you encounter frequent drops, check the download_status field in your output records. If the error indicates a connection failure, consider reducing the number of concurrent processes if you are running multiple instances, or ensure your memory allocation is sufficient to maintain a fast download speed, thereby finishing the transfer before the proxy rotates.
How does the pricing model affect large-scale data collection?
The actor follows a PAY_PER_EVENT pricing model where event charges are determined by the count of specific charge events, with platform usage billed separately. Every "result" event in the default dataset is charged at $0.01, while every "Actor Start" event is charged at $0.05 per GB of memory allocated to the run. Your total expenditure is the sum of these events plus the separate platform usage fees billed according to your Apify plan.
The "result" event follows discount-tier pricing: FREE at $0.01, BRONZE at $0.00833, SILVER at $0.00667, GOLD at $0.005, PLATINUM at $0.005, and DIAMOND at $0.005. Because the number of events scales with the count of videos downloaded, you should configure your maxVideosPerInput field carefully to prevent costs from exceeding your expectations when scraping large channels or playlists. Always calculate your potential costs based on the expected number of successful result items.
How should you handle concurrent processing requests?
You must avoid attempting to process a single request queue with multiple simultaneous runs, as the platform does not support fan-out across one shared queue. If you have a massive list of URLs, the correct approach is to partition your source data into distinct batches and assign each batch to a separate, independent run. This ensures that each actor instance has its own isolated request queue and avoids race conditions during data extraction.
Furthermore, remember that if you are using the Apify API to trigger these runs, you must provide the full input object manually. The prefill values configured in the Console UI are ignored during API calls, so relying on them will result in unexpected behavior or default settings being applied to your partitioned batches. Always pass an explicit input dictionary via the API to maintain consistency.
curl -X POST "https://api.apify.com/v2/acts/crawlerbros~youtube-video-downloader/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"], "proxyCountry": "GB"}'
Which limitations exist for private or restricted content?
The actor cannot access members-only or age-restricted videos without valid authentication credentials provided via the cookies input field. Even with cookies, which must be in the Netscape format, there is no guarantee of access if the video is restricted by specific regional policies that your proxyCountry cannot resolve.
Additionally, the actor relies on public YouTube accessibility, meaning if a video is hidden or requires special account permissions, the download attempt will likely return a failed status. Ensure your cookies are managed through a secure key-value store rather than hardcoded in your scripts to prevent credential leakage.
Can you avoid manual polling for status updates?
You can avoid manual polling by utilizing webhooks that trigger a POST request to your backend service the moment the actor run completes. Polling the status endpoint is inefficient as it consumes unnecessary requests and increases the complexity of your application logic. By using a webhook, your server stays idle until it receives the completion signal, which contains the necessary run details to immediately trigger your next workflow step.
This event-driven model is essential for long-running batches where the completion time is variable. By integrating webhooks, you effectively decouple your application's state from the actor's execution, ensuring that you only proceed with database updates or data processing once the download is confirmed successful in the dataset storage.
Is memory allocation important for performance?
Memory allocation directly impacts the stability of your actor runs, especially when processing high-definition video files. While 1024 MB is the default and typically suffices for standard 720p downloads, tasks involving high-resolution 1080p content or very long durations may require 2048 MB or more to avoid memory-related crashes.
If you are running metadata-only jobs, you can safely lower your memory allocation to 256 MB. Always check your run logs to identify if you are hitting memory limits. If the actor crashes or the process terminates prematurely, increasing the memory allocation is the first step in troubleshooting, though keep in mind that higher memory increases the "Actor Start" event cost.
{
"video_url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"video_id": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Music Video)",
"view_count": 1400000000,
"download_status": "success",
"download_url": "https://api.apify.com/v2/key-value-stores/.../records/yt_dQw4w9WgXcQ.mp4"
}
How to structure data for high-volume playlist processing?
Processing long playlists requires granular control over input to avoid exceeding memory or time limits per run. Using the maxVideosPerInput field is the most effective way to prevent a single run from ballooning in cost or duration, as it forces the actor to cap the number of videos processed before finishing the job.
If you have a playlist with 1,000 videos, splitting it into ten batches of 100 via multiple runs is safer than a single massive run. This approach ensures that if a proxy rotation or an unexpected network hiccup occurs, you only risk losing the progress of one small batch rather than the entire queue. Additionally, ensure you use named datasets for these batches to keep your data organized for eventual merging.
// Example of batching URLs for multiple API triggers
const playlistUrls = ["https://www.youtube.com/playlist?list=PLexample"];
const batchSize = 50;
for (let i = 0; i < playlistUrls.length; i += batchSize) {
const batch = playlistUrls.slice(i, i + batchSize);
// Call Apify API to start a run with the subset of data
console.log("Triggering run for batch:", batch);
}
What are the implications of unnamed storage expiration?
Unnamed storage on the Apify platform expires after a set period, and for users on the free plan, only the ten most recent runs are retained. If you rely on the default dataset without explicitly creating a named dataset, you risk losing your download history and metadata records after the platform purges old runs.
This behavior makes it critical to integrate your workflow with named datasets. A named dataset exists independently of the run that created it and is exempt from the deletion policies that affect unnamed, temporary storage. By defining a datasetId or using the default dataset naming convention and promptly migrating data to an external database, you ensure your work remains accessible for long-term audit or historical comparison.
How to troubleshoot failed download status codes?
When the download_status field returns failed, you should immediately inspect the accompanying error field to determine whether the issue stems from a network timeout, a geo-block, or a broken source link. Because the Youtube Video Downloader outputs a record even for failed attempts, you can easily filter your dataset to identify problematic URLs.
Common failures involve videos that have been set to private or age-restricted, which the actor cannot bypass without correct cookies. If the error mentions proxy connectivity, verify that the proxyCountry matches a region where the content is legally accessible. Using the extractMetadataOnly mode as a diagnostic step can also help confirm if the video exists and is accessible, effectively isolating download-specific errors from general availability issues.
Checked against the Actor's input schema and Apify docs on 2026-09-23.
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)