Darkweb Scraper Output Omits Empty Fields, Not Just Nulls
When working with external data extraction tools or APIs, a common assumption is that fields will always be present in the output, perhaps as null or empty arrays if no data is found. The Darkweb Scraper, an Apify Actor, presents a more nuanced behavior: any field that would be an empty array, object, or string is omitted entirely from the record. This is not merely an aesthetic choice; it's a critical detail for data engineers building robust downstream processing. Your data pipelines must defensively check for the presence of keys like emails or cryptoAddresses before attempting to access their values, rather than assuming they exist and might be empty. Failure to do so will result in KeyError or similar exceptions if you're not careful. These defensive programming techniques are general best practices for working with any external data source, not exclusive to this particular tool.
The implication is clear: a simple record.get('emails', []) in Python, or record?.emails || [] in JavaScript, becomes a necessity to safely access potentially missing fields. This design choice pushes error handling upstream into your data consumption logic, requiring explicit checks for field existence.
Consider the following output snippet. If emails or phones were not found, they simply wouldn't be present:
{
"url": "http://xjfbpuj56rdazx4iolylxplbvyft2onuerjeimlcqwaihp3s6r4xebqd.onion/",
"title": "Dark Market - Home",
"emails": ["contact@darkservice.onion"],
"cryptoAddresses": {
"bitcoin": ["1A1zP1eP5QGefi2DMPTfTL5SLmv7DivfNa"]
},
"scrapedAt": "2026-06-18T12:00:00.000000+00:00"
}
A Python script consuming this would need to handle the potential absence of phones or misc explicitly:
import json
sample_output = json.loads("""
{
"url": "http://xjfbpuj56rdazx4iolylxplbvyft2onuerjeimlcqwaihp3s6r4xebqd.onion/",
"title": "Dark Market - Home",
"emails": ["contact@darkservice.onion"],
"cryptoAddresses": {
"bitcoin": ["1A1zP1eP5QGefi2DMPTfTL5SLmv7DivfNa"]
},
"scrapedAt": "2026-06-18T12:00:00.000000+00:00"
}
""")
emails = sample_output.get("emails", [])
phones = sample_output.get("phones", []) # Will be an empty list if "phones" key is absent
social_handles = sample_output.get("misc", {}).get("telegram", []) # Nested check
print(f"Emails found: {emails}")
print(f"Phones found: {phones}")
print(f"Telegram handles found: {social_handles}")
Why Does the Darkweb Scraper's mode=search Return Zero Results?
The Darkweb Scraper Actor's mode=search configuration, which queries dark web search engines like Ahmia, Torch, and Haystack, is inherently a "best-effort" operation. These search engines frequently change their .onion addresses or go offline entirely. The Actor attempts multiple engines sequentially, but it's not uncommon for a run to yield no results if all engines are unreachable or fail to index relevant content at that moment.
This is a known operational characteristic, not a bug, and directly impacts the reliability of data acquisition when relying solely on search. To mitigate this, always prefer mode=crawl with explicitly known good startUrls for critical data acquisition. If search is necessary, implement retry logic in your orchestration for the entire Actor run, recognizing that ephemeral unavailability of search engines is a primary failure mode. The README explicitly states that mode=crawl is the "reliable default" and that "search mode may return zero results on any given run." Other dark web scraping tools may encounter similar issues with ephemeral search indexes.
How Does Input Schema Impact Run Reliability?
The input schema for the Darkweb Scraper explicitly defines constraints that dictate run reliability and expected outcomes. For instance, mode=crawl requires at least one entry in startUrls. If startUrls is empty while mode is set to crawl, the Actor run will likely fail quickly or produce zero results because it has no seed URLs to begin scraping. Conversely, mode=search requires a search keyword.
Attempting to run mode=search without a keyword will similarly lead to an unproductive or failed run. The searchAndCrawl mode requires at least one of these two fields (search or startUrls), allowing for a merged approach but still enforcing that at least one discovery mechanism is provided. These implicit constraints, derived from the required status and conditional dependencies in the Note section of the input documentation, tell you exactly what will break if your API calls don't respect them. Defensive code should validate these combinations before initiating an Actor run, avoiding unnecessary charges for failed executions.
Here's an example of how you might construct a valid input for mode=crawl, adhering to the schema requirements:
{
"mode": "crawl",
"startUrls": [
{ "url": "http://bbcnewsd73hkzno2ini43t4gblxvycyac5aw4gnv7t2rfl3d5jakber2iniad.onion/" }
],
"maxDepth": 1,
"maxPages": 5,
"maxItems": 5,
"extractEmails": true,
"extractPhones": false,
"extractCryptoAddresses": true,
"extractSocialHandles": true,
"extractApiKeys": false
}
What Happens When the Synchronous Run Endpoint Times Out?
Apify's synchronous run endpoint, used for simpler, faster Actor executions, has a hard-capped limit of 300 seconds (5 minutes). If a run exceeds this duration, the endpoint returns an HTTP 408 "Request Timeout" error. This does not mean the Actor run itself has stopped; it merely indicates that the HTTP connection to your client was severed. The Actor continues to execute in the background.
For the Darkweb Scraper, where "Tor is slow" and "each page to take 5-30 seconds to load," hitting this 300-second limit is a very real possibility, especially with higher maxPages or maxDepth values. To handle runs that might exceed 300 seconds, you must initiate the Actor run asynchronously by POSTing to /v2/acts/<actor>/runs and then polling the run's status or setting up a webhook for completion notifications. Relying on the synchronous endpoint for any non-trivial Darkweb Scraper use case is a recipe for silent failures in your client applications. This asynchronous handling is a general best practice for any long-running API operation.
Here's a Python example demonstrating an asynchronous run and polling:
import apify_client
import time
import os
# Initialize the ApifyClient with your API token
client = apify_client.ApifyClient(os.environ["APIFY_API_TOKEN"])
# Get the Actor by its ID or name
actor_id = "crawlerbros/darkweb-scraper"
actor = client.actor(actor_id)
# Input data for the Actor
run_input = {
"mode": "crawl",
"startUrls": [
{ "url": "http://bbcnewsd73hkzno2ini43t4gblxvycyac5aw4gnv7t2rfl3d5jakber2iniad.onion/" }
],
"maxDepth": 1,
"maxPages": 10, # Higher maxPages increases the chance of exceeding 300 seconds
"maxItems": 10
}
print(f"Starting Actor {actor_id} asynchronously...")
run = actor.call(run_input=run_input, wait_for_finish=0) # wait_for_finish=0 means async
run_id = run["id"]
print(f"Run started with ID: {run_id}. Polling for status...")
while True:
run_status = client.run(run_id).get()
status = run_status["status"]
print(f"Current run status: {status}")
if status in ["SUCCEEDED", "FAILED", "ABORTED"]:
print(f"Run {run_id} finished with status: {status}")
break
time.sleep(10) # Poll every 10 seconds
# Access results if successful
if status == "SUCCEEDED":
dataset = client.dataset(run["defaultDatasetId"])
print("Fetched results:")
for item in dataset.iterate_items():
print(item)
How Can I Prevent Accidental High Costs with maxTotalChargeUsd?
The ACTOR_MAX_TOTAL_CHARGE_USD environment variable (set via the maxTotalChargeUsd query parameter on run endpoints) is a crucial safety mechanism for controlling costs. When the cumulative cost of charged events and platform usage for an Actor run approaches or exceeds this cap, the run is signaled to terminate. While it's not an instantaneous kill and the run might briefly consume resources past the limit, it provides a strong safeguard against runaway expenses.
This cost control mechanism is generally available for Apify Actors.
For the Darkweb Scraper, which can be configured for deep crawls (maxDepth) and many pages (maxPages), costs can escalate. Combining maxTotalChargeUsd with reasonable maxPages and maxItems limits is a robust strategy. However, keep in mind the disclaimer that "Dark web sites are unreliable" and "Tor is slow." A run might spend significant time and therefore platform usage (which is billed separately from Actor events) attempting to access unresponsive .onion sites before reaching its defined maxPages or maxItems, potentially hitting the cost cap earlier than expected due to platform usage.
Here's an example of how you might include maxTotalChargeUsd when starting a run (note: this is a conceptual parameter for the Apify API call; its value isn't directly part of the run_input JSON):
# Assuming 'actor' object from previous example
# This 'maxTotalChargeUsd' is passed as a parameter to the Actor.call method, not within run_input
run = actor.call(
run_input=run_input,
wait_for_finish=0,
max_total_charge_usd=0.50 # Cap the run at 50 cents USD for total cost
)
Are Apify's prefill and default Input Schema Values the Same?
No, the prefill and default fields in an Apify Actor's input schema serve different purposes, and this distinction is crucial when interacting via the API. The prefill value is solely a convenience for the Apify Console UI; it populates the input fields when a user first opens the Actor's input form. It has no effect whatsoever on API calls or existing Actor tasks.
In contrast, the default value is applied when an input field is not explicitly provided in an API call or Actor task run. If you omit an input field for the Darkweb Scraper (e.g., maxDepth), it will revert to its default value (e.g., 1). This means that when programmatically initiating runs, you must always provide an explicit input dictionary for all fields where you want to override the default, or for any field that is conditionally required (like startUrls for mode=crawl), even if you see a value pre-filled in the Console. Relying on prefill for API-driven workflows is a common pitfall that can lead to unexpected Actor behavior. This distinction applies across various Apify Actors, not just the Darkweb Scraper.
Consider this example where maxDepth is explicitly set, overriding its default:
import apify_client
import os
client = apify_client.ApifyClient(os.environ["APIFY_API_TOKEN"])
actor = client.actor("crawlerbros/darkweb-scraper")
# Input explicitly setting maxDepth
explicit_input = {
"mode": "crawl",
"startUrls": [
{ "url": "http://examplev3onyc4x.onion/" }
],
"maxDepth": 2, # Overrides the default of 1
"maxPages": 5
}
# Input omitting maxDepth, it will use the default of 1
default_input = {
"mode": "crawl",
"startUrls": [
{ "url": "http://examplev3onyc4x.onion/" }
],
"maxPages": 5
}
print("Running with explicit maxDepth=2...")
run_explicit = actor.call(run_input=explicit_input, wait_for_finish=10) # Run for 10 seconds
print(f"Explicit run finished with status: {run_explicit['status']}")
print("Running with default maxDepth=1 (omitted from input)...")
run_default = actor.call(run_input=default_input, wait_for_finish=10) # Run for 10 seconds
print(f"Default run finished with status: {run_default['status']}")
Understanding Darkweb Scraper's Named Storage Expiration
When running Actors on Apify, the storage objects they produce (datasets, key-value stores, request queues) can be either unnamed or named. The Darkweb Scraper, by default, uses unnamed storage for its dataset output. This has important implications for data retention, especially for users on the Apify Free plan. Unnamed storages expire, and on the free plan, only the 10 most recent runs are retained, for a maximum of 4 months. If you rely on the data extracted by the Darkweb Scraper for long-term analysis or compliance, this default behavior can lead to unexpected data loss.
To prevent your scraped dark web data from expiring, you must explicitly name your storage. This is typically done through the Apify Console or by specifying datasetId when initiating a run via the API. Named storages are always exempt from deletion policies that affect unnamed ones. For critical threat intelligence or security research data, ensuring named storage is used is a non-negotiable step.
Here's how you might create a named dataset using the Apify Client:
import apify_client
import os
client = apify_client.ApifyClient(os.environ["APIFY_API_TOKEN"])
actor = client.actor("crawlerbros/darkweb-scraper")
# Create a named dataset before the run
dataset_name = "my-darkweb-intel-dataset"
named_dataset = client.datasets().get_or_create(name=dataset_name)
print(f"Using named dataset: {named_dataset['id']}")
run_input = {
"mode": "crawl",
"startUrls": [
{ "url": "http://bbcnewsd73hkzno2ini43t4gblxvycyac5aw4gnv7t2rfl3d5jakber2iniad.onion/" }
],
"maxDepth": 1,
"maxPages": 5,
"maxItems": 5
}
# Start the Actor run, linking it to the named dataset
run = actor.call(run_input=run_input, dataset_id=named_dataset['id'], wait_for_finish=300)
print(f"Run {run['id']} finished with status: {run['status']}")
print(f"Results stored in dataset: {named_dataset['id']}")
What Does the Darkweb Scraper Cost?
The Darkweb Scraper uses a PAY_PER_EVENT pricing model, meaning its cost is calculated based on specific events generated during the run, in addition to general Apify platform usage. It's important to understand that the event charges are separate from platform usage, which is billed at your Apify plan's rates (e.g., for memory, storage, bandwidth, etc.).
The charged events are:
-
result(apify-default-dataset-item): Charged at $0.002 per event. This event occurs for each single result record pushed into the default dataset.- Discount tiers apply: FREE $0.002, BRONZE $0.00167, SILVER $0.00133, GOLD $0.001, PLATINUM $0.001, DIAMOND $0.001. Your cost per result will vary based on your Apify user tier.
- The
maxItemsinput field directly controls the maximum number ofresultevents, thereby capping this part of the cost.
-
Actor Start(apify-actor-start): Charged at $0.05 per GB of memory allocated to the run. This is a flat per-event price based on memory, charged once per run. The minimum charge is $0.05 per GB of memory.
The event-based portion of your cost scales with the number of results generated (maxItems is your primary control here) and the memory allocated to the run. The platform usage cost scales with the overall resources consumed (duration, data transfer, etc.).
What Are the Core Limitations of Darkweb Scraper?
Understanding the Darkweb Scraper's limitations is as important as knowing its capabilities to avoid frustration and build realistic expectations for your data pipelines.
Firstly, the Actor can only access .onion (Tor hidden service) URLs. It will not crawl or return data from regular clearnet websites, even if they are linked from a .onion page. Any attempt to provide non-onion startUrls will be ignored or result in an error.
Secondly, the reliability of mode=search is fundamentally constrained by the dark web search engines it queries. These indexes are often incomplete, frequently offline, or move addresses without warning. This means that not all .onion sites are discoverable via search, and a search run may legitimately return zero results. Other dark web scraping tools may face similar challenges with search engine reliability.
Thirdly, Tor network connections are inherently slow. The Actor bundles its own Tor daemon, which simplifies setup, but the network itself adds significant latency. Expect individual page loads to take anywhere from 5 to 30 seconds. This directly impacts the overall run duration, making quick, large-scale crawls impractical compared to scraping the regular internet.
Finally, the Actor does not render JavaScript. Many modern websites, including some on the dark web, rely heavily on client-side JavaScript to load and display content. If a .onion site requires JavaScript for its core content, the Darkweb Scraper will only see the initial HTML, potentially returning incomplete or empty data for that page. This is a significant limitation for dynamic sites. More complex (and often more costly) scraping approaches, such as using full browser automation tools like Playwright or Puppeteer, would be required to handle JavaScript rendering, but these are outside the scope and design of this specific Actor. Additionally, if a site requires a CAPTCHA or complex anti-bot measures, the scraper may also be prevented from accessing content, leading to skipped pages or incomplete data. These are intrinsic characteristics of scraping the dark web, not bugs in the Actor.
Checked against the Actor's input schema and Apify docs on 2026-09-24.
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)