DEV Community

Cover image for Why TikTok Ads Library Scraper Pro Returns Empty Media URLs
Crawler Bros
Crawler Bros

Posted on

Why TikTok Ads Library Scraper Pro Returns Empty Media URLs

As data engineers, we often integrate with third-party tools and APIs, navigating their stated capabilities and their undocumented constraints. The real skill isn't just knowing what a tool does, but understanding how it fails, and how to build defensively against those failure modes. This isn't always obvious from product overviews or simple "getting started" guides.

Let's look at a specific example: the TikTok Ads Library Scraper Pro on Apify. This tool, found at https://apify.com/crawlerbros/tiktok-ads-library-scraper-pro, provides access to TikTok's public ad transparency library. While its functionality for scraping ad data by query, advertiser, region, and date range is well-documented, the subtle implications of its input schema and operational environment tell a deeper story about where things can go wrong.

What Input Schema Constraints Mean for Data Completeness?

The tiktok-ads-library-scraper-pro input schema contains implicit contracts about what combinations of inputs will yield complete data, and what will result in partial or empty fields. Consider the quickSearch boolean field. When true, it explicitly skips "ad-detail fetch and only return data from the listing card." The README further clarifies that this omission includes "full targeting breakdowns."

This means if your downstream application expects fields like targetingByLocation, ageBuckets, or genderSelection, setting quickSearch to true will lead to missing data for those specific fields. This isn't a failure in the traditional sense, but a deliberate design choice with consequences for data completeness.

To avoid this, always explicitly set quickSearch: false if your data model requires the full targeting breakdown. If you're building a generalized data pipeline, you might need a conditional branch to handle the two different output shapes, or ensure all runs that feed into a particular downstream consumer consistently request full details.

{
  "query": "sustainability",
  "region": "all",
  "quickSearch": false,
  "maxItems": 100
}
Enter fullscreen mode Exit fullscreen mode

In the above input, quickSearch: false signals the intent to retrieve full targeting breakdowns. Neglecting this could silently introduce None or empty values for detail-rich fields, leading to breaks in downstream analytics or reporting that assumes their presence.

Why Does TikTok Ads Library Scraper Pro Sometimes Return Empty Media URLs?

A common scenario where mediaUrl might be None or an empty string, despite the ad appearing in the results, is when an ad's compliance status changes. The FAQ states, "Ads with auditStatus: rejected or removed are still listed in the transparency database but have their creative assets stripped." This means you can get valid adId, adText, and targeting data, but no actual creative media to download.

This behavior implies a need for defensive programming when consuming the mediaUrl field. Instead of assuming its presence, always check for null or empty values before attempting to process or download the media. Your code should explicitly handle auditStatus values that indicate the media might be unavailable, perhaps by logging the adId for manual review or by skipping media-dependent processing steps.

Here's a Python snippet demonstrating how to detect and handle this in your data processing:

import apify_client

# Initialize the Apify client
client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

# Assume we have a run ID or directly fetch dataset items
# This example fetches the latest run's dataset
actor_id = "crawlerbros/tiktok-ads-library-scraper-pro"
run = client.actor(actor_id).call(run_input={
    "query": "shoes",
    "quickSearch": false,
    "maxItems": 20
})

# Access the dataset and iterate items
dataset = client.dataset(run["defaultDatasetId"])
for item in dataset.iterate_items():
    ad_id = item.get("adId")
    media_url = item.get("mediaUrl")
    audit_status = item.get("auditStatus")

    if audit_status in ["rejected", "removed"] and not media_url:
        print(f"Ad {ad_id}: Media URL is empty due to audit status '{audit_status}'. Skipping media download.")
    elif media_url:
        print(f"Ad {ad_id}: Media URL available: {media_url}")
        # Proceed with media download or processing
    else:
        print(f"Ad {ad_id}: Media URL is unexpectedly empty. Audit status: {audit_status}")

Enter fullscreen mode Exit fullscreen mode

This approach safeguards your pipeline against unexpected None values for media and allows you to build logic around the known reasons for their absence.

How Do Proxy Limits Affect Data Collection Stability?

The useApifyProxy input parameter, defaulting to true, implies that network stability and rate limiting are real concerns when interacting with TikTok's transparency library. The description for this field explicitly states, "Recommended on cloud, avoids IP rate-limiting. The actor automatically rotates the proxy session on errors." While robust, proxy sessions have finite lifespans: datacenter proxies persist for around 26 hours, while residential proxies last approximately 30 minutes.

For long-running or frequently scheduled data collection jobs, a prolonged run could theoretically exhaust the stability of a single proxy session, even with automatic rotation. If a run attempts to use a stale session for too long without successful rotation, it could encounter prolonged rate limiting or connection errors, leading to incomplete datasets or failed runs.

While the actor handles rotation, understanding this underlying mechanism is crucial for diagnosing inexplicable connection issues on very long runs or when observing performance degradation over extended periods. For most typical runs, the default proxy handling is sufficient. However, if you are collecting millions of records over many hours, monitor run logs for frequent proxy rotation messages or networking errors, which could indicate stress on the proxy pool.

# Example of explicitly enabling proxy usage, though it's the default
# If you were to override with a custom proxy (not recommended unless
# you have a specific, known-good proxy setup), you'd define proxyConfiguration
# For general use, rely on the default `useApifyProxy: true`.
run_input = {
    "query": "fashion",
    "region": "all",
    "quickSearch": false,
    "useApifyProxy": True, # Explicitly stating the default for clarity
    "maxItems": 5000
}
Enter fullscreen mode Exit fullscreen mode

This configuration is generally robust. However, if your data pipeline involves very high-volume, continuous scraping against TikTok's ad library, be aware that even managed proxy services have underlying session durations. If a run lasts for days, you might consider splitting it into smaller, sequential runs or monitoring for proxy-related warnings.

When Does a Run Exceed the Synchronous 300-Second Cap?

When calling an Actor via the Apify platform's synchronous POST /v2/acts/<actor_id>/run endpoint, there's a hard timeout of 300 seconds (5 minutes). If the Actor run doesn't complete and return its results within this window, the API call will return an HTTP 408 error. This doesn't mean the Actor itself stops; it continues running in the background, but your client application loses its synchronous connection to the result.

For tiktok-ads-library-scraper-pro, several input parameters can increase run duration, making it prone to hitting this cap:

  • maxItems: A large number of items can significantly extend the run time, especially when quickSearch is false (which requires an additional request per ad).
  • quickSearch: false: Each ad requires a separate detail page fetch, multiplying the request count and thus the duration.
  • requestDelaySecs: Increasing this value (the pause between successive API calls) directly extends run time, though it reduces rate-limit risk.

If your application needs to retrieve more than a few hundred items with quickSearch: false, or if you're using a higher requestDelaySecs, you should use the asynchronous POST /v2/acts/<actor>/runs endpoint. This allows you to start the run, get an immediate runId, and then poll for its completion or use webhooks.

import apify_client

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
actor_id = "crawlerbros/tiktok-ads-library-scraper-pro"

# This run is likely to exceed 300 seconds if maxItems is large and quickSearch is false
run_input = {
    "query": "sustainable fashion",
    "quickSearch": False, # Each ad needs a detail page fetch
    "maxItems": 5000,     # Many items
    "requestDelaySecs": 2 # Slightly slower requests
}

# Start the run asynchronously
# Note: client.actor(actor_id).call() defaults to asynchronous if needed,
# but it's good practice to understand the underlying mechanism for long runs.
run = client.actor(actor_id).call(run_input=run_input)
print(f"Run started with ID: {run['id']}")

# Now, poll for the run's status
run_id = run["id"]
while True:
    run_info = client.run(run_id).get()
    if run_info["status"] == "SUCCEEDED":
        print(f"Run {run_id} finished successfully.")
        break
    elif run_info["status"] in ["FAILED", "ABORTED"]:
        print(f"Run {run_id} failed or was aborted.")
        break
    print(f"Run {run_id} is still {run_info['status']}... Waiting.")
    # In a real application, implement a sensible sleep/retry strategy
    import time
    time.sleep(10)

# Once finished, you can access the dataset
dataset = client.dataset(run_info["defaultDatasetId"])
# ... process data ...
Enter fullscreen mode Exit fullscreen mode

This asynchronous approach prevents your client from timing out while the Actor diligently scrapes in the background. It's a fundamental pattern for any data extraction job that might take longer than a few minutes.

What Are the Cost Implications of Input Parameters?

Understanding the cost structure of tiktok-ads-library-scraper-pro requires looking at its "PAY_PER_EVENT" model. The Actor's primary charge event is "result" (apify-default-dataset-item), costing $0.002 per event. There's also an "Actor Start" event, billed at $0.005 per GB of memory allocated to the run, which is a flat per-event charge for run initiation. On top of these, you also pay for Apify platform usage, billed separately.

The maxItems input field directly multiplies the number of "result" events. Each ad emitted into the default dataset incurs this per-item cost. Therefore, setting maxItems to a high number will proportionally increase the event-based cost.

Similarly, quickSearch: false effectively doubles the amount of work the Actor does per ad, by making an additional detail page request. While the pricing model doesn't explicitly state a separate charge for "detail page fetches," it is reasonable to infer that the increased complexity and resources consumed by quickSearch: false contribute to the overall platform usage, even if the "result" event count remains tied to maxItems.

Consider the tiers for the "result" event: FREE $0.002, BRONZE $0.00167, SILVER $0.00133, GOLD $0.001, PLATINUM $0.001, DIAMOND $0.001. These indicate that unit costs decrease for users in higher Apify tiers. However, the event count itself is driven by your input parameters.

To manage costs effectively:

  • Use maxItems conservatively, especially during development and testing.
  • Only set quickSearch: false when you genuinely need the full targeting breakdown, as it increases resource consumption and thus platform usage.
  • Monitor your run's item count in the Apify Console.
# A simple representation of how input choices affect event counts
def estimate_event_cost(max_items, quick_search_enabled):
    # This is a simplified model for illustration.
    # Actual platform usage adds to this, and specific API behavior might vary.
    result_event_price = 0.002 # Using the FREE tier price for example
    actor_start_cost_per_gb = 0.005 # per GB

    results_cost = max_items * result_event_price
    # Assume 1GB memory for simplicity, so one Actor Start event.
    start_cost = 1 * actor_start_cost_per_gb

    # If quickSearch is false, it uses more resources, increasing platform usage,
    # though not necessarily directly more "result" events.
    # The README implies resource consumption, which is tied to platform usage.
    if not quick_search_enabled:
        print("Note: quickSearch: false increases resource consumption (platform usage).")

    total_event_cost = results_cost + start_cost
    return total_event_cost

# Example cost estimation:
# Scenario 1: Quick search, few items
print(f"Estimated event cost for 50 items (quickSearch=true): ${estimate_event_cost(50, True):.4f}")

# Scenario 2: Detailed search, more items
print(f"Estimated event cost for 1000 items (quickSearch=false): ${estimate_event_cost(1000, False):.4f}")
Enter fullscreen mode Exit fullscreen mode

This demonstrates how maxItems directly drives result event costs, while quickSearch impacts the resource-based platform usage component of the total bill.

Can a Single Request Queue Fan-Out to Multiple Runs?

A common pattern for distributing workloads is to have multiple workers pull from a single queue. However, Apify's RequestQueue has a specific constraint: "A request queue can only be PROCESSED by one Actor or task run at a time, though multiple runs may add to it." This means you cannot fan out processing from a single shared RequestQueue across multiple concurrent tiktok-ads-library-scraper-pro runs.

If your use case requires processing segments of a large ad list concurrently (e.g., scraping specific ad IDs across different regions in parallel), you would need to:

  1. Divide your adIds list into separate chunks.
  2. Start distinct tiktok-ads-library-scraper-pro runs, each with its own segment of adIds in the input. Each run would then manage its own implicit RequestQueue internally.
  3. Collect results from each run's separate default dataset.

Attempting to configure multiple Actor runs to pull from the same explicit RequestQueue for processing will lead to runs waiting for exclusive access, negating the benefits of parallelization.

# Example of splitting ad IDs for parallel runs (conceptual Python)
def trigger_parallel_ad_id_scrapes(ad_ids_list, client):
    chunk_size = 100 # Example chunk size
    chunks = [ad_ids_list[i:i + chunk_size] for i in range(0, len(ad_ids_list), chunk_size)]

    run_ids = []
    for i, chunk in enumerate(chunks):
        run_input = {
            "adIds": chunk,
            "quickSearch": False,
            "region": "all" # Or make this dynamic per chunk
        }
        print(f"Starting run {i+1} with {len(chunk)} ad IDs.")
        run = client.actor("crawlerbros/tiktok-ads-library-scraper-pro").call(run_input=run_input)
        run_ids.append(run["id"])
    return run_ids

# Example usage:
# all_ad_ids = ["id1", "id2", ..., "idN"] # Imagine a large list
# client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
# started_runs = trigger_parallel_ad_id_scrapes(all_ad_ids, client)
# print(f"Started {len(started_runs)} parallel runs: {started_runs}")
# You would then monitor and collect results from each of these runs independently.
Enter fullscreen mode Exit fullscreen mode

This pattern ensures each run operates with its own distinct input, effectively circumventing the single-processor constraint of a shared RequestQueue.

What Happens to Unnamed Storages and How Does It Affect Data Retention?

When you initiate a run of tiktok-ads-library-scraper-pro without explicitly creating and naming a Dataset or KeyValueStore, the platform uses unnamed storages. A critical limitation for data retention is that "unnamed storages expire." Specifically, on the Free plan, "only the 10 most recent runs are retained, for 4 months." After this, your data can be deleted.

For any production-grade data pipeline built around tiktok-ads-library-scraper-pro, you absolutely must ensure your datasets are explicitly named or immediately fetched and stored elsewhere. If you rely on the default, unnamed datasets, you risk losing valuable historical ad data, especially if you run the Actor frequently (pushing older runs out of the "10 most recent" window) or if you need data older than four months.

This is particularly important for competitive intelligence, compliance monitoring, or market sizing use cases, where historical trend analysis is key.

import apify_client
import json

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")
actor_id = "crawlerbros/tiktok-ads-library-scraper-pro"

# Define a named dataset for long-term storage
dataset_name = "tiktok-ads-library-scraper-pro-my-project-ads-2023-10-09"
dataset = client.datasets().get_or_create(name=dataset_name)

run_input = {
    "query": "electric vehicles",
    "region": "all",
    "quickSearch": false,
    "maxItems": 500
}

# When calling the Actor, specify the output dataset to be the named one
# The ApifyClient library handles passing this internally when you call the actor
run = client.actor(actor_id).call(run_input=run_input, dataset_id=dataset["id"])

print(f"Run started. Results will be stored in named dataset: {dataset['name']} (ID: {dataset['id']})")
print(f"Run ID: {run['id']}")

# After the run completes, you can fetch from the named dataset
# For simplicity, we'll just show fetching from the named dataset
# In a real scenario, you'd wait for the run to complete as shown in the async example
# items = client.dataset(dataset["id"]).iterate_items()
# for item in items:
#     print(json.dumps(item, indent=2))
Enter fullscreen mode Exit fullscreen mode

By explicitly creating and linking to a named dataset, you guarantee data persistence, regardless of how many subsequent runs occur or how old the data becomes. This is a fundamental safeguard against accidental data loss.

Limitations and Caveats

While tiktok-ads-library-scraper-pro is a powerful tool, it operates within specific boundaries that developers must acknowledge:

  • Region Coverage is Limited: The transparency library, by design, covers EU/EEA countries, UK, Switzerland, and Turkey due to DSA mandates. If your use case requires data from other global regions (e.g., US, Canada, APAC), this Actor will not fulfill that requirement. The region input field explicitly lists these supported regions, and attempting to specify an unsupported region will likely yield no results or an error.
  • Impression Data is Bucketed: TikTok provides impression data in ranges like 1K-10K or 1M-10M, not exact counts. The minImpressions Pro filter operates on these parsed numeric values. This means granular impression analysis requiring precise numbers is not possible; your metrics will need to adapt to these bucketed ranges.
  • No Native AWS S3 or Slack Integration: For integrating results into broader data ecosystems, remember that the Apify platform does not offer native AWS S3 or Slack integrations. You will need to route data through webhooks, or use external tools like n8n, Make, or Zapier to connect the Actor's output to these services. This adds a layer of orchestration if your existing stack heavily relies on these services.
  • Schedules are Disabled by Default: If you plan to automate tiktok-ads-library-scraper-pro runs with schedules, be aware that new schedules are created in a DISABLED state. They also require the Actor to have run at least once before they can be scheduled. This means an extra step of enabling the schedule and an initial manual (or API-triggered) run are necessary during setup.

These points highlight areas where the Actor's capabilities are constrained by its source data or platform design, requiring forethought in your integration strategy.

Checked against the Actor's input schema and Apify docs on 2026-10-09.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

Top comments (0)