DEV Community

Cover image for Why Handling Apify Run Timeouts Prevents Double-Processing
Crawler Bros
Crawler Bros

Posted on

Why Handling Apify Run Timeouts Prevents Double-Processing

Orchestrating Apify Actor Runs into External Systems

Building robust data pipelines often involves more than just running a scraper; it's about reliably moving the extracted data into downstream systems, whether that's a webhook, a data warehouse, or a local Pandas script. A common challenge arises when orchestrating these pipelines: how do you ensure data integrity, particularly preventing duplicate processing when a run needs to be re-executed, potentially covering data already handled?

This article focuses on integrating the Website Image Scraper Actor, not as a standalone tool, but as a component within a larger data flow. We'll explore strategies for managing state to avoid double-processing and ensure that your downstream applications, like an n8n workflow or a data warehouse ingest, receive only new or unique data, even when Actor runs are re-triggered.

What happens if an Actor run times out?

Apify's synchronous run endpoint has a hard cap of 300 seconds. If a run exceeds this 5-minute limit, it returns an HTTP 408 error. This doesn't necessarily mean the Actor failed or produced no data; it simply means the synchronous response channel timed out. The run might still be processing in the background, pushing results to its dataset.

When orchestrating a pipeline, relying solely on a synchronous response is risky for longer-running jobs. Instead, for anything potentially exceeding this limit, you should initiate the run asynchronously via the /v2/acts/<actor>/runs endpoint and then either poll for its status or, more efficiently, use webhooks. If you're building a pipeline that expects a complete dataset upon a run's return, a 408 can lead to an incomplete picture, forcing you to re-run and introducing the problem of handling existing data.

How can you prevent double-processing on re-runs?

To prevent double-processing when re-running an Actor, you need a mechanism to track which data items have already been successfully processed by your downstream system. The Website Image Scraper Actor, like many Apify Actors, deduplicates results by absolute URL within a single run. However, it doesn't track across different runs.

The key is to maintain an external record of processed url fields from the Actor's output. Before pushing a new batch of urls to your n8n workflow or data warehouse, check them against your historical record. This can be a simple database table, a persistent key-value store, or even a file in cloud storage. Any image url already in your record is skipped, ensuring only new images are processed. This strategy is essential for idempotency, allowing you to re-run the Actor safely without adverse effects on your downstream systems.

What is the output shape of the Website Image Scraper?

The Website Image Scraper produces one record per unique image URL. Each record includes granular details about the image's discovery and properties. Understanding this output shape is crucial for effective downstream processing and for designing your state management.

Here's an example of a single output record:

{
  "url": "https://apify.com/static/hero.jpg",
  "sourcePage": "https://apify.com/",
  "pageTitle": "Apify ยท The full-stack web-scraping & automation platform",
  "alt": "Apify hero image",
  "hasAltText": true,
  "title": "Apify",
  "width": 1200,
  "height": 600,
  "extension": "jpg",
  "discoveredVia": "img-tag",
  "mimeTypeHint": "image/jpeg",
  "crawlDepth": 0,
  "scrapedAt": "2024-12-16T14:23:11+00:00"
}
Enter fullscreen mode Exit fullscreen mode

The url field is paramount for deduplication across runs. Other fields like sourcePage, pageTitle, alt, and extension provide rich metadata that can be used for content audits, asset inventories, or SEO analysis. For instance, filtering by hasAltText: false can quickly flag accessibility issues.

Integrating with n8n via webhooks

For event-driven workflows, n8n's Apify Trigger node is ideal. It fires upon Actor run completion, eliminating the need for manual polling. However, connecting this to a deduplication strategy requires careful thought.

Instead of directly processing every record from the Apify dataset, your n8n workflow should perform a lookup. Here's a conceptual flow:

  1. Apify Trigger: An n8n "Apify Trigger" node initiates when the Website Image Scraper run completes.
  2. Fetch Results: Use an Apify node in n8n to fetch the dataset items from the completed run.
  3. Deduplication Logic: For each url in the fetched dataset, query your external state store (e.g., a database, Google Sheets, or another n8n workflow that manages a persistent list).
    • If the url already exists, discard the item.
    • If the url is new, process it (e.g., send to Slack, push to a data warehouse) and then add it to your external state store.

This ensures that even if the Website Image Scraper run is re-triggered and returns some previously seen image URLs, n8n only acts on the truly new ones.

Here's a simplified Python example showing how you might interact with the Apify API to fetch results for processing in an external system, assuming you've already handled the run execution and have a dataset_id:

import apify_client
import os

# Initialize the ApifyClient with your API token
# Set it as an environment variable or replace 'YOUR_APIFY_API_TOKEN'
client = apify_client.ApifyClient(os.environ.get("APIFY_API_TOKEN"))

# Replace with the actual dataset ID from your completed run
# This would typically come from your webhook payload or polling result
dataset_id = "YOUR_DATASET_ID" 

print(f"Fetching items from dataset: {dataset_id}")

all_images = []
# Iterate over all items in the dataset
for item in client.dataset(dataset_id).iterate_items():
    all_images.append(item)

print(f"Found {len(all_images)} image records.")

# Example of how you might then process these,
# assuming a simple set for tracking already processed URLs
processed_urls_db = set() # In a real scenario, this would be a persistent database

new_images_to_process = []
for image_record in all_images:
    image_url = image_record.get("url")
    if image_url and image_url not in processed_urls_db:
        new_images_to_process.append(image_record)
        # In a real scenario, you'd add image_url to your persistent DB *after* successful downstream processing
        # processed_urls_db.add(image_url) 

print(f"Identified {len(new_images_to_process)} new images for downstream processing.")

# Now, 'new_images_to_process' can be sent to your n8n webhook,
# inserted into a database, or processed by Pandas.
# For a webhook: POST each record to your n8n receive URL
# For a database: batch insert into your images table
Enter fullscreen mode Exit fullscreen mode

Using persistent storage for state management

To reliably manage state across Actor runs, particularly on the Apify platform, you need to use named storages. Unnamed storages, which are the default for Actor runs, expire and are only retained for a limited time (e.g., the 10 most recent runs for 4 months on the free plan). This makes them unsuitable for long-term deduplication.

Instead, create a named dataset or key-value store specifically for tracking processed image URLs. You can then configure your Website Image Scraper Actor to either:

  1. Read from a named key-value store at the start of each run to load previously processed URLs into memory for in-Actor filtering (if your dataset isn't too large).
  2. Post-process the results by reading the full dataset and then comparing against a named key-value store or a dedicated database table in your external system (as outlined in the n8n example). This is often more scalable, as the Actor itself doesn't need to hold the full state.

For instance, you could store a set of all previously processed image URLs in a JSON file within a named key-value store. Before processing new results, you fetch this file, update your set, and then upload the updated set.

import apify_client
import os
import json

client = apify_client.ApifyClient(os.environ.get("APIFY_API_TOKEN"))

# Define a named key-value store for tracking processed URLs
processed_urls_kv_store_name = "website-image-scraper-processed-urls"
processed_urls_key = "processed-image-urls"

# --- Fetch existing processed URLs ---
try:
    processed_urls_record = client.key_value_store(processed_urls_kv_store_name).get_record(processed_urls_key)
    if processed_urls_record:
        existing_processed_urls = set(json.loads(processed_urls_record["value"]))
        print(f"Loaded {len(existing_processed_urls)} existing processed URLs.")
    else:
        existing_processed_urls = set()
        print("No existing processed URLs found.")
except Exception as e:
    print(f"Error loading processed URLs from KV store: {e}. Starting with empty set.")
    existing_processed_urls = set()

# --- Simulate new image data from an Actor run (replace with actual dataset fetch) ---
# In a real pipeline, this would come from `client.dataset(dataset_id).iterate_items()`
new_image_data_simulated = [
    {"url": "https://example.com/image1.jpg", "sourcePage": "https://example.com/page1"},
    {"url": "https://example.com/image2.png", "sourcePage": "https://example.com/page1"},
    {"url": "https://example.com/image1.jpg", "sourcePage": "https://example.com/page2"}, # Duplicate
    {"url": "https://example.com/image3.webp", "sourcePage": "https://example.com/page3"}
]

# Add some 'old' URLs to simulate previous runs
existing_processed_urls.add("https://example.com/image1.jpg")
print(f"Simulating existing processed URLs: {existing_processed_urls}")

# --- Process new data, filtering out already processed URLs ---
newly_found_urls = set()
for item in new_image_data_simulated:
    url = item.get("url")
    if url:
        newly_found_urls.add(url)

urls_to_process_downstream = []
for url in newly_found_urls:
    if url not in existing_processed_urls:
        urls_to_process_downstream.append(url)
        # Mark as processed immediately for this run, before sending downstream
        # In production, this update might happen *after* successful downstream delivery
        existing_processed_urls.add(url)

print(f"URLs to send downstream: {urls_to_process_downstream}")

# --- Update the named key-value store with the new set of all processed URLs ---
client.key_value_store(processed_urls_kv_store_name).set_record(
    key=processed_urls_key,
    value=json.dumps(list(existing_processed_urls)),
    content_type="application/json"
)
print(f"Updated KV store '{processed_urls_kv_store_name}' with {len(existing_processed_urls)} processed URLs.")
Enter fullscreen mode Exit fullscreen mode

Integrating into a Pandas DataFrame for analysis

For data engineers and analysts, pulling the Website Image Scraper's output directly into a Pandas DataFrame is a common first step. This allows for powerful local analysis, aggregation, and transformation before loading into a data warehouse.

The workflow involves:

  1. Running the Actor: Initiate the website-image-scraper Actor.
  2. Fetching Results: Once the run completes, fetch the dataset items.
  3. Loading into Pandas: Convert the list of dictionaries into a DataFrame.
  4. Deduplication: Leverage Pandas' capabilities for deduplication. This is where your external state management becomes critical.

Let's assume you have a database table processed_images with a url column.

import apify_client
import os
import pandas as pd
# In a real application, you'd use a database connector like psycopg2 or sqlalchemy
# from your_db_connector import fetch_processed_urls, mark_urls_as_processed

client = apify_client.ApifyClient(os.environ.get("APIFY_API_TOKEN"))

# Example: Run the Actor (replace with your actual run logic if not hardcoding)
# This snippet assumes a run has already completed and we have its ID
# To start a run:
# run_input = {
#     "startUrl": "https://dev.to/",
#     "maxCrawlDepth": 1,
#     "maxTotalImages": 500
# }
# run = client.actor("crawlerbros/website-image-scraper").call(run_input=run_input)
# dataset_id = run["defaultDatasetId"]

# For demonstration, let's use a placeholder dataset ID
dataset_id = "YOUR_ACTOR_RUN_DATASET_ID" # Replace with a real dataset ID

print(f"Fetching data from dataset ID: {dataset_id}")
all_records = []
try:
    for item in client.dataset(dataset_id).iterate_items():
        all_records.append(item)
except Exception as e:
    print(f"Could not fetch dataset items. Error: {e}")
    # Handle the error appropriately, e.g., retry or exit
    exit()

if not all_records:
    print("No image records found in the dataset.")
    df = pd.DataFrame()
else:
    df = pd.DataFrame(all_records)
    print(f"DataFrame loaded with {len(df)} records.")

    # In a real scenario, you'd fetch this from your database
    # For simulation, let's create a list of already processed URLs
    # This list would contain URLs from previous successful runs
    simulated_processed_urls_from_db = [
        "https://dev.to/assets/webpack/image.png",
        "https://dev.to/assets/webpack/logo.svg"
    ]
    # For a real DB, you'd do:
    # processed_urls_from_db = fetch_processed_urls() # Returns a list or set of URLs

    # Convert to a set for efficient lookup
    processed_urls_set = set(simulated_processed_urls_from_db)

    # Filter out already processed URLs using Pandas
    initial_count = len(df)
    df_new_images = df[~df['url'].isin(processed_urls_set)].copy() # Use .copy() to avoid SettingWithCopyWarning

    print(f"Original records: {initial_count}")
    print(f"New images after deduplication: {len(df_new_images)}")

    if not df_new_images.empty:
        print("First 5 new images:")
        print(df_new_images[['url', 'sourcePage', 'extension']].head())

        # Now, df_new_images contains only unique URLs not previously processed.
        # You can perform further transformations or load into a data warehouse.
        # After successful ingestion, you'd mark these URLs as processed in your DB.
        # mark_urls_as_processed(df_new_images['url'].tolist())
    else:
        print("No new images found to process.")

# Example of a simple transformation: count images by extension
if not df.empty:
    print("\nImage count by extension (all records):")
    print(df['extension'].value_counts())
Enter fullscreen mode Exit fullscreen mode

This Pandas-based approach provides flexibility. You can add complex filtering, data cleaning, or merge with other datasets before final storage. The crucial part remains the external state management to ensure df_new_images truly represents novel data.

Considerations for scheduling and costs

When wiring the Website Image Scraper into a pipeline, understanding its pricing model and scheduling nuances is vital for cost management and reliable execution.

The Website Image Scraper uses a PAY_PER_EVENT pricing model. This means you pay for specific events emitted by the Actor, plus the general Apify platform usage.
The charged events are:

  • "result" (apify-default-dataset-item): $0.002 per event for each single result record pushed to the default dataset. This price has discount tiers: FREE $0.002, BRONZE $0.00167, SILVER $0.00133, GOLD $0.001, PLATINUM $0.001, DIAMOND $0.001.
  • "Actor Start" (apify-actor-start): $0.005 per GB of memory allocated to the run.

The maxTotalImages input field directly caps the number of "result" events, giving you strong control over the primary cost driver. If maxTotalImages is 1000, you will incur at most 1000 "result" events, plus "Actor Start" events based on memory allocated. Higher maxCrawlDepth values can lead to more pages being crawled, which in turn can hit maxTotalImages faster or simply lead to more "result" events if maxTotalImages is sufficiently high. Disabling includeBackgroundImages could reduce the total number of images found, potentially lowering "result" events.

Schedules on Apify are created DISABLED by default and require the Actor to have run at least once before they can be activated. This is a common gotcha. Ensure your Actor is tested with a manual run before relying on a schedule. The minimum interval for a cron schedule is 10 seconds. For pipelines involving the Website Image Scraper, setting up a daily or weekly schedule is a common pattern for monitoring changes on target websites.

{
  "startUrl": "https://apify.com/store",
  "maxCrawlDepth": 1,
  "maxImagesPerPage": 100,
  "maxTotalImages": 500,
  "imageExtensions": ["jpg", "png", "webp", "svg"],
  "includeBackgroundImages": true
}
Enter fullscreen mode Exit fullscreen mode

This example input will crawl https://apify.com/store, follow internal links one level deep, extract up to 100 images per page, and stop once 500 total images are found across the entire run. This limits the cost from "result" events to at most 500 records at your applicable tier price.

Limitations and caveats

While powerful, the Website Image Scraper has specific limitations that impact pipeline design:

  • HTTP-only: This Actor is HTTP-only, meaning it only parses the server-rendered HTML. It does not execute JavaScript. If a website heavily relies on client-side JavaScript (e.g., React, Vue, Angular SPAs) to lazy-load images or dynamically render image tags, this Actor may miss those images, only seeing placeholders or initial HTML. For such cases, you would need a Playwright-based Actor. The output field discoveredVia gives a hint about the source, but it won't tell you if images were missed entirely due to JS rendering.
  • No image binary download: The Actor only collects URLs and metadata (like alt text, width, height, extension). It explicitly does not download the image binaries. If your downstream pipeline requires the actual image files, you'll need to add a separate step (e.g., another Apify Actor like "URL list" or a custom script) to download them using the collected urls.
  • Internal links only: The maxCrawlDepth parameter only applies to internal links, specifically those on the same host as the startUrl. It will not follow external links. This is by design to keep the crawl scope and cost predictable.
  • Sentinel record on no images: If a crawled page (or entire site within the maxCrawlDepth) yields no images, the Actor will still emit a single sentinel record: {"type": "website_image_scraper_error", "reason": "no_images_found"}. Your pipeline should be prepared to handle this, recognizing it as a successful run that simply found no relevant data.
  • Request Queue limitation: An Apify request queue can only be processed by one Actor or task run at a time. While multiple runs can add to a queue, fan-out across a single shared queue for consumption is not supported. This means if your pipeline involves multiple scraping steps, each consuming from a queue, you'll need separate queues or sequential processing for distinct consumption stages.

Always consider these limitations when designing your pipeline, especially if dealing with complex, JavaScript-heavy sites or if direct image file access is required.

Checked against the Actor's input schema and Apify docs on 2026-10-02.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

Top comments (0)