Why n8n Workflows Drop Idealista Data When Runs Exceed 300 Seconds
When orchestrating web scrapers inside a production data pipeline, developers often default to synchronous webhook patterns. You configure an HTTP request node in n8n, point it to the Apify API run-sync endpoint, and wait for the payload to return. This simple pattern breaks down when scraping real estate listings from platforms like Idealista.
The idealista-scraper is designed to handle pagination, mirror-first routing, and proxy fallback strategies to safely extract listing data across Spain, Italy, and Portugal. However, these operations take time.
The Apify synchronous run endpoint enforces a hard timeout cap of 300 seconds. If the scraper takes longer than 5 minutes to paginate through listing cards, the API gateway terminates the client connection and returns an HTTP 408 status code. This causes your downstream n8n workflow or pandas pipeline to fail with a timeout error, even though the scraping container is still running in the background.
To build a production-grade real estate pipeline, you must shift from a synchronous wait-and-return architecture to an event-driven, asynchronous orchestration pattern. This article explains how to configure your runs, handle platform limits, manage state across geographic targets, and structure your pipelines to prevent data loss.
How Do You Avoid the 300 Second Synchronous Run Timeout?
You can avoid the 300-second synchronous timeout by initiating an asynchronous run via a POST request and using Apify's native webhook system or the n8n Trigger node to handle data delivery. Instead of waiting for the API response, your orchestrator starts the run and immediately frees up execution threads. Downstream systems only process the data once the run transitions to a finished state.
Here is the curl request to initiate an asynchronous run. Notice that we do not use the /run-sync endpoint, but rather the standard POST /runs endpoint which returns a run ID immediately:
curl --request POST \
--url "https://api.apify.com/v2/acts/crawlerbros~idealista-scraper/runs" \
--header "Content-Type: application/json" \
--header "Authorization: Bearer <YOUR_APIFY_TOKEN>" \
--data '{
"location": "barcelona-barcelona",
"operation": "sale",
"propertyType": "homes",
"country": "es",
"maxItems": 500,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}'
This request returns a JSON payload containing the id of the run. Your pipeline can store this run ID and poll for status changes, or better yet, rely on a webhook triggered directly by the platform when the run finishes.
What is the Correct Way to Configure n8n for Long Runs?
The correct way to configure n8n for long-running scrapers is to use the dedicated Apify Trigger node rather than generic HTTP Request nodes. The n8n Trigger node registers a webhook on the Apify platform that listens for run completion events, which avoids polling and prevents execution timeouts in your workflow.
If you are running n8n self-hosted, you cannot use OAuth2 credentials, but you can authenticate using your Apify API key. Here is a declarative JSON representation of an n8n workflow designed to trigger the idealista-scraper and process the resulting dataset asynchronously:
{
"nodes": [
{
"parameters": {},
"id": "trigger-node-id",
"name": "On Scrape Completed",
"type": "n8n-nodes-base.apifyTrigger",
"typeVersion": 1,
"position": [250, 300],
"credentials": {
"apifyApi": {
"id": "credential-id",
"name": "Apify API Token"
}
}
},
{
"parameters": {
"url": "={{ $json.body.resource.defaultDatasetId }}",
"options": {}
},
"id": "fetch-dataset-node",
"name": "Fetch Scraped Properties",
"type": "n8n-nodes-base.httpRequest",
"typeVersion": 4,
"position": [480, 300]
}
],
"connections": {
"On Scrape Completed": {
"main": [
[
{
"node": "Fetch Scraped Properties",
"type": "main",
"index": 0
}
]
]
}
}
}
This node structure guarantees that n8n remains idle while the scraper navigates through search results. Only when the run finishes does the webhook send the metadata, including the defaultDatasetId, directly to the next stage of your pipeline. Note that n8n trigger nodes utilize Apify webhooks, which support exactly one action: POSTing to a URL. This is the only event-driven primitive supported natively by the platform.
Why Do API Runs Ignore Your Console Prefill Values?
API runs ignore UI prefill values because prefill configurations are purely presentational cues for the Apify Console user interface, whereas the API runtime strictly evaluates input defaults from the schema. If you trigger a run via the API or an external orchestrator without passing a payload, only the properties defined as default in the schema are evaluated.
Checked against the Actor's input schema and Apify docs on 2026-10-03.
When calling the API, you must always explicitly construct and pass the entire input payload. If you rely on the platform UI prefill values, your API-initiated runs will fall back to the default schema settings (such as "location": "madrid-madrid", "operation": "sale", "propertyType": "homes", and "country": "es").
Here is a Python integration utilizing the official client to safely execute the scraper with a fully defined input payload:
from apify_client import ApifyClient
# Explicitly defining the input schema payload to bypass UI dependency issues
run_input = {
"location": "lisboa",
"operation": "rent",
"propertyType": "homes",
"country": "pt",
"maxItems": 150,
"proxyConfiguration": {
"useApifyProxy": True,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}
client = ApifyClient("YOUR_APIFY_API_TOKEN")
# Initiate the actor and wait asynchronously without relying on platform defaults
actor_run = client.actor("crawlerbros/idealista-scraper").call(run_input=run_input)
# Retrieve the dataset items from the default storage location
dataset_items = client.dataset(actor_run["defaultDatasetId"]).list_items().items
print(f"Scraped {len(dataset_items)} property records from Lisbon.")
How Do You Prevent Double Billing on Pipeline Re Runs?
To prevent double billing when a downstream pipeline step fails, you must isolate your scraping step from your data transformations and leverage stateful processing. Re-running a full pipeline that launches a new scraper execution will trigger new event charges and consume platform resources. Instead, check for existing, valid dataset runs before starting a new one.
Since unnamed storages on the free plan are retained only for the 10 most recent runs over a 4-month window, you should name your datasets if you want to reuse them across pipeline retries. Named storages are always exempt from automatic deletion.
Here is a JavaScript strategy that checks if a named dataset containing today's listings exists before initializing a costly new run:
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_API_TOKEN' });
async function getOrRunScraper(locationSlug) {
const todayStr = new Date().toISOString().slice(0, 10);
const datasetName = `idealista-scrape-${locationSlug}-${todayStr}`;
try {
// Attempt to access an existing named dataset from a previous run
const existingDataset = await client.dataset(datasetName).get();
if (existingDataset) {
console.log(`Found cached dataset: ${datasetName}. Skipping scraping run.`);
const { items } = await client.dataset(datasetName).listItems();
return items;
}
} catch (error) {
// A 404 error means the named dataset does not exist yet
console.log(`No cached data found for ${todayStr}. Starting scraper...`);
}
// Run the scraper and target a named dataset for storage
const run = await client.actor('crawlerbros/idealista-scraper').call({
location: locationSlug,
operation: 'sale',
propertyType: 'homes',
country: 'es',
maxItems: 100
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const targetDataset = await client.datasets().getOrCreate(datasetName);
await client.dataset(targetDataset.id).pushItems(items);
return items;
}
How Do You Parse and Map the Output Schema to a Database?
When building local real estate databases, you must map the flat JSON listings emitted by the scraper into relational tables. The scraper returns reliable core fields, location metrics, and media flags from the index cards, but leaves out detail-only fields like coordinates or deep media configurations. Your database migration must handle nullable boolean flags, extract nested objects like the parking structure, and normalize localized floor strings without throwing SQL exceptions during bulk insertions.
The nested hasParkingSpace object contains two boolean flags that describe whether parking is available and if it is included in the price. Similarly, the contactInfo object nests agency details. To store this efficiently, you can flatten these nested structures into a single database row. The floor field often contains descriptive text such as "4ยช planta exterior con ascensor", which you can parse to extract the structural floor number and elevator status.
Below is a complete Python pipeline using SQLAlchemy to define a schema, flatten the scraper's nested JSON output, and upsert listings into a relational SQLite database:
import json
import re
from sqlalchemy import create_engine, Column, String, Integer, Boolean, DateTime
from sqlalchemy.orm import declarative_base, sessionmaker
Base = declarative_base()
class PropertyListing(Base):
__tablename__ = "properties"
property_code = Column(String, primary_key=True)
url = Column(String, nullable=False)
price = Column(Integer, nullable=False)
price_by_area = Column(Integer)
currency = Column(String, default="EUR")
size = Column(Integer)
rooms = Column(Integer)
bathrooms = Column(Integer)
floor_raw = Column(String)
floor_level = Column(Integer) # Parsed floor level
exterior = Column(Boolean)
description = Column(String)
address = Column(String)
municipality = Column(String)
district = Column(String)
country = Column(String)
thumbnail = Column(String)
num_photos = Column(Integer)
has_lift = Column(Boolean)
has_terrace = Column(Boolean)
has_swimming_pool = Column(Boolean)
has_air_conditioning = Column(Boolean)
has_parking_space = Column(Boolean)
parking_included = Column(Boolean)
agency_name = Column(String)
property_type = Column(String)
operation = Column(String)
new_development = Column(Boolean)
scraped_at = Column(String)
def parse_floor(floor_str):
if not floor_str:
return None
match = re.search(r"(\d+)\ยช", floor_str)
return int(match.group(1)) if match else None
def map_and_save_listings(raw_json_data, db_url="sqlite:///idealista.db"):
engine = create_engine(db_url)
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine)
session = Session()
listings = json.loads(raw_json_data)
for item in listings:
parking = item.get("hasParkingSpace") or {}
contact = item.get("contactInfo") or {}
listing_record = PropertyListing(
property_code=item.get("propertyCode"),
url=item.get("url"),
price=item.get("price"),
price_by_area=item.get("priceByArea"),
currency=item.get("currency", "EUR"),
size=item.get("size"),
rooms=item.get("rooms"),
bathrooms=item.get("bathrooms"),
floor_raw=item.get("floor"),
floor_level=parse_floor(item.get("floor")),
exterior=item.get("exterior"),
description=item.get("description"),
address=item.get("address"),
municipality=item.get("municipality"),
district=item.get("district"),
country=item.get("country"),
thumbnail=item.get("thumbnail"),
num_photos=item.get("numPhotos"),
has_lift=item.get("hasLift"),
has_terrace=item.get("hasTerrace"),
has_swimming_pool=item.get("hasSwimmingPool"),
has_air_conditioning=item.get("hasAirConditioning"),
has_parking_space=parking.get("hasParkingSpace"),
parking_included=parking.get("isParkingSpaceIncludedInPrice"),
agency_name=contact.get("commercialName"),
property_type=item.get("propertyType"),
operation=item.get("operation"),
new_development=item.get("newDevelopment"),
scraped_at=item.get("scrapedAt")
)
session.merge(listing_record)
session.commit()
session.close()
Designing a State Machine in Python to Manage Scrape State
When constructing long-running real estate ETL pipelines, running your ingestion inside a robust state machine ensures that network disruptions do not force you to restart your scrapers. Idealista enforces a strict limit of 1,800 listings per search query (60 pages). To scrape larger areas, you must split your target geographic regions into smaller location slugs and loop through them.
If your script crashes halfway through processing 10 different municipal areas, you should only resume scraping the remaining areas.
The following Python state machine uses local JSON files to manage state across runs and prevents duplicate execution of the scraping phase:
import json
import os
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_API_TOKEN")
STATE_FILE = "pipeline_state.json"
TARGET_LOCATIONS = ["madrid-madrid", "barcelona-barcelona", "valencia-valencia", "sevilla-sevilla"]
def load_state():
if os.path.exists(STATE_FILE):
with open(STATE_FILE, "r") as f:
return json.load(f)
return {"completed_locations": {}, "results_count": 0}
def save_state(state):
with open(STATE_FILE, "w") as f:
json.dump(state, f, indent=4)
def run_geographical_pipeline():
state = load_state()
for location in TARGET_LOCATIONS:
if location in state["completed_locations"]:
print(f"Location {location} already processed. Skipping.")
continue
print(f"Scraping listings for: {location}")
try:
# Run the scraper for a single district
run = client.actor("crawlerbros/idealista-scraper").call(run_input={
"location": location,
"operation": "sale",
"propertyType": "homes",
"country": "es",
"maxItems": 100
})
# Store the dataset ID associated with this specific location
state["completed_locations"][location] = {
"dataset_id": run["defaultDatasetId"],
"status": "success"
}
save_state(state)
except Exception as e:
print(f"Failed to scrape {location}: {str(e)}")
state["completed_locations"][location] = {
"dataset_id": None,
"status": "failed",
"error": str(e)
}
save_state(state)
break # Break execution to allow debugging without losing current progress
if __name__ == "__main__":
run_geographical_pipeline()
Handling Location Input Failures and Empty Results
When scraping Idealista, passing an incorrect location slug is the most common reason for a run returning zero results. For instance, if you pass "rome" instead of "roma-roma", or "lisbon" instead of "lisboa", the platform will fail to match the query and yield zero results.
To prevent silent failures where your downstream pipeline processes empty datasets, your pipeline code must validate the run status message and dataset item count. The scraper returns specific status strings depending on why zero items were found.
Here is a Node.js script that demonstrates how to inspect a finished run, check the status message, and handle failures gracefully before downloading data:
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_API_TOKEN' });
async function validateAndFetchProperties(runId) {
const run = await client.run(runId).get();
// Check if the run completed but found nothing
const datasetInfo = await client.dataset(run.defaultDatasetId).get();
if (datasetInfo.itemCount === 0) {
const statusMessage = run.statusMessage || '';
console.warn(`Warning: Run ${runId} returned 0 listings.`);
if (statusMessage.includes('No listings extracted')) {
throw new Error('Scraping failed: Invalid location slug. Verify against regional names.');
} else if (statusMessage.includes('All attempts blocked')) {
throw new Error('Scraping failed: The target site blocked all proxy attempts. Retry later.');
} else if (statusMessage.includes('No URLs to process')) {
throw new Error('Scraping failed: Direct URLs provided did not contain search index cards.');
} else {
throw new Error(`Scraping failed with message: ${statusMessage}`);
}
}
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Successfully fetched ${items.length} validated items.`);
return items;
}
What Are the Real World Limitations of the Scraper?
Understanding the physical boundaries of the crawler prevents pipeline errors down the road. The scraper operates on a mirror-first cloud routing approach, falling back to proxied requests when direct connections are blocked.
First, Idealista imposes a platform-level cap of approximately 1,800 listings (60 pages) per search query. If your target is a large metropolitan area, the actor cannot bypass this cap because the underlying site will not generate pagination buttons past page 60. You must segment your searches into smaller districts or neighborhoods using multiple location slugs or startUrls targets.
Second, this actor extracts data visible on the search results index pages. Highly specific fields such as coordinate attributes (latitude, longitude), province lists, and deep media tours are not exposed on search cards. If your downstream database requires precise coordinates or full photo galleries, you cannot rely on this search-card output schema alone.
Third, the platform enforces storage limits of 400 requests per second for dataset item pushes and request queue operations. If you attempt to process multiple runs in parallel using a single shared queue, your run will fail. A request queue can only be processed by one actor run at a time.
Fourth, when relying on proxies for mirror bypasses, be aware that datacenter proxy sessions persist for 26 hours, whereas residential proxy sessions rotate every 30 minutes. If your run takes longer than 30 minutes, your scraper must handle proxy rotation without dropping connections.
Finally, schedules on the platform are created disabled by default. Also, the actor must have run successfully at least once before you can assign it to a schedule. If you try to deploy a schedule programmatically without an initial manual run, the scheduler will reject the request.
How Do the Event Charges Scale on Pay-Per-Event Billing?
The cost structure of the Idealista Scraper is governed by a PAY_PER_EVENT billing model, combined with standard platform usage costs. This means your financial projection scales with the scale of the scrape, not just execution time.
On top of the event prices, you also pay the Apify platform usage the run consumes, billed separately at your Apify plan's rates. The events themselves scale directly with the size of your requests.
The charged events for this actor are:
-
Actor Start (
apify-actor-start): Charged once per run at a flat rate of $0.05 per GB of memory allocated to the run. Number of events charged depends on Actor memory (one event per GB, minimum one event). -
result (
apify-default-dataset-item): Charged at $0.002 per event. Each single property listing successfully saved to your default dataset represents one result event.
The per-event price for results scales down based on your Apify user discount tier. The enterprise PLATINUM and DIAMOND tiers, along with the GOLD tier, lower the price to $0.001 per result.
Here is the pricing structure across different platform tiers:
| Apify Discount Tier | Result Event Price (USD) |
|---|---|
| FREE | $0.002 |
| BRONZE | $0.00167 |
| SILVER | $0.00133 |
| GOLD | $0.001 |
| PLATINUM | $0.001 |
| DIAMOND | $0.001 |
The event part of the cost scales with the number of events a run emits, which corresponds to the scraped property listings, while the platform-usage part scales with the resources the run consumes. Platform usage costs scale primarily with CPU allocation and memory consumption, meaning larger allocations of memory to handle fast concurrent scraping requests will increase your separate platform bill even if the event prices remain constant.
To control your financial exposure, you can append the query parameter maxTotalChargeUsd to your run initiation API endpoints. This query parameter maps to the internal environment variable ACTOR_MAX_TOTAL_CHARGE_USD inside the running container. When this limit is reached, the run terminates gracefully, giving the system a 30-second window to finish saving current items.
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)