DEV Community

Cover image for Why Post-Fetch Filtering on LinkedIn Guest Pages Drops Your Dataset Yield
Crawler Bros
Crawler Bros

Posted on

Why Post-Fetch Filtering on LinkedIn Guest Pages Drops Your Dataset Yield

The Illusion of Server-Side Filters in Guest Sessions

Data engineers integrating the linkedin-jobs-scraper often assume that selecting an input parameter means LinkedIn's database executes that query. It does not. Because this Actor operates without a login, cookies, or an API key, it relies entirely on LinkedIn's public-facing guest search pages. Over the last few years, LinkedIn has systematically stripped filtering capabilities from its guest-facing search endpoints. To compensate, the Actor performs post-fetch filtering. This architectural reality introduces major downstream implications for your data pipelines.

When you pass parameters like experienceLevel, jobType, industry, or function, the Actor cannot apply these to the initial search query. Instead, it must fetch the job details first and then discard the records that do not match your criteria. If you set maxItems to 100, and a large portion of the fetched jobs do not match your post-fetch filters, your pipeline will receive far fewer records than expected, despite the Actor scanning a massive volume of pages behind the scenes.

How do you programmatically detect when post-fetch filters discard most of your job dataset?

To detect when post-fetch filters discard too many jobs, compare the final clean dataset item count against your target yield threshold after the run completes. If the final count of items in your default dataset is lower than your business requirements, your orchestrator can trigger an automated alert. This ensures you know immediately when overly restrictive filters or layout changes are starving your database.

By retrieving the run object and the default dataset details through the Apify API client, you can evaluate the volume of delivered records. This programmatic approach ensures you are notified immediately when an overly restrictive filter combination starves your downstream database of fresh job listings.

import os
from apify_client import ApifyClient

def audit_dataset_yield(run_id: str, expected_minimum: int = 10):
    client = ApifyClient(os.environ.get("APIFY_TOKEN"))

    run_details = client.run(run_id).get()
    dataset_id = run_details.get("defaultDatasetId")

    dataset_info = client.dataset(dataset_id).get()
    clean_item_count = dataset_info.get("cleanItemCount", 0)

    if clean_item_count < expected_minimum:
        print(f"Warning: Run {run_id} yielded only {clean_item_count} items.")
    else:
        print(f"Run {run_id} yielded {clean_item_count} items, meeting the minimum of {expected_minimum}.")
Enter fullscreen mode Exit fullscreen mode

The Structural Speed and Payload Differences of Detail Enrichment

The input parameter enrichmentDepth defaults to "standard". If you change this to "full", the Actor transitions from a lightweight guest job page scraper to an aggressive multi-request scraper. According to the schema documentation, the "full" enrichment depth is approximately five to seven times slower per job.

Crucially, setting the enrichment depth to "full" completely drops the applyUrl field from the output. In standard mode, the Actor grabs the static redirection link from the Apply button. In full mode, LinkedIn renders this button via client-side JavaScript, which does not expose a static link.

Furthermore, company page enrichment triggers an extra web request for every unique company identifier found in your job list. If your search keywords return fifty jobs from fifty different employers, the Actor must execute fifty additional concurrent requests to resolve company profiles. If those fifty jobs are all at one enterprise company, it only executes one. This makes run times highly variable and dependent on search results rather than input parameters.

What is the structured schema difference between Standard and Full enrichment?

Standard enrichment includes candidate application pathways like the outbound apply URL and the resolved ATS platform name while omitting internal taxonomy IDs. Full enrichment drops the static application link but appends similar job arrays, related searches, and internal numerical category labels. Choosing full enrichment when your pipeline requires apply links will break your workflow.

Understanding this division is essential because selecting "full" enrichment when your pipeline requires candidate application tracking routing will break your automation. If your downstream database maps directly to custom outreach tools, you must stick with the "standard" configuration depth.

Here is an example JSON input configuration demonstrating how to request standard enrichment to preserve the application tracking system data:

{
    "keywords": "data engineer",
    "location": "United States",
    "maxItems": 50,
    "scrapeJobDetails": true,
    "enrichmentDepth": "standard",
    "proxyConfiguration": {
        "useApifyProxy": true
    }
}
Enter fullscreen mode Exit fullscreen mode

The matching output object payload structure varies significantly. Below is the exact JSON shape to expect when parsing standard outputs:

{
    "id": "3876543210",
    "title": "Senior Data Engineer",
    "companyName": "Tech Corp",
    "companyUrl": "https://www.linkedin.com/company/techcorp",
    "location": "Austin, TX",
    "applyUrl": "https://careers.techcorp.com/jobs/101",
    "atsPlatform": "Greenhouse",
    "salaryRange": {
        "min": 120000,
        "max": 160000,
        "currency": "USD",
        "interval": "yearly"
    },
    "scrapedAt": "2026-10-02T14:32:01.000Z"
}
Enter fullscreen mode Exit fullscreen mode

Sentinel Values and Omitted Keys in Job Output Shapes

A common mistake when writing parsers for scraped data is expecting a static JSON output model. The linkedin-jobs-scraper explicitly states that empty or unavailable fields are omitted entirely from the output dataset rather than returned as null values.

If a job listing does not state a salary, the salary and salaryRange keys are completely missing from the dictionary. If a company page fails to load or lacks a description, fields like companyWebsite, companySpecialties, and companyDescription will not exist in the output JSON.

Your parsing code must utilize defensive lookup strategies. If you write item["salaryRange"]["min"] without checking if the parent key exists, your pipeline will throw a KeyError and crash mid-run.

How do you safely parse highly volatile scraped outputs without throwing runtime exceptions?

To parse volatile outputs without exceptions, you must use safe dictionary lookups that provide default fallback values when expected keys are absent. In Python, utilizing the .get() method is the standard way to retrieve optional fields or nested dictionaries. This prevents unhandled lookup errors from crashing your processing loop when encountering incomplete postings.

Using safe fallback defaults prevents runtime errors from crashing your indexing scripts mid-batch. It allows you to build a structured, unified database from highly non-uniform raw data.

Here is a robust Python parser designed to ingest these irregular schema records without failing:

def normalize_job_record(raw_item: dict) -> dict:
    salary_data = raw_item.get("salaryRange") or {}

    return {
        "job_id": raw_item.get("id", "UNKNOWN"),
        "title": raw_item.get("title", "Untitled Position"),
        "company": raw_item.get("companyName", "Unknown Employer"),
        "company_url": raw_item.get("companyWebsite", ""),
        "apply_link": raw_item.get("applyUrl", ""),
        "ats": raw_item.get("atsPlatform", "Direct"),
        "min_salary": salary_data.get("min", None),
        "max_salary": salary_data.get("max", None),
        "currency": salary_data.get("currency", "USD"),
        "has_visa_support": raw_item.get("visaSponsorshipMentioned", False),
        "collected_at": raw_item.get("scrapedAt", "")
    }
Enter fullscreen mode Exit fullscreen mode

Handling Non-English Locations and LinkedIn Translation Quirks

The location and locations input fields carry a strict design constraint: they must always be written in English. LinkedIn's cookieless guest search engine does not resolve non-English geographical place names.

If you input "Deutschland" instead of "Germany", or "Londres" instead of "London", LinkedIn's public endpoint will either fail to resolve the location entirely, default back to a global search, or silently return empty results.

Because the Actor does not validate the spelling or language of your inputs before sending them to LinkedIn, this failure presents as a successful run that mysteriously returns zero results.

When does the 300-second synchronous cap kill your data pipeline?

The Apify platform imposes a hard timeout limit of 300 seconds on all synchronous API calls, resulting in an HTTP 408 gateway timeout. If you trigger an Actor run synchronously and the crawl execution takes longer than 5 minutes, the connection closes abruptly. For complex scraper runs, you must execute the Actor asynchronously and poll the status.

For scrapers, reaching this limit is incredibly easy. Standard searches with multiple locations or deep company page enrichment regularly scale past 5 minutes.

To prevent broken pipelines, you must trigger your jobs asynchronously. Instead of waiting for the connection to close, initiate the run, obtain the run ID, and then poll the execution status or deploy webhooks to receive the final dataset payload when ready.

# Initiate an asynchronous run on the Apify platform
curl -X POST "https://api.apify.com/v2/acts/crawlerbros~linkedin-jobs-scraper/runs?token=$APIFY_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{
       "keywords": "security engineer",
       "location": "United States",
       "maxItems": 200
     }'
Enter fullscreen mode Exit fullscreen mode

This request returns a JSON response containing a run ID immediately. You can track this ID to monitor its status, avoiding HTTP 408 errors entirely.

{
  "data": {
    "id": "abc123xyz789run",
    "actId": "crawlerbros~linkedin-jobs-scraper",
    "status": "RUNNING",
    "startedAt": "2026-10-02T14:30:00.000Z",
    "defaultDatasetId": "dataset-abc123xyz"
  }
}
Enter fullscreen mode Exit fullscreen mode

The Single Consumer Request Queue Bottleneck

The Apify platform utilizes a system called the Request Queue to coordinate which URLs need to be scraped during a run. A fundamental architectural rule of this storage system is that a request queue can only be actively processed by one Actor run or task run at a time.

You cannot run multiple concurrent instances of the scraper pointing to the same shared request queue to try and speed up your scraping operations. Attempting to share a queue across concurrent runs will lock the resource and result in serialization errors.

If you need to scale up your scraping throughput, you must configure completely independent tasks. Each task must initialize its own isolated queue. To deduplicate results across these distinct tasks, you can utilize the excludeJobIds input array, passing the collected IDs of previous runs to the input of subsequent executions.

Real-World Cost Analysis of the Pay-Per-Event Billing Model

This Actor is billed under the PAY_PER_EVENT pricing model, which charges you based on specific actions emitted during the run. On top of those event prices, the Apify platform usage for the run is billed separately at the reader's Apify plan rates. You must consider both components, as the event prices do not constitute the whole cost of a run, and you should never total a run's cost from event prices alone. No subscription-plan rate applies to these event prices.

The specific event charges for this Actor are:

  • "Actor Start" (apify-actor-start): $0.05 per GB of memory allocated to the run. This is a flat per-event price charged once per run at initialization.
  • "result" (apify-default-dataset-item): $0.002 per event. A result event represents a single job listing item pushed to the default dataset.

The pricing of the "result" event scales based on your Apify user account's discount-tier. These tiers are explicitly defined with the following per-event rates:

  • FREE: $0.002
  • BRONZE: $0.00167
  • SILVER: $0.00133
  • GOLD: $0.001
  • PLATINUM: $0.001
  • DIAMOND: $0.001

The event part of the cost scales directly with the number of events a run emits (such as the number of result records pushed to the dataset); the platform-usage part scales with the resources the run consumes. If you configure your search filters too broadly, the Actor may fetch and discard hundreds of irrelevant jobs to find the matching subset. While you are charged the event fee only for successful "result" events written to the final dataset, those filtering operations still consume separate platform usage during their processing window.

What are the limitations and failure modes of this scraper?

This scraper operates on cookieless guest pages, which restricts it from viewing private or network-only job listings. Additionally, because LinkedIn does not expose stable metadata fields on public interfaces, filters for work type, experience level, and industry must be applied post-fetch, which can drastically reduce yield. Lastly, standard datacenter proxy sessions persist for 26 hours, whereas residential sessions rotate or expire after approximately 30 minutes, risking timeouts on long runs.

Below are the key limitations and technical caveats you must design your pipeline to handle:

  • No Private or Network Listings: Because this scraper does not use a login, it cannot access job listings restricted to logged-in users, private groups, or personal networks. It only sees what is visible to the public.
  • Proxy Lifespan Limitations: Apify datacenter proxies persist for 26 hours, while residential sessions last around 30 minutes. If your run has a large maxItems value or targets multiple locations, forcing the run past 30 minutes, residential connections will drop, leading to sudden request timeouts.
  • Work Type Inaccuracy: LinkedIn no longer exposes a reliable on-site, remote, or hybrid signal on its public guest pages. As a result, the workType input filter is a best-effort field and rarely narrows down search results accurately.
  • Transient Captchas: Public search endpoints are subject to aggressive IP rate limits. If the Actor cannot escalate requests through Apify's Unblocker proxy successfully, requests will be blocked and return zero results.
  • Unstable Experience and Job Type Filters: Because LinkedIn does not filter experience level or job type server-side for cookieless searches, the Actor must post-filter. This means it may fetch hundreds of items to return only a handful of matches, driving up platform usage without generating results.

Defensive Pipeline Design and Schema Auditing

To deploy this scraper reliably in production, your orchestrator must handle missing schema attributes, enforce API call fallbacks, and manage proxy lifespans. Apify datacenter proxies persist for 26 hours, whereas residential proxy sessions generally die after 30 minutes. If your scraping run is configured with massive item limits that push execution times past 30 minutes, residential proxy connections will drop, leading to sudden request timeouts.

Checked against the Actor's input schema and Apify docs on 2026-10-02.

The following script integrates all of these concepts into a production-grade execution loop. It starts the Actor asynchronously on the Apify platform, waits for completion without risking HTTP 408 timeouts, and processes the output dataset using safe dictionary access:

import os
import time
from apify_client import ApifyClient

def run_safe_linkedin_pipeline():
    token = os.environ.get("APIFY_TOKEN")
    if not token:
        raise ValueError("Missing APIFY_TOKEN environment variable")

    client = ApifyClient(token)

    run_input = {
        "keywords": "reliability engineer",
        "location": "United States",
        "maxItems": 150,
        "scrapeJobDetails": True,
        "enrichmentDepth": "standard",
        "proxyConfiguration": {
            "useApifyProxy": true
        }
    }

    # Start asynchronously to bypass the 300s sync limit
    run = client.actor("crawlerbros/linkedin-jobs-scraper").call(
        run_input=run_input,
        wait_secs=0
    )

    run_id = run["id"]
    print(f"Started run {run_id}. Polling status...")

    while True:
        status_details = client.run(run_id).get()
        status = status_details.get("status")
        print(f"Current status: {status}")

        if status in ["SUCCEEDED", "FAILED", "ABORTED", "TIMED-OUT"]:
            break
        time.sleep(15)

    if status != "SUCCEEDED":
        raise RuntimeError(f"Actor run failed with status: {status}")

    dataset_id = status_details["defaultDatasetId"]
    items = client.dataset(dataset_id).list_items().items

    processed_jobs = []
    for item in items:
        # Avoid KeyErrors with dict.get() for optional keys
        salary_info = item.get("salaryRange") or {}

        processed_jobs.append({
            "job_id": item.get("id"),
            "position": item.get("title"),
            "employer": item.get("companyName"),
            "apply_url": item.get("applyUrl"),
            "salary_min": salary_info.get("min"),
            "salary_max": salary_info.get("max"),
            "visa_sponsorship": item.get("visaSponsorshipMentioned"),
            "scraped_time": item.get("scrapedAt")
        })

    print(f"Successfully processed {len(processed_jobs)} records.")
    return processed_jobs

if __name__ == "__main__":
    run_safe_linkedin_pipeline()
Enter fullscreen mode Exit fullscreen mode

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

Top comments (0)