<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Devil Scrapes</title>
    <description>The latest articles on DEV Community by Devil Scrapes (@devil_scrapes).</description>
    <link>https://dev.to/devil_scrapes</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3960872%2Ffa930ad0-5ebc-4ca7-b894-6bc6fb3e2b40.png</url>
      <title>DEV Community: Devil Scrapes</title>
      <link>https://dev.to/devil_scrapes</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devil_scrapes"/>
    <language>en</language>
    <item>
      <title>Bayt Jobs Scraper: the proxy country pin that was the block</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Sat, 26 Sep 2026 10:06:59 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/bayt-jobs-scraper-the-proxy-country-pin-that-was-the-block-4281</link>
      <guid>https://dev.to/devil_scrapes/bayt-jobs-scraper-the-proxy-country-pin-that-was-the-block-4281</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;We pinned our Bayt.com scraper's proxy to the country it scrapes, and that pin was the block. On Apify residential exits in the UAE, cloud QA landed &lt;strong&gt;0 rows in 15 attempts&lt;/strong&gt;. We changed one field to a UK residential exit and the same build returned &lt;strong&gt;20 rows&lt;/strong&gt;, then a full 60-row run with no blocked pages. The data stayed genuinely Emirati because Bayt decides the country from the URL path (&lt;code&gt;/en/uae/jobs/...&lt;/code&gt;), not from where your IP sits. The &lt;a href="https://apify.com/DevilScrapes/bayt-jobs-scraper" rel="noopener noreferrer"&gt;Bayt Jobs Scraper&lt;/a&gt; now defaults to &lt;code&gt;RESIDENTIAL&lt;/code&gt; / &lt;code&gt;GB&lt;/code&gt;, and you pick the market with the &lt;code&gt;country&lt;/code&gt; input.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shouldn't a proxy match the country you're scraping?
&lt;/h2&gt;

&lt;p&gt;Usually, yes. A geo-random exit can quietly give you plausible but wrong data: prices in another currency, or another region's catalogue. Pinning the exit is normally the safe default, and it's our fleet rule for exactly that reason. The rule assumes the site uses your IP to pick the country. Bayt doesn't. It scopes every search by a path segment, so &lt;code&gt;/en/uae/jobs/accountant-jobs/&lt;/code&gt; returns UAE jobs whatever exit the request comes from. That leaves the pin with a cost and no benefit. The UAE residential pool was the one Bayt's Cloudflare layer challenged hardest, so pinning to it made us look &lt;em&gt;more&lt;/em&gt; suspicious.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you tell "the pin is the block" apart from "the site blocks us"?
&lt;/h2&gt;

&lt;p&gt;Change one variable. Same build, same input, same residential tier, and only the exit country different: AE went 0 for 15, GB cleared on the first attempt and kept clearing over three spaced runs. Swapping tiers or TLS profiles at the same time would have hidden which change did the work. Before you shelve a target as unreachable, try a neutral exit.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the site picks its country from the URL, the exit country is just one more fingerprint, and the local one can be the most-watched.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per job: &lt;code&gt;job_id&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;company&lt;/code&gt;, &lt;code&gt;company_url&lt;/code&gt;, &lt;code&gt;location&lt;/code&gt;, &lt;code&gt;career_level&lt;/code&gt;, &lt;code&gt;salary&lt;/code&gt; (only on postings that publish it), a &lt;code&gt;remote&lt;/code&gt; flag, &lt;code&gt;summary&lt;/code&gt;, &lt;code&gt;posted_at_raw&lt;/code&gt; (e.g. "3 days ago"), the absolute &lt;code&gt;detail_url&lt;/code&gt;, and the &lt;code&gt;keyword&lt;/code&gt; that found it. Pass several &lt;code&gt;keywords&lt;/code&gt; in one run, and set &lt;code&gt;country&lt;/code&gt; to &lt;code&gt;uae&lt;/code&gt;, &lt;code&gt;saudi-arabia&lt;/code&gt;, &lt;code&gt;egypt&lt;/code&gt; or any other Bayt market path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;List-page fields only, with no follow-through to detail pages. The posted date is Bayt's own relative text, kept verbatim rather than guessed into a timestamp. Salary appears on a minority of postings, and we return null rather than invent it. A search with no matches is a successful run.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need a Bayt account?&lt;/strong&gt;&lt;br&gt;
No. It reads Bayt's public search result pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why a UK proxy for UAE jobs?&lt;/strong&gt;&lt;br&gt;
Because Bayt picks the market from the URL, and UK residential exits cleared where UAE ones were challenged. The rows are still the UAE (or whichever &lt;code&gt;country&lt;/code&gt; you set).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What about blocks and captchas?&lt;/strong&gt;&lt;br&gt;
That's our job. We rotate browser fingerprints, retry with backoff on rate limits and challenges, and rotate residential sessions on every block.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.002 per job row: &lt;strong&gt;$2.20 per 1,000 results&lt;/strong&gt;. A zero-match search succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/bayt-jobs-scraper" rel="noopener noreferrer"&gt;Bayt Jobs Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We handle the fingerprints, the proxies and the retries, and we test the proxy default we ship against the target before a customer ever runs it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>apify</category>
      <category>proxy</category>
    </item>
    <item>
      <title>InfoJobs Spain Scraper: the proxy default that 407'd every run</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:11:38 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/infojobs-spain-scraper-the-proxy-default-that-407d-every-run-4oao</link>
      <guid>https://dev.to/devil_scrapes/infojobs-spain-scraper-the-proxy-default-that-407d-every-run-4oao</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;We shipped a proxy default that read as reasonable and failed every single run. The config was &lt;code&gt;{"useApifyProxy": true, "apifyProxyCountry": "ES"}&lt;/code&gt; — a bare &lt;code&gt;useApifyProxy&lt;/code&gt; with no named group. That resolves to Apify's &lt;em&gt;datacenter&lt;/em&gt; tier, and our datacenter pack is a static US-only pool, so pinning &lt;code&gt;ES&lt;/code&gt; against it asked for an exit country the pack simply doesn't own. Cloud QA caught it before any customer did: &lt;code&gt;CONNECT 407&lt;/code&gt;, proxy auth rejected. The fix was naming the group — &lt;code&gt;apifyProxyGroups: ["RESIDENTIAL"]&lt;/code&gt; alongside the country pin — and the &lt;a href="https://apify.com/DevilScrapes/infojobs-spain-jobs-scraper" rel="noopener noreferrer"&gt;InfoJobs Spain Jobs Scraper&lt;/a&gt; now ships with both, not just one.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you pin a proxy country, do you also need to name the group?
&lt;/h2&gt;

&lt;p&gt;Yes, and skipping it isn't a degraded result — it's a guaranteed tunnel failure. A country pin tells Apify Proxy &lt;em&gt;which&lt;/em&gt; exit to hand you; the group tells it &lt;em&gt;which pool&lt;/em&gt; to hand it from. Name only the country and Apify defaults you into datacenter, because that's the cheaper tier absent other instructions. If your account's datacenter pack has no IPs in the country you pinned — ours is a static US-only 27-IP pack — the tunnel refuses to open at all. Every InfoJobs run against the old default would have failed before fetching a single page. This is a config-shape bug, not a scraping-difficulty bug, and it's exactly the class of thing a five-minute local test never surfaces because local runs don't route through Apify's proxy infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did the shipped defaults actually get exercised in the cloud before going live?
&lt;/h2&gt;

&lt;p&gt;No — and that was the second, separate finding on this Actor. The prefill that ships with every new Actor is what Apify's own automated QA runs, and what a customer sees pre-filled if they click Start without touching the form. Ours was &lt;code&gt;maxResults: 5, maxPages: 1&lt;/code&gt; — a single page. Pagination is the entire point of a job-board scraper, and it had never actually run in the cloud, by us or by Apify. We widened it to &lt;code&gt;maxResults: 40, maxPages: 2&lt;/code&gt; and re-ran: 40 rows landed, status message read "2/2 page(s) parsed". Two separate bugs, same root cause — a config nobody had actually pointed at the live target.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A prefill that returns one page is a prefill that never tests pagination — no matter how much the pagination code itself has been reviewed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per deduplicated job posting: &lt;code&gt;job_id&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;description&lt;/code&gt;, &lt;code&gt;city&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;contract_type&lt;/code&gt;, &lt;code&gt;salary_min&lt;/code&gt;/&lt;code&gt;salary_max&lt;/code&gt;/&lt;code&gt;salary_period&lt;/code&gt;/&lt;code&gt;salary_currency&lt;/code&gt; (null when unpublished), &lt;code&gt;workday&lt;/code&gt;, &lt;code&gt;teleworking&lt;/code&gt;, &lt;code&gt;published_at&lt;/code&gt; (ISO-8601), &lt;code&gt;company_name&lt;/code&gt;, &lt;code&gt;company_logo_url&lt;/code&gt;, &lt;code&gt;company_url&lt;/code&gt;, &lt;code&gt;states&lt;/code&gt;, and &lt;code&gt;is_executive&lt;/code&gt;. Province filters accept either a plain name ("Madrid") or InfoJobs' own numeric id ("33") — both resolve against the site's own province table, not a hand-typed lookup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;One &lt;code&gt;keyword&lt;/code&gt; plus an optional &lt;code&gt;province&lt;/code&gt; per run — no multi-keyword batching yet. &lt;code&gt;maxResults&lt;/code&gt; and &lt;code&gt;maxPages&lt;/code&gt; are independent caps; whichever fires first stops the run, and the status message says which. List-page fields only, no detail-page follow-through. A narrow search returning zero rows is a successful run, not a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why did the Actor 407 on every run before this fix?&lt;/strong&gt;&lt;br&gt;
Because the default proxy config named a country but not a group, which silently resolved to a US-only datacenter pack that has no Spanish exits. Naming &lt;code&gt;apifyProxyGroups: ["RESIDENTIAL"]&lt;/code&gt; fixed it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need an InfoJobs account or API key?&lt;/strong&gt;&lt;br&gt;
No — this reads InfoJobs' public search result pages, parsed from the page's own &lt;code&gt;window.__INITIAL_PROPS__&lt;/code&gt; JSON blob.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I filter by province?&lt;/strong&gt;&lt;br&gt;
Yes — pass a plain Spanish province name or its numeric InfoJobs id; the input param is &lt;code&gt;provinceIds&lt;/code&gt; under the hood, resolved by testing live requests rather than guessed from the UI.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.0025 per deduplicated job row — &lt;strong&gt;$2.70 per 1,000 results&lt;/strong&gt;. A zero-match search still succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/infojobs-spain-jobs-scraper" rel="noopener noreferrer"&gt;InfoJobs Spain Jobs Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints, retry with backoff, and route every request through Apify Proxy — and we test our own proxy configuration against the pack we actually own before we ship it as a default.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>apify</category>
      <category>proxy</category>
    </item>
    <item>
      <title>Fotocasa Scraper: the null field that hid every discounted listing</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:11:06 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/fotocasa-scraper-the-null-field-that-hid-every-discounted-listing-4610</link>
      <guid>https://dev.to/devil_scrapes/fotocasa-scraper-the-null-field-that-hid-every-discounted-listing-4610</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;We assumed &lt;code&gt;reducedPrice&lt;/code&gt; was a raw integer, because in the single listing we'd probed it, the field was &lt;code&gt;null&lt;/code&gt; — and &lt;code&gt;null&lt;/code&gt; has no type to infer from. A mandated live run against real listings surfaced the truth: when a Fotocasa listing actually carries a discount, &lt;code&gt;reducedPrice&lt;/code&gt; is a &lt;strong&gt;formatted string&lt;/strong&gt;, &lt;code&gt;"75.000 €"&lt;/code&gt;, the exact same shape as the regular &lt;code&gt;price&lt;/code&gt; field. Ship the integer assumption and the parser silently skips every discounted listing — no error, no warning, just a quietly incomplete dataset on precisely the rows a buyer cares about most. The &lt;a href="https://apify.com/DevilScrapes/fotocasa-spain-property-scraper" rel="noopener noreferrer"&gt;Fotocasa Spain Property Scraper&lt;/a&gt; parses it as a formatted string, same as price, before converting to &lt;code&gt;reduced_price_eur&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you infer a field's type from a &lt;code&gt;null&lt;/code&gt; sample?
&lt;/h2&gt;

&lt;p&gt;No — and that's the whole lesson here. A field that reads &lt;code&gt;null&lt;/code&gt; in your one probed sample has no observable shape; you cannot know from an absence whether the real value, when it exists, is an integer, a formatted string, or something else entirely. The only fix is running against enough live data that the field actually appears populated at least once. We did — and &lt;code&gt;reducedPrice&lt;/code&gt; turned out to share the exact same &lt;code&gt;"75.000 €"&lt;/code&gt;-style formatting as the primary &lt;code&gt;price&lt;/code&gt; field, not the raw integer a schema guess would produce. &lt;code&gt;reduced_price_eur&lt;/code&gt; in our output is legitimately &lt;code&gt;null&lt;/code&gt; on listings with no discount; it's only populated, as an integer in euros, on the listings that actually have one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does this scraper need a browser to run?
&lt;/h2&gt;

&lt;p&gt;No. Fotocasa.es server-renders each listing's full structured data into a &lt;code&gt;&amp;lt;script id="__initial_props__"&amp;gt;&lt;/code&gt; block on the search-results page itself — price, address, coordinates, features, photo URLs, and the description all arrive in that one JSON payload. No separate API calls, no client-side rendering to wait on. We read it with &lt;code&gt;curl-cffi&lt;/code&gt; under browser TLS impersonation, no headless browser required, which keeps the run cheap and the dataset arriving fast.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A field that's &lt;code&gt;null&lt;/code&gt; in your sample carries no type information at all — the only way to learn its real shape is a live run against data where it's actually populated.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One validated row per listing: price in euros (&lt;code&gt;price_eur&lt;/code&gt;, plus the human-formatted &lt;code&gt;price_display&lt;/code&gt;) and &lt;code&gt;reduced_price_eur&lt;/code&gt; when discounted, &lt;code&gt;building_type&lt;/code&gt;/&lt;code&gt;building_subtype&lt;/code&gt;, &lt;code&gt;rooms&lt;/code&gt;, &lt;code&gt;bathrooms&lt;/code&gt;, &lt;code&gt;surface_m2&lt;/code&gt;, full address (&lt;code&gt;address_province&lt;/code&gt;, &lt;code&gt;address_city&lt;/code&gt;, &lt;code&gt;address_district&lt;/code&gt;, &lt;code&gt;address_neighborhood&lt;/code&gt;, &lt;code&gt;address_zip_code&lt;/code&gt;), &lt;code&gt;latitude&lt;/code&gt;/&lt;code&gt;longitude&lt;/code&gt;, &lt;code&gt;dynamic_features&lt;/code&gt;, up to ~18 &lt;code&gt;image_urls&lt;/code&gt;, the full Spanish-language &lt;code&gt;description&lt;/code&gt;, and &lt;code&gt;agency_alias&lt;/code&gt;/&lt;code&gt;agency_type&lt;/code&gt;/&lt;code&gt;phone&lt;/code&gt; when the seller publishes them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;Search-results-page fields only in v1 — no per-listing detail-page fetch, since the results page already embeds everything above. Rent (&lt;code&gt;alquiler&lt;/code&gt;) URLs are supported best-effort; sale (&lt;code&gt;comprar&lt;/code&gt;) is the wire-confirmed path. Price filters apply after fetching, so a narrow band still costs the request for out-of-range listings on that page. &lt;code&gt;propertyTypeId&lt;/code&gt; passes through as Fotocasa's raw internal ID — there's no confirmed ID-to-label table yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why did the original spec get &lt;code&gt;reducedPrice&lt;/code&gt;'s type wrong?&lt;/strong&gt;&lt;br&gt;
The only sample probed at spec time had it as &lt;code&gt;null&lt;/code&gt;, and &lt;code&gt;null&lt;/code&gt; doesn't reveal a type. A live run against listings that actually carry a discount showed it's a formatted string like price, not a raw integer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need a Fotocasa account or API key?&lt;/strong&gt;&lt;br&gt;
No — Fotocasa publishes no public API for this data; this Actor reads the same public search-results pages a browser loads and parses the JSON already embedded in them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is the proxy pinned to residential Spain (&lt;code&gt;ES&lt;/code&gt;) by default?&lt;/strong&gt;&lt;br&gt;
Because a Spain-only portal can render the wrong currency, or block outright, from a foreign exit — we force the ES residential exit regardless of other settings.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.004 per validated listing — &lt;strong&gt;$4.20 per 1,000 results&lt;/strong&gt;. No data, no charge beyond the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/fotocasa-spain-property-scraper" rel="noopener noreferrer"&gt;Fotocasa Spain Property Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints, rotate residential proxies, retry with backoff, and ship clean typed rows — and we don't guess a field's shape from a sample where it's empty.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>apify</category>
      <category>realestate</category>
    </item>
    <item>
      <title>XING's job data ships as a JS object, not JSON. Here's how we parse it without breaking on the first undefined.</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:56:33 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/xings-job-data-ships-as-a-js-object-not-json-heres-how-we-parse-it-without-breaking-on-the-12n8</link>
      <guid>https://dev.to/devil_scrapes/xings-job-data-ships-as-a-js-object-not-json-heres-how-we-parse-it-without-breaking-on-the-12n8</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;XING ships a complete Apollo GraphQL cache inline in every search-result page — &lt;code&gt;&amp;lt;script id="runtime-config"&amp;gt;window.crate={...}&lt;/code&gt; — so the job data never needs DOM scraping. The catch: that blob is a JavaScript object literal, not JSON, so a naive &lt;code&gt;json.loads&lt;/code&gt; throws on it. We normalize it first, then resolve &lt;code&gt;ROOT_QUERY&lt;/code&gt;'s &lt;code&gt;jobSearchByQuery&lt;/code&gt; collection into typed &lt;code&gt;VisibleJob&lt;/code&gt;/&lt;code&gt;Company&lt;/code&gt; nodes. Verified cloud run: 50 deduplicated rows across two keywords ("developer" 40, "marketing" 10), Berlin, pagination confirmed across 2 pages per keyword. The &lt;a href="https://apify.com/DevilScrapes/xing-jobs-scraper" rel="noopener noreferrer"&gt;XING Jobs Scraper&lt;/a&gt; is built on that resolved graph, not a screen-scrape.&lt;/p&gt;

&lt;h2&gt;
  
  
  Isn't &lt;code&gt;window.crate={...}&lt;/code&gt; just JSON you can &lt;code&gt;json.loads&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;No, and treating it like JSON is the fastest way to ship a scraper that dies on page two. &lt;code&gt;window.crate&lt;/code&gt; is a live JavaScript object literal assigned in a &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tag — it's what the browser's own runtime evaluates, not a JSON serialization of it. It contains bare &lt;code&gt;undefined&lt;/code&gt; tokens in place of missing fields, which are valid JavaScript but not valid JSON syntax; &lt;code&gt;json.loads&lt;/code&gt; throws on the first one it hits. We normalize the blob — swapping the JS-only tokens for their JSON equivalents — before we ever try to parse it, which is the difference between "this works on the one listing I tested" and "this works on the thousandth listing that happens to have a hole in it."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where does the company name actually come from?
&lt;/h2&gt;

&lt;p&gt;Not where you'd expect. The Apollo cache is normalized: every job entity holds a reference like &lt;code&gt;Company:&amp;lt;id&amp;gt;&lt;/code&gt; instead of the company's data inline, and resolving that reference on its own gets you a logo URL and nothing else — no name. The name lives on a &lt;em&gt;different&lt;/em&gt; node entirely, &lt;code&gt;companyInfo.companyNameOverride&lt;/code&gt;, which we join back onto the resolved job. Skip that join and you ship a scraper that returns a logo and a blank company column — technically a "success," practically useless to a recruiter filtering by employer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is &lt;code&gt;salaryMedian&lt;/code&gt; sparser than &lt;code&gt;salaryMin&lt;/code&gt;/&lt;code&gt;salaryMax&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;Because XING serves two structurally different salary shapes under one field name, and only one of them carries a median. Its own &lt;code&gt;SalaryEstimate&lt;/code&gt; type includes a computed median; an employer-stated &lt;code&gt;SalaryRange&lt;/code&gt; gives you a min and max only, no median, because the employer never published one. Counted live on a single result page: 44 &lt;code&gt;SalaryEstimate&lt;/code&gt; nodes against 28 &lt;code&gt;SalaryRange&lt;/code&gt; nodes. So on a page where roughly two-thirds of salaried postings carry a median, expect that ratio to hold in your dataset too — the sparsity isn't a scraping gap, it's what XING itself published.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The blob at &lt;code&gt;window.crate&lt;/code&gt; looks like JSON and isn't — bare &lt;code&gt;undefined&lt;/code&gt; tokens break &lt;code&gt;json.loads&lt;/code&gt; before you ever get to the graph you actually want.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One deduplicated row per job across one or more keywords: &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;slug&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;companyName&lt;/code&gt;, &lt;code&gt;companyLogoUrl&lt;/code&gt;, &lt;code&gt;locationCity&lt;/code&gt;, &lt;code&gt;locations&lt;/code&gt;, &lt;code&gt;employmentType&lt;/code&gt;, &lt;code&gt;salaryMin&lt;/code&gt;/&lt;code&gt;salaryMax&lt;/code&gt;, &lt;code&gt;salaryMedian&lt;/code&gt; (when the posting is a &lt;code&gt;SalaryEstimate&lt;/code&gt;), &lt;code&gt;salaryCurrency&lt;/code&gt;, &lt;code&gt;keyResponsibilities&lt;/code&gt;, &lt;code&gt;refreshedAt&lt;/code&gt;, &lt;code&gt;activeUntil&lt;/code&gt;, &lt;code&gt;paid&lt;/code&gt;, &lt;code&gt;topJob&lt;/code&gt;, plus &lt;code&gt;sourceKeyword&lt;/code&gt;/&lt;code&gt;sourceLocation&lt;/code&gt; tagging which search found each row. Run-wide dedup means the same job id never appears twice, even across pages or across keywords batched into one run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;List-page fields only — full job-description text isn't fetched; follow the returned &lt;code&gt;url&lt;/code&gt; for a detail-page job if you need it. XING ranks by relevance, not strict keyword match, so a niche keyword can surface adjacent roles instead of an empty page. A narrow search legitimately returning zero rows is a successful run, not a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need a XING account or API key?&lt;/strong&gt;&lt;br&gt;
No — this reads XING's public search result pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I search multiple keywords in one run?&lt;/strong&gt;&lt;br&gt;
Yes — pass an array to &lt;code&gt;keywords&lt;/code&gt;; every row is tagged via &lt;code&gt;sourceKeyword&lt;/code&gt;, and dedup applies across the whole run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did I get fewer rows than &lt;code&gt;maxResults&lt;/code&gt;?&lt;/strong&gt;&lt;br&gt;
Either the search genuinely has fewer matches, or &lt;code&gt;maxPagesPerKeyword&lt;/code&gt; stopped the run first — the run's status message says which.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.0025 per deduplicated job row — &lt;strong&gt;$2.70 per 1,000 results&lt;/strong&gt;. A zero-match search still succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/xing-jobs-scraper" rel="noopener noreferrer"&gt;XING Jobs Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate Chrome/Firefox TLS fingerprints, retry with backoff, and route every request through residential proxy pinned to Germany — and we normalize XING's inline cache into a real object graph instead of guessing at a screen-scrape.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>graphql</category>
      <category>apify</category>
    </item>
    <item>
      <title>PropertyFinder.ae's category parameter wasn't what the spec assumed. We tested it before shipping.</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:56:03 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/propertyfinderaes-category-parameter-wasnt-what-the-spec-assumed-we-tested-it-before-shipping-523j</link>
      <guid>https://dev.to/devil_scrapes/propertyfinderaes-category-parameter-wasnt-what-the-spec-assumed-we-tested-it-before-shipping-523j</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;We tested PropertyFinder.ae's &lt;code&gt;c=&lt;/code&gt; search parameter before trusting it, and it wasn't what the spec assumed. The written spec read &lt;code&gt;c=&lt;/code&gt; as a plain residential-vs-commercial switch. The wire disagreed: &lt;code&gt;c=1&lt;/code&gt; is residential-sale, &lt;code&gt;c=2&lt;/code&gt; is residential-rent, &lt;code&gt;c=3&lt;/code&gt; is commercial-sale, &lt;code&gt;c=4&lt;/code&gt; is commercial-rent — one parameter carries both category &lt;em&gt;and&lt;/em&gt; transaction type. Shipped on the original reading, the &lt;a href="https://apify.com/DevilScrapes/propertyfinder-uae-property-scraper" rel="noopener noreferrer"&gt;PropertyFinder UAE Property Scraper&lt;/a&gt; would have quietly handed a customer asking for for-sale listings a page of for-rent ones. We caught it before it shipped, not after a refund request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does PropertyFinder.ae's &lt;code&gt;c=&lt;/code&gt; parameter just mean residential vs commercial?
&lt;/h2&gt;

&lt;p&gt;No, and assuming so is the mistake worth flagging for anyone building against this target directly. &lt;code&gt;c=&lt;/code&gt; is a combined category-and-transaction axis, not two independent filters bolted together — &lt;code&gt;category&lt;/code&gt; and &lt;code&gt;purpose&lt;/code&gt; in this Actor's input map onto that single parameter internally so you never have to know the encoding. The listings themselves live in the page's embedded &lt;code&gt;__NEXT_DATA__&lt;/code&gt; JSON blob — roughly 584,000 characters of it — at &lt;code&gt;props.pageProps.searchResult.listings&lt;/code&gt;, which is a far more reliable source than parsing rendered DOM, but the query encoding sitting in front of that JSON still has to be verified against real requests, not inferred from the parameter's name.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when you ask for a bedroom filter that doesn't exist on the wire?
&lt;/h2&gt;

&lt;p&gt;There is no server-side bedroom parameter on this target — we checked, and it isn't there. That left two options: invent a query param the server silently ignores, which looks like it worked and does nothing, or filter client-side on the parsed rows and say so plainly in the docs. We chose the second. An unverified filter is worse than a missing one, because a filter the server ignores returns a plausible, confidently wrong result set, and nothing in the output tells you it was never applied. We'd rather ship an honest client-side filter than a fake server-side one. We also confirmed the failure mode on the other end: an absurd price range returns a genuine zero-match &lt;code&gt;SUCCEEDED&lt;/code&gt; run with a status message — not a crash dressed up as a bug.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An unverified filter is worse than a missing one — a filter the server ignores returns a plausible, confidently wrong result set, and nothing in the output tells the buyer it was never applied.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per listing: &lt;code&gt;listing_id&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;price_amount&lt;/code&gt; + &lt;code&gt;price_currency&lt;/code&gt;, &lt;code&gt;property_type&lt;/code&gt;, &lt;code&gt;bedrooms&lt;/code&gt;, &lt;code&gt;bathrooms&lt;/code&gt;, &lt;code&gt;size_value&lt;/code&gt; + &lt;code&gt;size_unit&lt;/code&gt;, &lt;code&gt;location&lt;/code&gt; and &lt;code&gt;location_hierarchy&lt;/code&gt; (city → community → sub-community), &lt;code&gt;agent_name&lt;/code&gt;, &lt;code&gt;agency_name&lt;/code&gt;, &lt;code&gt;listing_url&lt;/code&gt;, &lt;code&gt;image_urls&lt;/code&gt;, &lt;code&gt;posted_at&lt;/code&gt; / &lt;code&gt;updated_at&lt;/code&gt;, and &lt;code&gt;latitude&lt;/code&gt;/&lt;code&gt;longitude&lt;/code&gt; where PropertyFinder's own JSON carries them. A real sampled row: a 4-bedroom villa in Meadows 7, Meadows, Dubai, AED 23,500,000, &lt;code&gt;location_hierarchy&lt;/code&gt; &lt;code&gt;["Dubai", "Meadows", "Meadows 7"]&lt;/code&gt;, with agent name and coordinates attached. Proxy is &lt;code&gt;RESIDENTIAL&lt;/code&gt;, pinned to &lt;code&gt;AE&lt;/code&gt; — forced in code, not trusted from whatever the run's prefill happens to say, because a geo-random exit on a country-scoped portal is a known way to get wrong-currency or wrong-inventory data back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;v1 is search-results only — full description text, amenities lists, and floor plans live on the detail page and aren't fetched. &lt;code&gt;location_hierarchy&lt;/code&gt;, &lt;code&gt;latitude&lt;/code&gt;/&lt;code&gt;longitude&lt;/code&gt;, and &lt;code&gt;agency_name&lt;/code&gt; are null when PropertyFinder's own JSON doesn't carry them for a given listing — we don't backfill guesses. &lt;code&gt;bedrooms&lt;/code&gt; is applied client-side, as covered above, not as a confirmed server-side filter. Non-UAE PropertyFinder country portals are out of scope for this Actor.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does &lt;code&gt;category&lt;/code&gt;/&lt;code&gt;purpose&lt;/code&gt; map onto one parameter or two on the real site?&lt;/strong&gt;&lt;br&gt;
One — &lt;code&gt;c=&lt;/code&gt; is a combined category-and-transaction axis (&lt;code&gt;1&lt;/code&gt;–&lt;code&gt;4&lt;/code&gt;), not two independent switches. This Actor encodes it correctly so you just pick &lt;code&gt;category&lt;/code&gt; and &lt;code&gt;purpose&lt;/code&gt; and never touch the raw parameter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the bedroom filter applied by PropertyFinder's server or by the Actor?&lt;/strong&gt;&lt;br&gt;
By the Actor, client-side, on the parsed rows — there's no confirmed server-side bedroom parameter on this target, and we'd rather filter honestly after the fact than send a query param the server silently ignores.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a zero-result search count as a failed run?&lt;/strong&gt;&lt;br&gt;
No — an absurd price range verified as a genuine &lt;code&gt;SUCCEEDED&lt;/code&gt; run with zero rows and a status message, not a crash.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.008 per listing row — &lt;strong&gt;$8.20 per 1,000 results&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/propertyfinder-uae-property-scraper" rel="noopener noreferrer"&gt;PropertyFinder UAE Property Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt; 😈. We rotate Chrome/Firefox TLS fingerprints, retry with backoff, pin the proxy exit country where the target needs it, and verify every filter against the real wire before we ship it as working.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>apify</category>
      <category>realestate</category>
    </item>
    <item>
      <title>A green QA run hid an untested pagination path. Here's the check that caught it before we shipped.</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:55:32 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/a-green-qa-run-hid-an-untested-pagination-path-heres-the-check-that-caught-it-before-we-shipped-3m7e</link>
      <guid>https://dev.to/devil_scrapes/a-green-qa-run-hid-an-untested-pagination-path-heres-the-check-that-caught-it-before-we-shipped-3m7e</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;Our own QA harness silently clamped &lt;code&gt;maxResults&lt;/code&gt; from the prefill's 48 down to 3. The run passed, returned rows, and looked completely green — but it had only ever fetched page one, so pagination shipped untested behind a passing check. We caught it because the harness prints a DEPTH UNPROVEN warning when a run's actual depth falls short of its own prefill, and we fired a deliberate deep run at the prefill's own value: 48 results, charged &lt;code&gt;result: 48&lt;/code&gt; times, pagination proven across pages before this Actor shipped. The &lt;a href="https://apify.com/DevilScrapes/realtor-ca-listings-scraper" rel="noopener noreferrer"&gt;Realtor.ca Listings Scraper&lt;/a&gt; sells depth — a curated city or bounding box of normalized Canadian listings — so a capped smoke test couldn't be allowed to stand in for proof of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can a "successful" QA run still hide an untested feature?
&lt;/h2&gt;

&lt;p&gt;Yes, and that's the uncomfortable part — success and coverage are two different claims, and a green checkmark only speaks to the first one. A run that exits 0, writes rows to the dataset, and matches its expected schema has proven the code executes without crashing. It has not proven the code executes the &lt;em&gt;path&lt;/em&gt; a customer will actually exercise. Here, the harness's own safety clamp — meant to keep smoke tests cheap — quietly capped &lt;code&gt;maxResults&lt;/code&gt; at 3 instead of running the prefill's real value of 48. Three results fit on page one of Realtor.ca's &lt;code&gt;api2.realtor.ca&lt;/code&gt; search response. The pagination logic that walks page two, page three, and beyond simply never ran, and nothing about the run's exit code said so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What made this catchable?
&lt;/h2&gt;

&lt;p&gt;A second signal that doesn't just ask "did it succeed" but "did it prove what it's supposed to prove." The harness compares the run's actual depth against the prefill's own stated depth and prints DEPTH UNPROVEN when they diverge — in this case, 3 delivered against 48 promised. That's the difference between a binary pass/fail and a coverage check: it forced the question "did this run test the feature we're charging for," not just "did it crash."&lt;/p&gt;

&lt;h2&gt;
  
  
  What did the deep run actually prove?
&lt;/h2&gt;

&lt;p&gt;We reran at the prefill's real value — 48 results, no clamp — and the run billed &lt;code&gt;result&lt;/code&gt; 48 times and returned distinct listings spanning multiple response pages, confirming the pagination path that page-one-only testing had never touched. A sampled row from that run: a Markham, Ontario condo listed at CAD 990,000, 5+1 bedrooms, with the agent's name and the brokerage's name and address attached — the full shape of what this Actor is built to return, not just what page one happens to contain.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A green run tells you the code ran, not that it ran the feature you are selling. If an Actor's product is depth, a capped smoke test cannot prove it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One normalized row per listing, deduplicated by &lt;code&gt;id&lt;/code&gt;: &lt;code&gt;mlsNumber&lt;/code&gt;, &lt;code&gt;price&lt;/code&gt;, &lt;code&gt;currency&lt;/code&gt;, &lt;code&gt;transactionType&lt;/code&gt;, &lt;code&gt;streetAddress&lt;/code&gt;, &lt;code&gt;latitude&lt;/code&gt;/&lt;code&gt;longitude&lt;/code&gt;, &lt;code&gt;propertyType&lt;/code&gt;, &lt;code&gt;bedrooms&lt;/code&gt;, &lt;code&gt;bathroomTotal&lt;/code&gt;, &lt;code&gt;sizeInterior&lt;/code&gt;, &lt;code&gt;amenities&lt;/code&gt;, &lt;code&gt;photoUrls&lt;/code&gt;, &lt;code&gt;agentName&lt;/code&gt;, &lt;code&gt;agentPhone&lt;/code&gt;, &lt;code&gt;brokerageName&lt;/code&gt;, &lt;code&gt;brokerageAddress&lt;/code&gt;, &lt;code&gt;listingDetailUrl&lt;/code&gt;, plus &lt;code&gt;searchCity&lt;/code&gt;, &lt;code&gt;page&lt;/code&gt;, and &lt;code&gt;scrapedAt&lt;/code&gt; tagging where and when each row was found.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;One search — one city or bounding box, one transaction type — per run, capped at the server's own 600-result ceiling. Residential listings only; commercial and land property groups are out of scope for v1. No agent email and no listing description text, because Realtor.ca's search response carries neither (we don't emit empty columns pretending otherwise). The curated city list covers 10 major Canadian cities; anywhere else needs an explicit bounding box.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need a Realtor.ca account or API key?&lt;/strong&gt;&lt;br&gt;
No — this reads Realtor.ca's own publicly listed search-result data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I search a custom area instead of a city?&lt;/strong&gt;&lt;br&gt;
Yes — pass an explicit bounding box (&lt;code&gt;latitudeMax&lt;/code&gt;/&lt;code&gt;latitudeMin&lt;/code&gt;/&lt;code&gt;longitudeMax&lt;/code&gt;/&lt;code&gt;longitudeMin&lt;/code&gt;) instead of a city name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you know pagination actually works, not just page one?&lt;/strong&gt;&lt;br&gt;
We measured it directly: a deep run at 48 results billed 48 result events and returned listings spanning multiple response pages, after our QA harness flagged a shallower run as DEPTH UNPROVEN.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.003 per deduplicated listing row — &lt;strong&gt;$3.20 per 1,000 results&lt;/strong&gt;. A zero-match search still succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/realtor-ca-listings-scraper" rel="noopener noreferrer"&gt;Realtor.ca Listings Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate residential proxies pinned to Canada, retry with backoff, and — before we call a run "proven" — we check that it actually exercised the depth we're charging for.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>qa</category>
      <category>apify</category>
    </item>
    <item>
      <title>A probe and a production scraper saw two different pages for the same missing Snapchat user</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:08:36 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/a-probe-and-a-production-scraper-saw-two-different-pages-for-the-same-missing-snapchat-user-53lk</link>
      <guid>https://dev.to/devil_scrapes/a-probe-and-a-production-scraper-saw-two-different-pages-for-the-same-missing-snapchat-user-53lk</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;A plain &lt;code&gt;curl&lt;/code&gt; at a nonexistent Snapchat username — no browser impersonation, no proxy — returns a bare &lt;code&gt;404&lt;/code&gt;. The same username, fetched through the path the &lt;a href="https://apify.com/DevilScrapes/snapchat-profile-scraper" rel="noopener noreferrer"&gt;Snapchat Profile Scraper&lt;/a&gt; actually ships (impersonated Chrome/Firefox TLS, US-proxied), returns &lt;code&gt;200&lt;/code&gt; with a doubly nested &lt;code&gt;__NEXT_DATA__&lt;/code&gt; payload whose real signal is buried at &lt;code&gt;pageProps.pageProps.pageMetadata.pageType == "NOT_FOUND"&lt;/code&gt;. Those are two different documents describing the same fact. Our classifier was originally written against the curl-shaped probe and called the production response MALFORMED instead of NOT_FOUND — a clean "this user doesn't exist" was misreported as a parse failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why would the same missing username produce two different response shapes?
&lt;/h2&gt;

&lt;p&gt;Because "the page your probe sees" and "the page your scraper sees" are only the same document if you fetched them the same way, and a bare curl and an impersonated, proxied browser session are not the same way. Snapchat serves a lighter, flatter error response to unauthenticated, unfingerprinted requests — a &lt;code&gt;404&lt;/code&gt; status with a small body — and serves the full Next.js application shell, &lt;code&gt;__NEXT_DATA__&lt;/code&gt; and all, to requests that look like a real browser hitting the page. Once you're getting the full shell, the "not found" state isn't a status code anymore, it's a field three levels deep inside a JSON payload that otherwise looks exactly like a live profile's payload structure. If your test fixtures were captured with a convenient &lt;code&gt;curl&lt;/code&gt; command during development and your classifier logic was written against &lt;em&gt;that&lt;/em&gt; shape, it will correctly handle a probe and incorrectly handle the thing you actually ship, because production traffic never takes the path you tested.&lt;/p&gt;

&lt;p&gt;The nesting itself is a small, specific trap: &lt;code&gt;pageProps&lt;/code&gt; appears twice, not once. The outer &lt;code&gt;pageProps&lt;/code&gt; is the Next.js page wrapper; the inner &lt;code&gt;pageProps&lt;/code&gt; is where Snapchat puts the actual page data, including &lt;code&gt;pageMetadata.pageType&lt;/code&gt;. Reading &lt;code&gt;data.pageProps.pageMetadata&lt;/code&gt; instead of &lt;code&gt;data.pageProps.pageProps.pageMetadata&lt;/code&gt; gets you &lt;code&gt;None&lt;/code&gt; on every single request, live profile or not — which looks exactly like a malformed-page symptom, not a missing-key symptom, and sends you debugging the wrong layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's the second finding, on top of the fixture mismatch?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;userProfile&lt;/code&gt; in that same payload is a discriminated union keyed on a &lt;code&gt;$case&lt;/code&gt; field, not a single fixed shape. The obvious-looking implementation assumes &lt;code&gt;publicProfileInfo&lt;/code&gt; is always present and reads it directly — which works for the common case and silently drops every profile whose &lt;code&gt;$case&lt;/code&gt; resolves to something else, private accounts and non-standard profile types included. Switching on &lt;code&gt;$case&lt;/code&gt; explicitly, the way a discriminated union is meant to be consumed, is what keeps those profiles correctly classified (and reported) instead of vanishing from the run with no trace they were ever requested.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Capture your fixtures from the path you actually ship — impersonation, proxy, headers, all of it — not from whatever curl command was fastest to type while developing. A probe and a production request can return structurally different documents for the identical fact, and a classifier trained on the wrong one fails exactly where it matters most: distinguishing "doesn't exist" from "couldn't be read."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per resolved username: &lt;code&gt;is_public_profile&lt;/code&gt;, &lt;code&gt;display_name&lt;/code&gt;, &lt;code&gt;bio&lt;/code&gt;, &lt;code&gt;website_url&lt;/code&gt;, &lt;code&gt;subscriber_count&lt;/code&gt;, &lt;code&gt;profile_picture_url&lt;/code&gt;, &lt;code&gt;snapcode_image_url&lt;/code&gt;, &lt;code&gt;has_active_story&lt;/code&gt;, &lt;code&gt;story_snap_count&lt;/code&gt;, and &lt;code&gt;story_snaps&lt;/code&gt; (per-snap &lt;code&gt;media_url&lt;/code&gt;, &lt;code&gt;media_preview_url&lt;/code&gt;, &lt;code&gt;media_type&lt;/code&gt;, &lt;code&gt;posted_at&lt;/code&gt;) when the account currently has a live story. &lt;code&gt;page_title&lt;/code&gt; and &lt;code&gt;profile_url&lt;/code&gt; are included for traceability. A private, removed, or nonexistent username is a per-item skip reported in the run's status message — it never fails the whole run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;v1 covers the profile card and current story only — no lenses, no curated/Spotlight highlights, no tagged-tab search, no pagination. &lt;code&gt;story_snaps[].media_url&lt;/code&gt; and &lt;code&gt;media_preview_url&lt;/code&gt; are Snapchat CDN links that &lt;strong&gt;expire&lt;/strong&gt; — Snapchat stories are ephemeral by design, so treat story media as a snapshot at scrape time, not a durable asset you can fetch later. No authenticated views, DMs, or friend-graph data — public profile pages only.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why did a username that clearly exists come back as not found in my own quick test?&lt;/strong&gt;&lt;br&gt;
Check how you tested it. A bare, unimpersonated request to Snapchat's profile URL can return a different response shape than the one this Actor uses in production — the production path is impersonated and proxied specifically because that's the path that reliably resolves real profiles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need a Snapchat account?&lt;/strong&gt;&lt;br&gt;
No — this reads only the public profile page at &lt;code&gt;snapchat.com/@&amp;lt;username&amp;gt;&lt;/code&gt;, no login, no API key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I rely on &lt;code&gt;story_snaps&lt;/code&gt; media URLs after the run finishes?&lt;/strong&gt;&lt;br&gt;
Not indefinitely. They're Snapchat CDN links tied to ephemeral story content and will expire — download or re-host anything you need to keep before the story itself expires on Snapchat's side.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus ~$0.003 per resolved public profile — &lt;strong&gt;~$3.00 per 1,000 profiles&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/snapchat-profile-scraper" rel="noopener noreferrer"&gt;Snapchat Profile Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints, pin the proxy exit the payload actually needs, and retry with backoff — and we test our classifiers against the page our own scraper sees, not the page a quick curl sees.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>apify</category>
      <category>nextjs</category>
    </item>
    <item>
      <title>Per-profile or per-link: the billing unit on a link-in-bio scraper differs by 10x</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:08:05 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/per-profile-or-per-link-the-billing-unit-on-a-link-in-bio-scraper-differs-by-10x-242g</link>
      <guid>https://dev.to/devil_scrapes/per-profile-or-per-link-the-billing-unit-on-a-link-in-bio-scraper-differs-by-10x-242g</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;On a link-in-bio page, "one row per profile" and "one row per link" aren't a formatting preference — they're two different products, and they diverge by roughly 10x on a typical creator profile. A profile with 68 links bills as 1 row under the first model and 68 rows under the second. The &lt;a href="https://apify.com/DevilScrapes/linktree-profile-scraper" rel="noopener noreferrer"&gt;Linktree Profile Scraper&lt;/a&gt; bills per link, on purpose: the link is the unit you actually analyze — "where does this creator send traffic" — not the profile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the billing unit matter more than the price-per-result headline?
&lt;/h2&gt;

&lt;p&gt;Because "price per result" is meaningless until you know what a result is. A scraper billing $1/1,000 &lt;em&gt;profiles&lt;/em&gt; and a scraper billing $4/1,000 &lt;em&gt;links&lt;/em&gt; look like the first is 4x cheaper — until you fetch the same 68-link profile from both and realize the first one gave you one flat JSON blob you'll spend engineering time re-parsing into link-level rows yourself, while the second one already gave you 68 clean, typed, filterable rows. The per-profile scraper's "1,000 results" is 1,000 profiles, which could be tens of thousands of underlying links; the per-link scraper's "1,000 results" is 1,000 analyzable facts. Compare the cost per &lt;em&gt;useful fact&lt;/em&gt;, not the cost per row, and the ranking can flip entirely.&lt;/p&gt;

&lt;p&gt;We chose per-link billing for this Actor because a link-in-bio page's value is almost entirely in its outbound destinations — link type, position, group membership, gate status — not in the container profile around them. &lt;code&gt;link_url&lt;/code&gt;, &lt;code&gt;link_type&lt;/code&gt;, &lt;code&gt;position&lt;/code&gt;, &lt;code&gt;group_id&lt;/code&gt;/&lt;code&gt;group_title&lt;/code&gt;, and &lt;code&gt;gated&lt;/code&gt;/&lt;code&gt;gate_type&lt;/code&gt; are all first-class columns precisely because those are the fields a link-in-bio analysis actually filters and joins on. Denormalizing the profile identity fields (&lt;code&gt;username&lt;/code&gt;, &lt;code&gt;profile_title&lt;/code&gt;, &lt;code&gt;profile_is_verified&lt;/code&gt;, &lt;code&gt;profile_tier&lt;/code&gt;) onto every link row means you never have to join back to a separate profile table to answer "which verified creators link to a specific commerce product."&lt;/p&gt;

&lt;h2&gt;
  
  
  What's the second finding, below the billing question?
&lt;/h2&gt;

&lt;p&gt;Linktree's &lt;code&gt;__NEXT_DATA__&lt;/code&gt; payload — the JSON blob the page ships instead of server-rendered HTML — carries a per-link &lt;code&gt;rules.gate&lt;/code&gt; object: &lt;code&gt;age&lt;/code&gt;, &lt;code&gt;passcode&lt;/code&gt;, &lt;code&gt;password&lt;/code&gt;, or &lt;code&gt;NFT&lt;/code&gt;. That means some links inside a single profile's payload are ones a visitor literally cannot click through without clearing a gate, and the payload doesn't mark that at the top level — it's per link, buried in &lt;code&gt;rules&lt;/code&gt;. A scraper that emits &lt;code&gt;link_url&lt;/code&gt; and &lt;code&gt;link_title&lt;/code&gt; without reading &lt;code&gt;rules.gate&lt;/code&gt; hands you a row that looks exactly like every other clickable link and isn't one. We surface that as &lt;code&gt;gated&lt;/code&gt; (boolean) and &lt;code&gt;gate_type&lt;/code&gt; (which gate) on every row, without attempting to clear the gate — you get an accurate picture of what a real visitor can and can't reach, not a false completeness count.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Before comparing "price per result" across two scrapers on the same target, check what a result &lt;em&gt;is&lt;/em&gt;. A cheaper price on a coarser unit can cost more per fact you actually need, and it's the buyer's job to check the unit — a listing rarely states it plainly.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per link (not per profile): &lt;code&gt;link_id&lt;/code&gt;, &lt;code&gt;link_title&lt;/code&gt;, &lt;code&gt;link_url&lt;/code&gt;, &lt;code&gt;link_type&lt;/code&gt; (&lt;code&gt;CLASSIC&lt;/code&gt;, &lt;code&gt;GROUP&lt;/code&gt;, &lt;code&gt;YOUTUBE_VIDEO&lt;/code&gt;, &lt;code&gt;COMMERCE_PRODUCT&lt;/code&gt;, and any newer type Linktree adds — never dropped), &lt;code&gt;position&lt;/code&gt;, &lt;code&gt;group_id&lt;/code&gt;/&lt;code&gt;group_title&lt;/code&gt; for child-of-group links, &lt;code&gt;thumbnail_url&lt;/code&gt;, &lt;code&gt;gated&lt;/code&gt;/&lt;code&gt;gate_type&lt;/code&gt;, &lt;code&gt;meta_title&lt;/code&gt;/&lt;code&gt;meta_description&lt;/code&gt;, and &lt;code&gt;scraped_at&lt;/code&gt;. Every row also carries the parent profile's &lt;code&gt;profile_title&lt;/code&gt;, &lt;code&gt;profile_description&lt;/code&gt;, &lt;code&gt;profile_avatar_url&lt;/code&gt;, &lt;code&gt;profile_is_verified&lt;/code&gt;, &lt;code&gt;profile_tier&lt;/code&gt;, and &lt;code&gt;profile_has_password&lt;/code&gt;, denormalized so you never need a join to answer a profile-level question from a link-level table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;Public profiles only — no login-required data, analytics, or edit-mode fields. Gated links are flagged, never bypassed; this Actor tells you a gate exists, it doesn't try to get past one. Group membership is a flat &lt;code&gt;group_id&lt;/code&gt;/&lt;code&gt;group_title&lt;/code&gt; join, not a nested tree — reconstruct hierarchy yourself if you need it. Zero-link profiles emit no rows for that profile (logged, not a failure); the run only fails if every requested profile yields zero rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does this bill per link instead of per profile?&lt;/strong&gt;&lt;br&gt;
Because the link is the analyzable unit on a link-in-bio page — you're asking "where does this account send traffic," and a per-profile price hides how many or how few links you actually got for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this need a Linktree account or API key?&lt;/strong&gt;&lt;br&gt;
No — it reads the same public &lt;code&gt;__NEXT_DATA__&lt;/code&gt; JSON payload the profile page itself serves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens to age/passcode/password/NFT-gated links?&lt;/strong&gt;&lt;br&gt;
They're returned with &lt;code&gt;gated: true&lt;/code&gt; and the specific &lt;code&gt;gate_type&lt;/code&gt; set, but the gate itself is never cleared — you know a link exists and is protected, without us attempting to see behind it.&lt;/p&gt;

&lt;p&gt;At 1,000 link rows, that's &lt;strong&gt;$4.20 / 1,000&lt;/strong&gt; ($0.20 start + $4.00 for the rows).&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/linktree-profile-scraper" rel="noopener noreferrer"&gt;Linktree Profile Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints and keep the dataset clean — Pydantic-validated, ISO-8601 timestamps — so the unit you're billed on is exactly the unit you can build on.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>apify</category>
      <category>api</category>
    </item>
    <item>
      <title>We tested the 'always pin the proxy country' rule instead of trusting it. Same data came back from two exit countries.</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:07:34 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/we-tested-the-always-pin-the-proxy-country-rule-instead-of-trusting-it-same-data-came-back-from-d8l</link>
      <guid>https://dev.to/devil_scrapes/we-tested-the-always-pin-the-proxy-country-rule-instead-of-trusting-it-same-data-came-back-from-d8l</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;We tested "pin the proxy exit country on a country-scoped site" before paying for it, instead of assuming it. Same URL, same minute, two proxy tiers: a US-exit datacenter request returned &lt;code&gt;200&lt;/code&gt;, title "Python Jobs und Stellenangebote", &lt;code&gt;totalCount: 1727&lt;/code&gt;, top result inovex GmbH. A &lt;code&gt;RESIDENTIAL&lt;/code&gt;/&lt;code&gt;DE&lt;/code&gt;-pinned request returned &lt;code&gt;200&lt;/code&gt;, the same title, &lt;code&gt;totalCount: 1750&lt;/code&gt;, the &lt;em&gt;same&lt;/em&gt; top result. Identical German job data from two different exit countries. The &lt;a href="https://apify.com/DevilScrapes/stepstone-jobs-scraper" rel="noopener noreferrer"&gt;StepStone Jobs Scraper&lt;/a&gt; ships on the cheaper US datacenter pack by default, because the measurement — not the instinct — decided the proxy tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Doesn't a country-scoped site require a matching exit country?
&lt;/h2&gt;

&lt;p&gt;Often, yes — that's a good enough prior that we start every new country-scoped target by assuming it. It's also a prior worth pricing before you commit to it, because getting it wrong costs twice. First, this account's datacenter pack (&lt;code&gt;BUYPROXIES94952&lt;/code&gt;) is a static US-only pool — pinning &lt;code&gt;apifyProxyCountry: DE&lt;/code&gt; against it doesn't get you German data, it fails the proxy tunnel outright, because the pack simply has no German IPs to hand out. Second, escalating to &lt;code&gt;RESIDENTIAL&lt;/code&gt; to satisfy the pin isn't free: it's a more expensive proxy tier per request, and Apify's own review process treats an Actor that defaults to residential as belonging to the anti-bot-prone class, which comes with more publish scrutiny. So before eating either cost, we ran the actual comparison: identical query, identical minute, US-datacenter exit against &lt;code&gt;RESIDENTIAL&lt;/code&gt;/&lt;code&gt;DE&lt;/code&gt; exit. &lt;code&gt;totalCount&lt;/code&gt; differed by 23 out of 1,727 — well inside normal listing churn for a live job board between two requests — and the top-ranked result was identical on both. stepstone.de does not vary its result set by the requester's exit geography. The country pin was protecting against a failure mode that doesn't exist here.&lt;/p&gt;

&lt;h2&gt;
  
  
  So when should you actually pin a proxy country?
&lt;/h2&gt;

&lt;p&gt;When you've checked, not when you've assumed. A geo pin is a testable hypothesis about the target, not a blanket best practice — some sites genuinely do serve different prices, stock, or catalogs per exit country (we've shipped Actors where that's exactly the finding), and some just don't vary at all. The only way to know which one you're looking at is to run the same query through two tiers and diff the response. Guessing wrong costs you either a failed tunnel (wrong pack, right instinct) or a needlessly expensive and more heavily scrutinized default (right instinct, unnecessary cost) — two request's worth of measurement is cheaper than either.&lt;/p&gt;

&lt;p&gt;Below the proxy question, the parsing question had its own thing worth stating precisely: StepStone's job data lives in &lt;code&gt;window.__PRELOADED_STATE__["app-unifiedResultlist"]&lt;/code&gt;, not in server-rendered HTML you'd need to reconstruct from DOM structure. And pagination distinctness — the thing that actually protects your dataset from silent duplicate rows across pages — was verified directly: page 1 and page 2 share zero job ids, page 2 and page 3 share zero job ids.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A proxy country pin is a hypothesis about the target, and testing it costs one extra request. Skipping the test costs you either a failed tunnel against a pack that can't serve the country, or an unnecessarily expensive and more heavily scrutinized tier protecting against a difference that isn't there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One deduplicated row per job listing across one or more keywords: &lt;code&gt;id&lt;/code&gt; (dedupe key), &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;companyName&lt;/code&gt;, &lt;code&gt;companyId&lt;/code&gt;, &lt;code&gt;location&lt;/code&gt;, &lt;code&gt;datePosted&lt;/code&gt; (ISO-8601), &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;salary&lt;/code&gt; when published, &lt;code&gt;isSponsored&lt;/code&gt;, &lt;code&gt;isHighlighted&lt;/code&gt;, &lt;code&gt;isAnonymous&lt;/code&gt;, &lt;code&gt;postCode&lt;/code&gt;, &lt;code&gt;textSnippet&lt;/code&gt;, &lt;code&gt;skills&lt;/code&gt;, &lt;code&gt;labels&lt;/code&gt;, plus &lt;code&gt;sourceKeyword&lt;/code&gt; / &lt;code&gt;sourceLocation&lt;/code&gt; tagging which input search found each row. Run-wide dedup means the same job id never appears twice, even across pages or across multiple keywords batched into one run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;List-page fields only — full job-description text isn't fetched; use the returned &lt;code&gt;url&lt;/code&gt; for a follow-up detail-page job if you need it. &lt;code&gt;.de&lt;/code&gt; only in this version; &lt;code&gt;.at&lt;/code&gt; and &lt;code&gt;.ch&lt;/code&gt; StepStone variants aren't covered. A narrow keyword + location combination can legitimately return zero rows, and that's a successful run, not a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do you pin the proxy exit country to Germany by default?&lt;/strong&gt;&lt;br&gt;
No. We measured it: US-datacenter and &lt;code&gt;RESIDENTIAL&lt;/code&gt;/&lt;code&gt;DE&lt;/code&gt; returned the same &lt;code&gt;totalCount&lt;/code&gt;, the same top result, and the same title for an identical query. stepstone.de doesn't vary by exit geo, so the default ships on the cheaper, less-scrutinized US datacenter pack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need a StepStone account or API key?&lt;/strong&gt;&lt;br&gt;
No — this reads StepStone's public search result pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I search multiple keywords in one run?&lt;/strong&gt;&lt;br&gt;
Yes — pass an array to &lt;code&gt;keywords&lt;/code&gt;; every row is tagged with the keyword (and location) that produced it, and dedup applies across the whole run, not per keyword.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.003 per deduplicated job row — &lt;strong&gt;$3.20 per 1,000 results&lt;/strong&gt;. A zero-match search still succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/stepstone-jobs-scraper" rel="noopener noreferrer"&gt;StepStone Jobs Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints, retry with backoff, and route every request through Apify Proxy — and before we spend your money on a more expensive proxy tier, we measure whether the target actually needs it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>apify</category>
      <category>proxy</category>
    </item>
    <item>
      <title>nbHits said 6,360. The API only let us page through 1,000 of them.</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:07:03 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/nbhits-said-6360-the-api-only-let-us-page-through-1000-of-them-27nb</link>
      <guid>https://dev.to/devil_scrapes/nbhits-said-6360-the-api-only-let-us-page-through-1000-of-them-27nb</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;Querying Welcome to the Jungle's search index with no filter reports &lt;code&gt;nbHits: 6360&lt;/code&gt; in the response envelope — then refuses to page past result 1,000. Page 40 of that same query comes back &lt;code&gt;HTTP 200&lt;/code&gt; with a body that reads &lt;code&gt;{"message": "you can only fetch the 1000 hits for this query"}&lt;/code&gt;. That's not a rate limit and it's not a block; it's Algolia's hard pagination ceiling, and it's baked into every Algolia-backed site, not just this one. The &lt;a href="https://apify.com/DevilScrapes/welcome-to-the-jungle-jobs-scraper" rel="noopener noreferrer"&gt;Welcome to the Jungle Jobs Scraper&lt;/a&gt; treats that ceiling as a first-class fact: it surfaces it in the run's status message and gives you &lt;code&gt;query&lt;/code&gt;, &lt;code&gt;country&lt;/code&gt;, &lt;code&gt;contractType&lt;/code&gt;, and &lt;code&gt;remote&lt;/code&gt; as facets so you can slice a 6,360-hit corpus into several sub-1,000 queries instead of silently losing 84% of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does a 200 response contain an error?
&lt;/h2&gt;

&lt;p&gt;Because Algolia's own index — the one Welcome to the Jungle's frontend calls directly from the browser — enforces a 1,000-result cap per query on the free/standard tier, and it enforces it at the &lt;em&gt;pagination&lt;/em&gt; layer, not the &lt;em&gt;response status&lt;/em&gt; layer. The first page of any query looks completely healthy: &lt;code&gt;200&lt;/code&gt;, real JSON, a &lt;code&gt;nbHits&lt;/code&gt; field truthfully reporting the full match count, and 20 or so listings in &lt;code&gt;hits&lt;/code&gt;. Nothing about that first page tells you the query has more matches than the index will ever hand you. Only when you request page 40+ (at 25 hits/page, that's the 1,000-result boundary) does the response change shape — still &lt;code&gt;200&lt;/code&gt;, but now the payload &lt;em&gt;is&lt;/em&gt; the error: &lt;code&gt;you can only fetch the 1000 hits for this query&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A scraper that treats "response received, parse hits" as its only success condition will parse that error body as if it were a results page, get an empty or malformed &lt;code&gt;hits&lt;/code&gt; array, and stop. The run finishes green. The dataset looks complete because nothing in the run log says otherwise — it just quietly tops out at 1,000 rows on any query broad enough to have more. That's a worse failure than a crash: a crash gets investigated, a silently-truncated dataset gets shipped to a customer's pipeline and trusted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is the ceiling really about the search index and not the target blocking us?
&lt;/h2&gt;

&lt;p&gt;Yes, and the distinction matters because they call for opposite fixes. A block calls for better fingerprinting, proxy rotation, or backoff. A 1,000-hit ceiling calls for narrower queries. We confirmed which one we were looking at by checking the &lt;em&gt;same&lt;/em&gt; query with a tighter facet: narrowing by &lt;code&gt;country=FR&lt;/code&gt; plus a keyword cut &lt;code&gt;nbHits&lt;/code&gt; from 6,360 to under 1,000, and paging that narrower query all the way through returned every result with no error body — same site, same session, same browser fingerprint, no block anywhere in the run. The ceiling moved when the query got smaller; a block wouldn't have.&lt;/p&gt;

&lt;p&gt;The second finding sits one layer below the pagination cap: the endpoint gates on the request's &lt;code&gt;Origin&lt;/code&gt; / &lt;code&gt;Referer&lt;/code&gt; header, not on the User-Agent or TLS fingerprint. A request with no origin header returns &lt;code&gt;403 {"message":"Method not allowed with this referer"}&lt;/code&gt; regardless of which browser you're impersonating. Both &lt;code&gt;chrome131&lt;/code&gt; and &lt;code&gt;firefox133&lt;/code&gt; impersonation profiles pass cleanly once the header is set to the site's own origin — the fingerprint was never the gate here, the header was. Worth knowing before you spend a debugging session rotating TLS profiles against a header check.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A &lt;code&gt;200&lt;/code&gt; status code tells you the transport succeeded, not that the payload is a results page. On any Algolia-backed search, check whether page N+1 still parses as a hit list before you trust that a broad query returned everything it claimed to have.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per job listing, with the site's own facets available as scrape-time filters (not just export-time ones): &lt;code&gt;objectId&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;slug&lt;/code&gt;, &lt;code&gt;jobUrl&lt;/code&gt;, &lt;code&gt;companyName&lt;/code&gt;, &lt;code&gt;companySlug&lt;/code&gt;, &lt;code&gt;offices&lt;/code&gt; (city/state/country per office), &lt;code&gt;contractType&lt;/code&gt; + &lt;code&gt;contractTypeLabel&lt;/code&gt;, &lt;code&gt;remote&lt;/code&gt; policy, &lt;code&gt;publishedAt&lt;/code&gt; (ISO-8601), salary range and currency when the recruiter published one, sectors, profession, and department. Set &lt;code&gt;query&lt;/code&gt;, &lt;code&gt;country&lt;/code&gt;, &lt;code&gt;contractType&lt;/code&gt;, or &lt;code&gt;remote&lt;/code&gt; — alone or combined — to keep any single search's true match count under the 1,000-hit ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;p&gt;The 1,000-hit-per-query cap is real and not something we bypass — you narrow around it with facets, you don't get past it with retries or a different proxy. Full job-page HTML (benefits copy, application forms) isn't included; the search index already carries everything relevant to filtering and ranking, and a detail-page fetch would be a different, heavier Actor. Company-profile enrichment (reviews, culture pages) is out of scope for v1.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why did my broad search return exactly 1,000 rows when the listing count looked higher?&lt;/strong&gt;&lt;br&gt;
You hit the Algolia index's pagination ceiling. Split the query by &lt;code&gt;country&lt;/code&gt;, &lt;code&gt;contractType&lt;/code&gt;, or &lt;code&gt;remote&lt;/code&gt; and run it again — each narrower query gets its own 1,000-hit budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this need a login or API key?&lt;/strong&gt;&lt;br&gt;
No — listings are public, and we call the same index endpoint the site's own frontend calls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did my run return zero rows instead of failing?&lt;/strong&gt;&lt;br&gt;
A query + filter combination with no matches is a successful search, not a broken one. The run status message says exactly what was searched.&lt;/p&gt;

&lt;p&gt;$0.20 per run plus $0.0025 per result — &lt;strong&gt;$2.70 per 1,000 results&lt;/strong&gt;. A zero-match search still succeeds and costs only the start fee.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/welcome-to-the-jungle-jobs-scraper" rel="noopener noreferrer"&gt;Welcome to the Jungle Jobs Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We rotate browser fingerprints, retry with backoff, and route every request through Apify Proxy — and when a target's own search index caps what any single query can return, we tell you in the run status instead of shipping you a quietly truncated dataset.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webscraping</category>
      <category>apify</category>
      <category>algolia</category>
    </item>
    <item>
      <title>Filter the CFTC COT report for WHEAT and you get three different markets</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 03:32:39 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/filter-the-cftc-cot-report-for-wheat-and-you-get-three-different-markets-3l7n</link>
      <guid>https://dev.to/devil_scrapes/filter-the-cftc-cot-report-for-wheat-and-you-get-three-different-markets-3l7n</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;Ask the CFTC's Commitments of Traders data for &lt;code&gt;commodity_name = 'WHEAT'&lt;/code&gt; on the 2026-09-15 report and you don't get one market back. You get three: soft red winter wheat on the Chicago Board of Trade, hard red winter wheat on the same exchange, and hard red spring wheat on the MIAX Futures Exchange. They are different contracts with different open interest and different traders, and speculators can be positioned in opposite directions in them. Add them up and you get a number that describes no real market.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://apify.com/DevilScrapes/cftc-commitments-of-traders-scraper" rel="noopener noreferrer"&gt;CFTC Commitments of Traders Scraper&lt;/a&gt; pulls the COT reports straight from the CFTC's own public data portal. Most of the engineering is in the dull parts. Most of what you need to know to use the data well is in that one gotcha.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does one commodity filter return several markets?
&lt;/h2&gt;

&lt;p&gt;Because &lt;code&gt;commodity_name&lt;/code&gt; is a family, not a contract. In the CFTC's legacy dataset, the wheat contracts all carry &lt;code&gt;commodity_name: "WHEAT"&lt;/code&gt;. The field that tells them apart is &lt;code&gt;market_and_exchange_names&lt;/code&gt; (&lt;code&gt;WHEAT-SRW - CHICAGO BOARD OF TRADE&lt;/code&gt;), and the one that identifies a contract for good is &lt;code&gt;cftc_contract_market_code&lt;/code&gt; (&lt;code&gt;001602&lt;/code&gt; for SRW, &lt;code&gt;001612&lt;/code&gt; for HRW).&lt;/p&gt;

&lt;p&gt;Our own cloud QA run shows why this matters. It used the Actor's default input (wheat, Chicago Board of Trade, report dates 2026-06-30 to 2026-09-15) and returned 24 rows: 12 weekly reports × 2 contracts. Speculator net positioning (non-commercial long minus short), straight from those rows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;report date   contract          open interest   speculator net
2026-06-30    WHEAT-SRW (CBOT)        406,723          -55,004
2026-06-30    WHEAT-HRW (CBOT)        260,913          -10,339
2026-09-15    WHEAT-SRW (CBOT)        485,138           +1,228
2026-09-15    WHEAT-HRW (CBOT)        310,836          +30,335
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Speculators ended the window net long in both contracts, but by very different amounts. HRW moved to +30,335 while SRW barely crossed zero. A chart summed on &lt;code&gt;commodity_name&lt;/code&gt; blurs those two into one line. Group on &lt;code&gt;cftc_contract_market_code&lt;/code&gt; and you keep them apart.&lt;/p&gt;

&lt;p&gt;So the Actor's &lt;code&gt;commodity&lt;/code&gt; input is an exact match on the CFTC's family name, and &lt;code&gt;market&lt;/code&gt; is a substring match on the market-and-exchange name. You can use &lt;code&gt;market&lt;/code&gt; to narrow a family down to the contract you actually trade.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does the raw CFTC API actually hand you?
&lt;/h2&gt;

&lt;p&gt;Strings. This is a real row from the CFTC's Socrata endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"260630001612F"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"commodity_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"WHEAT"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"open_interest_all"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"260913"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"noncomm_positions_long_all"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"66398"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"report_date_as_yyyy_mm_dd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-06-30T00:00:00.000"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every position count is a JSON string, and the field named &lt;code&gt;..._yyyy_mm_dd&lt;/code&gt; holds a full timestamp. Both are small problems, and both break a spreadsheet or a pandas pipeline quietly: string columns sort lexically, and &lt;code&gt;"66398" &amp;gt; "100000"&lt;/code&gt;. The Actor coerces every position field to a number and trims the date to real &lt;code&gt;YYYY-MM-DD&lt;/code&gt;. If a value won't parse, it's set to &lt;code&gt;null&lt;/code&gt; and a warning goes in the log. We don't guess a number for you.&lt;/p&gt;

&lt;p&gt;The row &lt;code&gt;id&lt;/code&gt; isn't random, either. &lt;code&gt;260630001612F&lt;/code&gt; is the report date (&lt;code&gt;260630&lt;/code&gt;), the contract code (&lt;code&gt;001612&lt;/code&gt;), and &lt;code&gt;F&lt;/code&gt; for futures-only. It's stable, so you can dedupe on it when you re-pull overlapping windows.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much history is there?
&lt;/h2&gt;

&lt;p&gt;A lot. When we queried the legacy futures-only dataset it held 289,274 rows, running from 1986-01-15 to the 2026-09-15 report. The Actor covers all six COT variants: Legacy, Disaggregated, and Traders in Financial Futures, each as futures-only or futures-and-options-combined. You pick one with the &lt;code&gt;reportType&lt;/code&gt; field. A second cloud run backfilled wheat from 2024 onward, returning 100 rows with 100 distinct IDs: 50 consecutive weekly reports for each of the two CBOT contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One row per market per report date, with the Socrata &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;report_type&lt;/code&gt;, market-and-exchange name, report date, commodity name, contract market code, total open interest, non-commercial and commercial long and short positions, week-over-week change in open interest, and total trader count. You can filter by report type, commodity, market, and date range. There's no API key. We handle the pagination, retries with backoff on &lt;code&gt;408&lt;/code&gt; / &lt;code&gt;429&lt;/code&gt; / &lt;code&gt;503&lt;/code&gt;, and the type coercion, so the dataset exports to JSON, CSV, or Excel ready to chart.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A capped run returns the oldest weeks, not the newest.&lt;/strong&gt; The wheat backfill capped at 100 rows over 2024-01-01 to 2026-09-15 returned January through December 2024. To get this week's report, set &lt;code&gt;dateFrom&lt;/code&gt; close to today rather than relying on the cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pct_of_open_interest_all&lt;/code&gt; is always 100&lt;/strong&gt; in the legacy dataset. We checked all 289,274 rows and it held one value. It's passed through for completeness, but it tells you nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CFTC's schedule is the Actor's schedule.&lt;/strong&gt; Reports come out weekly and describe the prior Tuesday's positions. The Actor can't return a report before the CFTC publishes it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Raw positions only.&lt;/strong&gt; The Actor doesn't compute net positioning, COT index, or z-scores. Those take one subtraction or a rolling window downstream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use the Legacy report for position fields.&lt;/strong&gt; The v1 row shape follows the Legacy report. The Disaggregated and TFF datasets split traders into different categories (managed money, swap dealers, leveraged funds and so on) and don't publish the Legacy commercial/non-commercial columns. On those report types, you get open interest, the week-over-week change, and trader count, but the long/short position fields come back &lt;code&gt;null&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need a CFTC or Socrata API key?&lt;/strong&gt;&lt;br&gt;
No. The CFTC's public reporting portal is keyless open data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does it cost?&lt;/strong&gt;&lt;br&gt;
$3.20 per 1,000 rows under Pay-Per-Event: a $0.20 start fee plus $0.003 per row written to your dataset. A year of weekly reports for one contract is 52 rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did my &lt;code&gt;WHEAT&lt;/code&gt; query return more markets than I expected?&lt;/strong&gt;&lt;br&gt;
Because &lt;code&gt;WHEAT&lt;/code&gt; is a family of contracts. Add a &lt;code&gt;market&lt;/code&gt; filter, or group on &lt;code&gt;cftc_contract_market_code&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if my filters match nothing?&lt;/strong&gt;&lt;br&gt;
The run succeeds with zero rows and a status message listing the filters it used, and you're charged only the start fee. If the CFTC portal can't be reached at all, the run fails loudly instead of returning a quiet, empty success.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/cftc-commitments-of-traders-scraper" rel="noopener noreferrer"&gt;CFTC Commitments of Traders Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We handle the paging, retries and type coercion, so the numbers in your dataset are numbers.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>opendata</category>
      <category>finance</category>
      <category>api</category>
    </item>
    <item>
      <title>Our two-state UCC run returned 200 rows, 127 filings, and zero from Colorado</title>
      <dc:creator>Devil Scrapes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 03:32:08 +0000</pubDate>
      <link>https://dev.to/devil_scrapes/our-two-state-ucc-run-returned-200-rows-127-filings-and-zero-from-colorado-12l4</link>
      <guid>https://dev.to/devil_scrapes/our-two-state-ucc-run-returned-200-rows-127-filings-and-zero-from-colorado-12l4</guid>
      <description>&lt;h2&gt;
  
  
  Quick answer
&lt;/h2&gt;

&lt;p&gt;Before the &lt;a href="https://apify.com/DevilScrapes/ucc-lien-filing-scraper" rel="noopener noreferrer"&gt;UCC Lien Filing Scraper&lt;/a&gt; ever went public, a deep cloud run asked it for Connecticut &lt;strong&gt;and&lt;/strong&gt; Colorado filings from the last 120 days, capped at 200 rows. The run came back &lt;code&gt;SUCCEEDED&lt;/code&gt; with 200 rows. Every one of them was from Connecticut. Colorado — which had real filings in that exact window — contributed zero. And those 200 rows described only 127 distinct filings.&lt;/p&gt;

&lt;p&gt;Both numbers were telling us something. One was a bug. The other was the data being honest about its own shape, and the fix there was to our copy, not our code.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does a two-state run return rows from only one state?
&lt;/h2&gt;

&lt;p&gt;By handing the whole budget to whoever asks first. The first version of the run loop gave the entire remaining &lt;code&gt;maxResults&lt;/code&gt; to the first state it fetched. Connecticut went first, had well over 200 matching records in 120 days, and spent the whole cap before Colorado was ever queried.&lt;/p&gt;

&lt;p&gt;Nothing failed. There was no error to log. The run did exactly what it was told, and the result looked like "Colorado had nothing this quarter" — which is a plausible-sounding lie. We only caught it because we checked Colorado's portal directly for the same window and it returned real filings. Running the Connecticut fetch alone with a budget of 200 then reproduced the exact 200-rows / 127-filings result, which pinned it on the cap, not on a broken Colorado adapter.&lt;/p&gt;

&lt;p&gt;The fix: &lt;code&gt;maxResults&lt;/code&gt; is now split evenly across the states you request, up front, and any share a state doesn't use rolls over to the next one so the cap is never wasted. The same failing input now returns 100 Connecticut rows and 100 Colorado rows. Our post-fix cloud QA runs (&lt;code&gt;daysBack: 14&lt;/code&gt;, &lt;code&gt;maxResults: 25&lt;/code&gt;) landed 13 CT and 12 CO rows each time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why were 200 rows only 127 filings?
&lt;/h2&gt;

&lt;p&gt;Because a UCC filing can name more than one debtor, and Connecticut publishes one record &lt;strong&gt;per debtor per filing&lt;/strong&gt;. We checked filing &lt;code&gt;0005384782&lt;/code&gt; against &lt;code&gt;data.ct.gov&lt;/code&gt; by hand: four genuine source records, same filing number, four different debtors. In that same run, one filing — &lt;code&gt;0005384859&lt;/code&gt; — accounted for 43 rows on its own.&lt;/p&gt;

&lt;p&gt;Colorado was the side out of step. Its adapter joined each filing to its debtor dataset and kept only the first debtor it saw, silently discarding the rest. So the two states disagreed on what a row even meant.&lt;/p&gt;

&lt;p&gt;We had a choice: collapse Connecticut into one row per filing with a debtor list, or expand Colorado to one row per debtor. We expanded Colorado. Connecticut's native shape is per-debtor, so aggregating would mean inventing a structure the source doesn't publish. And the people who buy this data — lenders, collections shops, lead-gen teams — want each debtor as its own row with its own address, ready to import as a contact.&lt;/p&gt;

&lt;p&gt;That decision is visible on your bill, so we put it next to the price rather than in a footnote: a row is a filing–debtor pair. A filing with three debtors is three rows and three result events. Rows that share a &lt;code&gt;filing_id&lt;/code&gt; are the same lien.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does Colorado need a three-way join?
&lt;/h2&gt;

&lt;p&gt;Connecticut ships one flat dataset. Colorado splits the same information across three linked Socrata datasets — a filings index, a debtor table, and a secured-party table — joined on &lt;code&gt;fileid&lt;/code&gt;. The filings index alone held 2,594,585 rows when we measured it, so an unbounded pull is never acceptable. Every query is date-bounded server-side (&lt;code&gt;daysBack&lt;/code&gt;, or &lt;code&gt;dateFrom&lt;/code&gt;/&lt;code&gt;dateTo&lt;/code&gt;) and paged.&lt;/p&gt;

&lt;p&gt;Here is a real Colorado row from a cloud run, after the join:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"CO"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"filing_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2594901"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"filing_date"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-08T00:00:00.000"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"lapse_date"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2031-09-08T00:00:00.000"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"filing_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ucc"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"filing_description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"UCC financing statement"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"transaction_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"original"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"debtor_address"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5045 BEACH CT"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"debtor_city"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"DENVER"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"secured_party_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GoodLeap, LLC"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"secured_party_state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"CA"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"active"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Some Colorado filings join to nothing — a termination amendment we sampled (&lt;code&gt;2595634&lt;/code&gt;) had no debtor or secured-party rows at all. Those still ship, with the missing fields set to &lt;code&gt;null&lt;/code&gt;, instead of one empty join sinking the whole run.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Actor gives you
&lt;/h2&gt;

&lt;p&gt;One normalized row per filing–debtor pair across Connecticut and Colorado in a single run: &lt;code&gt;state&lt;/code&gt;, &lt;code&gt;filing_id&lt;/code&gt;, &lt;code&gt;filing_date&lt;/code&gt;, &lt;code&gt;lapse_date&lt;/code&gt;, &lt;code&gt;filing_type&lt;/code&gt;, &lt;code&gt;transaction_type&lt;/code&gt; / &lt;code&gt;is_amendment&lt;/code&gt;, debtor name and address, secured party name and address, a derived &lt;code&gt;status&lt;/code&gt; (&lt;code&gt;active&lt;/code&gt;, &lt;code&gt;lapsed&lt;/code&gt;, &lt;code&gt;terminated&lt;/code&gt;, &lt;code&gt;unknown&lt;/code&gt;), and a &lt;code&gt;source_record_url&lt;/code&gt; that queries the state portal for that exact filing. You can filter by debtor name with &lt;code&gt;debtorNameContains&lt;/code&gt;. There is no API key and no login. We handle the pacing, retries with backoff on &lt;code&gt;408&lt;/code&gt; / &lt;code&gt;429&lt;/code&gt; / &lt;code&gt;5xx&lt;/code&gt;, and the join between the two states' schemas, so what you get is one dataset rather than two government schemas to reconcile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations 🚧
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two states only.&lt;/strong&gt; Connecticut and Colorado. It is not a national UCC search. If a debtor only filed in a third state, you correctly get zero rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Colorado individual debtors come back without a name.&lt;/strong&gt; Colorado stores business debtors in &lt;code&gt;organizationname&lt;/code&gt; and people in separate &lt;code&gt;firstname&lt;/code&gt; / &lt;code&gt;lastname&lt;/code&gt; fields, and v1 reads only the first. In our two post-fix QA samples, 6 of 12 and 8 of 12 Colorado rows had an empty &lt;code&gt;debtor_name&lt;/code&gt;. The address is still there and business debtors are unaffected. We found this while writing this post and it's the next fix on the list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rows are not sorted newest-first.&lt;/strong&gt; In both 14-day QA runs, every row was dated the first day of the window. If &lt;code&gt;maxResults&lt;/code&gt; is smaller than the window, you get the oldest filings in it. To catch fresh filings, shrink &lt;code&gt;daysBack&lt;/code&gt;; don't just lower the cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dates are Socrata timestamps&lt;/strong&gt; (&lt;code&gt;2026-09-08T00:00:00.000&lt;/code&gt;), not bare &lt;code&gt;YYYY-MM-DD&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;UCC filings only. Colorado's IRS and hospital lien types are out of scope, and so are scanned UCC-1 images and collateral text.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need an API key or a state portal login?&lt;/strong&gt;&lt;br&gt;
No. Both states publish this as public, keyless open data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does it cost?&lt;/strong&gt;&lt;br&gt;
$5.20 per 1,000 rows under Pay-Per-Event: a $0.20 start fee plus $0.005 per row written to your dataset. A row is a filing–debtor pair, so budget by debtors, not filings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I monitor new liens against a specific business?&lt;/strong&gt;&lt;br&gt;
Yes. Set &lt;code&gt;debtorNameContains&lt;/code&gt; and a short &lt;code&gt;daysBack&lt;/code&gt;, then run it on a schedule. A new financing statement against a business means it just took on secured credit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if my search matches nothing?&lt;/strong&gt;&lt;br&gt;
The run succeeds with zero rows and a status message saying what was searched. An empty result is still an answer.&lt;/p&gt;

&lt;p&gt;→ &lt;a href="https://apify.com/DevilScrapes/ucc-lien-filing-scraper" rel="noopener noreferrer"&gt;UCC Lien Filing Scraper on Apify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/DevilScrapes" rel="noopener noreferrer"&gt;Devil Scrapes&lt;/a&gt;. We handle the pacing, retries and cross-state schema joins, and when a run returns something that looks plausible, we check it against the source before we believe it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>opendata</category>
      <category>webscraping</category>
      <category>api</category>
    </item>
  </channel>
</rss>
