<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Spicrawl</title>
    <description>The latest articles on DEV Community by Spicrawl (spicrawl).</description>
    <link>https://dev.to/spicrawl</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F15146%2Fabcba039-98bf-489b-aae1-77c2481fea45.png</url>
      <title>DEV Community: Spicrawl</title>
      <link>https://dev.to/spicrawl</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/spicrawl"/>
    <language>en</language>
    <item>
      <title>I Tested 16 Job Sources With a Scraping API, here's What Worked!</title>
      <dc:creator>Mike Alex</dc:creator>
      <pubDate>Fri, 09 Oct 2026 12:14:05 +0000</pubDate>
      <link>https://dev.to/spicrawl/i-tested-16-job-sources-with-a-scraping-api-heres-what-worked-4m51</link>
      <guid>https://dev.to/spicrawl/i-tested-16-job-sources-with-a-scraping-api-heres-what-worked-4m51</guid>
      <description>&lt;p&gt;&lt;strong&gt;I Tested 16 Job Sources With a Scraping API. Here's What Worked&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Job hunting can get messy quickly.&lt;/p&gt;

&lt;p&gt;You find a role on LinkedIn, another on a startup's career page, and a few more on remote job boards. Soon, you have dozens of tabs open and no single place to track everything 😅&lt;/p&gt;

&lt;p&gt;I started thinking about building a small tool that collects job listings from different sources and organizes them in one place.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvm7c3gmic3mx5y4m9zb3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvm7c3gmic3mx5y4m9zb3.gif" alt="coding" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As a product manager working with Spicrawl, I wanted to explore how its web scraping API could help.&lt;/p&gt;

&lt;p&gt;In our October 2026 tests, &lt;strong&gt;15 out of 16 job sources returned usable listings.&lt;/strong&gt; Here's what we learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Not every website needs a browser
&lt;/h2&gt;

&lt;p&gt;Some job sources provide data through public APIs or RSS feeds. Others rely on JavaScript to display listings.&lt;/p&gt;

&lt;p&gt;Here's a snapshot of our results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What worked&lt;/th&gt;
&lt;th&gt;Credits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Remote OK API&lt;/td&gt;
&lt;td&gt;100 job items&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remotive API&lt;/td&gt;
&lt;td&gt;Structured job data&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LinkedIn single posting&lt;/td&gt;
&lt;td&gt;Full job description&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LinkedIn job search&lt;/td&gt;
&lt;td&gt;Rendered listings&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Y Combinator jobs&lt;/td&gt;
&lt;td&gt;Rendered listings&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Glassdoor jobs&lt;/td&gt;
&lt;td&gt;Rendered listings&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Indeed search&lt;/td&gt;
&lt;td&gt;Listings after scrolling&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are results from our tests, not guaranteed results for every request.&lt;/p&gt;

&lt;p&gt;The takeaway? &lt;strong&gt;Start with the simplest request and only add browser rendering when you need it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Collect job listings with Spicrawl
&lt;/h2&gt;

&lt;p&gt;You'll need a Spicrawl API key, which you can get through &lt;a href="https://spicrawl.com/" rel="noopener noreferrer"&gt;Spicrawl&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For a JavaScript-rendered job page, you can use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.spicrawl.com/v1/scrape &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$SPICRAWL_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "url": "https://www.ycombinator.com/jobs",
    "js_render": true,
    "cache": false
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This loads the page in a browser so its JavaScript can build the listings.&lt;/p&gt;

&lt;p&gt;In our October tests, this returned around 28 KB of content for three credits.&lt;/p&gt;

&lt;p&gt;For a source that already provides JSON, you can use a plain fetch instead. For example, Remote OK returned 100 items for one credit.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What about pages that load more jobs while scrolling?
&lt;/h2&gt;

&lt;p&gt;This was one of the more interesting results from our tests.&lt;/p&gt;

&lt;p&gt;Indeed returned only a tiny, 60-byte page when we tried JavaScript rendering alone. The request technically succeeded, but the job listings weren't there.&lt;/p&gt;

&lt;p&gt;We added scrolling actions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://www.indeed.com/jobs?q=data+engineer&amp;amp;l=Remote"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"js_render"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"wait"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"actions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"scroll"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"to_bottom"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"scroll"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"to_bottom"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response grew to around 68 KB of listings. This request cost eight credits because browser actions use the full-browser rendering tier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A successful HTTP response doesn't always mean you received useful data.&lt;/strong&gt; Check the returned content and target status before saving results.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Turn scraped pages into a job tracker
&lt;/h2&gt;

&lt;p&gt;Once you've collected the page content, you can extract fields such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Job title and company&lt;/li&gt;
&lt;li&gt;Location and employment type&lt;/li&gt;
&lt;li&gt;Salary, when available&lt;/li&gt;
&lt;li&gt;Required skills&lt;/li&gt;
&lt;li&gt;Application URL&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Spicrawl's &lt;code&gt;autoparse: true&lt;/code&gt; option can return structured data embedded in a page, including schema.org &lt;code&gt;JobPosting&lt;/code&gt; data when available.&lt;/p&gt;

&lt;p&gt;You can store the results in a database or spreadsheet and build features such as job filters, application tracking, and AI-powered comparisons against your skills.&lt;/p&gt;

&lt;p&gt;Just remember that not every page contains every field, and some job listings may be outdated.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. What we learned
&lt;/h2&gt;

&lt;p&gt;A few lessons stood out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use APIs and feeds when available.&lt;/strong&gt; They're often cheaper and easier to process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Render JavaScript when needed.&lt;/strong&gt; Some listing pages don't contain their data in the initial HTML.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the actual response.&lt;/strong&gt; Empty pages can still return HTTP 200.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Respect website rules.&lt;/strong&gt; Check terms and &lt;code&gt;robots.txt&lt;/code&gt;, follow rate limits, and collect only the data you need.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What would you build?
&lt;/h2&gt;

&lt;p&gt;A job collector is just one use case. The same idea could help organize internships, scholarships, research opportunities, or listings for a job aggregator.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you were building a job-hunting assistant, what would you want it to do first?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Find relevant jobs, compare opportunities, match jobs to your skills, or track applications?&lt;/p&gt;

&lt;p&gt;I'm exploring these use cases while working on Spicrawl, and I'd genuinely love to hear your ideas.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Want to explore more job boards?&lt;/strong&gt; Check out the &lt;a href="https://docs.spicrawl.com/use-cases/job-postings" rel="noopener noreferrer"&gt;Spicrawl job-postings scraping guide&lt;/a&gt; for tested API requests, credit costs, JavaScript rendering, and tips for scraping different job sources.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://spicrawl.com/" rel="noopener noreferrer"&gt;Explore Spicrawl&lt;/a&gt;  currently free during beta, with no credit card required.&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>webdev</category>
      <category>discuss</category>
    </item>
  </channel>
</rss>
