<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Prime Sieve</title>
    <description>The latest articles on DEV Community by Prime Sieve (@primesieve).</description>
    <link>https://dev.to/primesieve</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4081693%2F1d5bd6c2-071c-4b66-9914-b251a6acbcc6.png</url>
      <title>DEV Community: Prime Sieve</title>
      <link>https://dev.to/primesieve</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/primesieve"/>
    <language>en</language>
    <item>
      <title>I Built an AI Agent in a Weekend — Here's the Stack That Gets You Hired in 2026</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Thu, 10 Sep 2026 02:15:32 +0000</pubDate>
      <link>https://dev.to/primesieve/i-built-an-ai-agent-in-a-weekend-heres-the-stack-that-gets-you-hired-in-2026-3fj4</link>
      <guid>https://dev.to/primesieve/i-built-an-ai-agent-in-a-weekend-heres-the-stack-that-gets-you-hired-in-2026-3fj4</guid>
      <description>&lt;h1&gt;
  
  
  I Built an AI Agent in a Weekend — Here's the Stack That Gets You Hired in 2026
&lt;/h1&gt;

&lt;p&gt;The job posting said: "Experience building agentic AI workflows required."&lt;/p&gt;

&lt;p&gt;I had none. I'd shipped APIs, worked with LLM APIs, done my share of prompt engineering — but an &lt;em&gt;agent&lt;/em&gt;? Something that plans, calls functions, retries, and produces a verifiable output? Never built one.&lt;/p&gt;

&lt;p&gt;So I built one in a weekend. Not a demo. A deployed system with a real endpoint, real error handling, a README a hiring manager would actually open, and an evaluation result I could quote in an interview.&lt;/p&gt;

&lt;p&gt;This article is that weekend, compressed. If you're a working developer who's read one too many "AI is eating the world" posts and wants the concrete version — this is the concrete version. You'll finish with a project you can point to, and the words to sell it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "agentic experience" is suddenly on every posting
&lt;/h2&gt;

&lt;p&gt;Quick context, because it changes what you build. The 2026 tech job market isn't one market — it's two. General software engineering is cold; AI-adjacent engineering is red-hot. LinkedIn ranks &lt;strong&gt;AI engineer as the #1 fastest-growing job in the US for the second consecutive year&lt;/strong&gt;. Stanford's 2026 AI Index recorded agentic AI job listings growing &lt;strong&gt;10,854% year over year&lt;/strong&gt;. Indeed's listing index for machine learning engineers sits near &lt;strong&gt;159&lt;/strong&gt; against a baseline of 100 from February 2020, while general software engineer listings sit near &lt;strong&gt;51&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The demand spike isn't for research scientists. It's for engineers who can &lt;em&gt;wire LLMs into products&lt;/em&gt; — companies have real systems and they need people who can build the plumbing around a model. "Agentic" is the current name for that plumbing.&lt;/p&gt;

&lt;p&gt;Here's the encouraging part, and the whole reason this weekend was possible: &lt;strong&gt;an agent is just a service with a loop in it.&lt;/strong&gt; If you can build a REST API, you already know 80% of the mechanics. What's left is a specific five-layer structure — and it fits in a weekend.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five-layer stack hiring managers actually check
&lt;/h2&gt;

&lt;p&gt;Article after article says "learn agents," then lists buzzwords. Here's the concrete version — the five layers every agentic system has, whether it's a startup's internal bot or a FAANG feature:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model access&lt;/strong&gt; — an LLM API (or hosted model) behind a thin service layer. Cost, rate limits, and a fallback understood.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context management&lt;/strong&gt; — a retrieval layer (embeddings + vector store, or a well-structured prompt cache). Raw context windows don't scale to real products.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability wiring&lt;/strong&gt; — letting the model invoke &lt;em&gt;your&lt;/em&gt; functions: search, CRUD, external APIs. This is where most "AI features" actually live.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails&lt;/strong&gt; — validating model output before it touches a database or a customer. Schemas, rejection rules, retry policies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation&lt;/strong&gt; — a fixed scored dataset so you can say "v2 answers correctly 92% of the time vs. 84% for v1" instead of "it feels better."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Notice something: four of those five are ordinary backend engineering in a new costume. The one genuinely new piece is layer 3 — giving the model a way to act. That's what "agentic" means to most hiring managers: &lt;em&gt;the model can call your backend and the loop runs until the job is done.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Pick one weekend, build all five layers, and you've demonstrated the entire skill a posting like mine was asking for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing the problem (the 30-minute trap to avoid)
&lt;/h2&gt;

&lt;p&gt;The biggest risk in a weekend build is scope. Do not build a "general AI assistant." General assistants have no finish line — you'll spend Sunday night tuning a personality and have nothing to show.&lt;/p&gt;

&lt;p&gt;Pick a small, boring, &lt;em&gt;completable&lt;/em&gt; problem where the value is obvious. Mine: &lt;strong&gt;an internal research agent that answers questions about my own codebase and related documentation, then drafts a summary file.&lt;/strong&gt; Boring. Useful. Demos well in three minutes.&lt;/p&gt;

&lt;p&gt;The rules for a good pick:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Has a clear output artifact.&lt;/strong&gt; A report, a converted file, a categorized list. Something you can show.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Needs at least one function call.&lt;/strong&gt; The agent must act, not just chat. If there's no capability wiring, it's not an agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fails sometimes.&lt;/strong&gt; Honestly, this is a feature. A system that can make an error and recover — retry, fall back, log — is more impressive than one that never fails because it never does anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Framework choice: LangGraph, CrewAI, or AutoGen?
&lt;/h2&gt;

&lt;p&gt;The question every "how do I build an agent" post raises, and the honest answer is: &lt;strong&gt;they're all fine, pick one, and don't spend the weekend switching.&lt;/strong&gt; The reason matters for the interview, not just the build. Hiring managers rarely care which framework you used; they care that you understand the five layers, because frameworks turn over every ~18 months and layers don't.&lt;/p&gt;

&lt;p&gt;That said, there are real differences, and knowing them costs you nothing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LangGraph&lt;/strong&gt; — graph-based orchestration. Explicit nodes, edges, state. Best fit if you come from a backend background and like to see control flow on a page. The largest ecosystem (LangChain), which means the most examples and the most job postings mentioning it by name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CrewAI&lt;/strong&gt; — role-based "crews." You define agents with roles and tasks, they cooperate. Fastest to a working prototype, very readable, a common choice for demo projects. Less explicit control over the state machine when things get complex.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AutoGen&lt;/strong&gt; (Microsoft) — conversation-based multi-agent framework. Strong for research-style, multi-agent conversations and for integrating with Microsoft tooling. Its design philosophy (agents talk to each other) is a different mental model from a graph.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a weekend build, my order of preference: &lt;strong&gt;LangGraph if you want the skill to transfer to the largest number of postings, CrewAI if you want the fastest readable demo, AutoGen if you live in the Microsoft ecosystem.&lt;/strong&gt; All three let you implement the five layers. The framework is a means; the layers are the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  The weekend, compressed
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Day 1: Model access + context + a working loop
&lt;/h3&gt;

&lt;p&gt;Morning: stand up the model layer. One module that wraps the LLM API, one function that handles the raw call, a small retry with backoff, and a couple of unit tests that hit the real API. The "thin service layer" from layer 1 is genuinely thin: a function, error types, a logger.&lt;/p&gt;

&lt;p&gt;Afternoon: add context. I used embeddings + a small local vector store over my docs, with a &lt;code&gt;retrieve(query) -&amp;gt; list[chunk]&lt;/code&gt; function. Then — the moment it becomes an agent — I wrote the loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_steps&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;steps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_steps&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# "which capability, if any?"
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;capabilities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;steps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;AgentTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;did not finish in &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;max_steps&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; steps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eighteen lines. That loop &lt;em&gt;is&lt;/em&gt; the agent. Everything else is engineering.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;capabilities&lt;/strong&gt; registry is layer 3, and it's where a weekend builder should spend the most time, because it's what makes the difference between a chatbot and an agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;CAPABILITIES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retrieve_docs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;read_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# guarded to an allow-list of paths
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;write_summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;write_summary&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# writes to an output dir only
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;search_github&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;github_search&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# read-only API call
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model picks an action, the registry executes it, the result goes back into context, repeat. Guardrails (layer 4) wrap the two write-capable functions: every file write goes through one narrow function that validates the path is inside the allow-list and the content matches a schema.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 2: Evaluation, deployment, and the README that matters
&lt;/h3&gt;

&lt;p&gt;Morning: evaluation. I built a fixed dataset of 20 tasks with known-good answers, ran the agent, scored it. Result: 17/20 correct end-to-end. That one number — precise, dated, reproducible — is worth more in an interview than ten "I'm passionate about AI" lines. I also logged the two failure modes I saw (one context-truncation miss, one capability the agent ignored) and wrote a one-paragraph "known limitations" note.&lt;/p&gt;

&lt;p&gt;Afternoon: deployment. I containerized the service and exposed a single HTTP endpoint, health check included, so it's &lt;em&gt;runnable&lt;/em&gt; — not a notebook. Then the README. This is the part most builders skip and it's the part that gets you hired:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# codebase-research-agent&lt;/span&gt;

An autonomous research agent that answers questions about this repo's
docs and drafts a summary file. Built the weekend of 2026-09-05.

&lt;span class="gu"&gt;## How it works&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; 5-layer stack: model access, context (embeddings), capability wiring,
  guardrails (path allow-list + schema validation), evaluation
&lt;span class="p"&gt;-&lt;/span&gt; 18-line core loop (see agent/loop.py)
&lt;span class="p"&gt;-&lt;/span&gt; 3 capabilities: retrieve_docs, read_file (allow-listed), write_summary

&lt;span class="gu"&gt;## Run it&lt;/span&gt;
docker compose up   # then POST /agent {"task": "..."}

&lt;span class="gu"&gt;## Evaluation&lt;/span&gt;
20-task fixed dataset: 17/20 correct. Failure modes logged in ./eval/.

&lt;span class="gu"&gt;## Architecture&lt;/span&gt;
[one ASCII diagram: user -&amp;gt; loop -&amp;gt; model &lt;span class="nt"&gt;&amp;lt;-&amp;gt;&lt;/span&gt; context; model -&amp;gt; capabilities]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That README answers every question a hiring manager has: &lt;em&gt;What is it? Does it work? Can I run it? How does it think?&lt;/em&gt; Three minutes to read, and it demonstrates layer 5 (evaluation) better than any bullet point.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 60-second interview pitch
&lt;/h2&gt;

&lt;p&gt;You built the thing. Here's how to make it land when someone asks. Structure it as &lt;strong&gt;problem → constraint → mechanism → result&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"My team kept answering the same architecture questions against a growing documentation set. I wanted to know if an agent could do the retrieval &lt;em&gt;and&lt;/em&gt; the summarization reliably. Constraint: a weekend, and she must act — fetch docs, read files, write the summary herself. So I built a one-loop agent over three capabilities with a path allow-list as the guardrail. I scored it on 20 fixed tasks: 17/20 correct, and I logged the two failure modes. The service is deployed with a health check, and the README walks through the architecture."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Notice what that does. It shows you understand &lt;strong&gt;scope&lt;/strong&gt; (one problem, one weekend), &lt;strong&gt;mechanism&lt;/strong&gt; (the loop, the capabilities, the guardrail), &lt;strong&gt;verification&lt;/strong&gt; (the 17/20 number), and &lt;strong&gt;delivery&lt;/strong&gt; (deployed, documented). That's the full stack a hiring manager screens for, compressed into a minute.&lt;/p&gt;

&lt;p&gt;Then, when they ask a follow-up — and they will — you have the architecture in your head. "Why three capabilities?" "Because each one is a boundary I can guard." "What would you add?" "A second evaluation pass on the retrieval layer, and a human-in-the-loop confirmation for the write capability." Those answers only exist because you actually built it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;A weekend build does not make you a principal AI engineer, and this article should not read as "one weekend and you're done." What it does is close the gap I described in my last post: the &lt;strong&gt;30-40% of what hiring managers want&lt;/strong&gt; that a traditional checklist misses. It moves you from "knows what agents are" to "has shipped one, can defend the design, has a number." For a working developer in the 2026 market, that's the single highest-leverage weekend available.&lt;/p&gt;

&lt;p&gt;Also honest: the framework landscape moves fast, and my pick (LangGraph) might not be yours — that's fine and I say so explicitly. The layers are the durable knowledge; the framework is a costume.&lt;/p&gt;

&lt;p&gt;If you build one this weekend, two requests:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Put the evaluation number in the README.&lt;/strong&gt; One reproducible number beats ten frameworks mentioned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tell me what you built&lt;/strong&gt; — reply here with the link. The most useful part of the agentic wave is that the bar for "shipped something real" is still low, and we're all raising it together.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;I track this kind of signal every week — which AI roles are growing fastest, what employers actually ask for, and where the good remote roles live. It goes out in the Remote Signal newsletter: one concise email, no spam. If you're navigating the 2026 market, it's worth your inbox space.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>career</category>
      <category>llm</category>
    </item>
    <item>
      <title>The 2026 Tech Hiring Split: Remote Is Scarce, AI Is Hot</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Wed, 09 Sep 2026 01:12:36 +0000</pubDate>
      <link>https://dev.to/primesieve/the-2026-tech-hiring-split-remote-is-scarce-ai-is-hot-5179</link>
      <guid>https://dev.to/primesieve/the-2026-tech-hiring-split-remote-is-scarce-ai-is-hot-5179</guid>
      <description>&lt;h1&gt;
  
  
  The 2026 Tech Hiring Split: Remote Is Scarce, AI Is Hot
&lt;/h1&gt;

&lt;p&gt;Two recruiters, same week, same search interface, same candidate pool.&lt;/p&gt;

&lt;p&gt;One says: "The market is great. I can't fill AI roles fast enough."&lt;/p&gt;

&lt;p&gt;The other says: "The market is dead. Generalist front-end requisitions sit open for months and hiring managers keep getting pickier."&lt;/p&gt;

&lt;p&gt;Both are telling the truth. That's the defining feature of tech hiring in 2026 — it isn't one market anymore. It's two, running in parallel, with completely different rules. If you're hunting for a remote role right now, knowing which market you're in is the single highest-leverage piece of information you can have.&lt;/p&gt;

&lt;p&gt;Here's the data that proves the split, what the remaining fully-remote jobs actually look like, and a three-move playbook to position yourself on the right side of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bifurcation, in numbers
&lt;/h2&gt;

&lt;p&gt;The most useful framing I've seen for the 2026 market comes from the Indeed Hiring Lab. They index tech job listings against their February 2020 baseline, which is set at 100. Here's where things stood in mid-2025:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Machine learning engineer listings: index 159&lt;/strong&gt; (+59% vs. baseline)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All tech listings: index 64&lt;/strong&gt; (-36%)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;General software engineer listings: index 51&lt;/strong&gt; (-49%)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read those numbers slowly. General software engineering is at half its 2020 level. Machine learning engineering is at 160% of it. That's a 108-point gap between the two — the market isn't contracting uniformly, it's rotating. The jobs being cut are not the jobs being created.&lt;/p&gt;

&lt;p&gt;Other sources agree. LinkedIn's Jobs on the Rise 2026 list ranks &lt;strong&gt;AI engineer as the #1 fastest-growing job in the United States for the second consecutive year&lt;/strong&gt;. Stanford's 2026 AI Index found agentic AI job listings grew &lt;strong&gt;10,854% year over year&lt;/strong&gt; — not a typo, ten thousand percent — while entry-level software developer employment fell roughly &lt;strong&gt;20% from its 2024 peak&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The entry-level collapse is the part most summaries skip, and it's the part that matters most for early-career devs. The floors of the pyramid are shrinking while the peak grows. If you graduated recently or you're self-taught, you're competing in a market where the traditional on-ramp roles are the ones disappearing.&lt;/p&gt;

&lt;p&gt;CompTIA's March 2026 data shows total tech openings finally turned positive again — 537,000+ active openings, up 9.7% month over month and 8.9% versus the prior year. But that growth is lopsided: CompTIA counts 275,000+ active US job postings referencing AI skills in January 2026, a &lt;strong&gt;153% jump from January 2024&lt;/strong&gt;, and data analyst and scientist roles are growing at 420% of the national rate while traditional infrastructure and helpdesk work sits at the bottom.&lt;/p&gt;

&lt;p&gt;Volume is rising. But it's rising into a market where demand has shifted hard toward AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The remote premium: 8% is the new scarce tier
&lt;/h2&gt;

&lt;p&gt;Here's the other number that reframes everything you've heard about remote work. Robert Half's 2026 data on US tech job ads:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Work arrangement&lt;/th&gt;
&lt;th&gt;Share of tech ads&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fully on-site&lt;/td&gt;
&lt;td&gt;74%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid&lt;/td&gt;
&lt;td&gt;18%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fully remote&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Tech still offers more fully-remote work than any other US sector — every other industry is below that — but the default for new tech hires is now in-person. Remote is no longer the baseline feature it was in 2021; it's the scarce tier.&lt;/p&gt;

&lt;p&gt;And scarcity creates a flood. Recruiting sources consistently report that fully-remote openings receive &lt;strong&gt;6–10× the applicant volume of comparable on-site listings&lt;/strong&gt;. An 8% share of listings pulling 6-10× the applications means the remote hunt is not soft competition for the leftovers — it's the most contested slice of the whole market.&lt;/p&gt;

&lt;p&gt;That changes your strategy in a concrete way. If you're applying to a fully-remote general software engineering role, you are statistically in the most applicant-dense corner of the entire tech job market: the shrinking category (general engineering) meeting the scarce category (fully remote). Which is exactly why the next section matters more than any resume trick.&lt;/p&gt;

&lt;h2&gt;
  
  
  What hiring managers actually screen for now
&lt;/h2&gt;

&lt;p&gt;The old checklist — "Python, plus a framework, plus some projects" — is quietly missing the point. One 2026 analysis from the recruiting side put it this way: recruiters who screen on the traditional checklist are missing 30-40% of what hiring managers now want to see in the first conversation.&lt;/p&gt;

&lt;p&gt;The new working definition, as one labor-market analyst phrased it, is: "Python plus a framework plus enough applied ML to ship a model into production plus enough prompt engineering to wire it into an agentic workflow."&lt;/p&gt;

&lt;p&gt;Notice what's in there:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Shipping, not studying.&lt;/strong&gt; "Enough applied ML to ship a model into production" is a deployment statement, not a theory statement. Hiring managers want evidence you've taken something from notebook to live system — observability, errors, latency, the unglamorous parts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic workflow wiring.&lt;/strong&gt; The 10,854% spike in agentic AI listings isn't job-post inflation, it's real: companies are wiring LLM calls into multi-step workflows, and they need people who can do the plumbing — memory, context management, retries, guardrails, evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breadth over depth.&lt;/strong&gt; The 30-40% gap comes from candidates who have deep theory or deep CRUD experience but can't cross the boundary between the two.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The salary signal reinforces the shift. Robert Half's 2026 Salary Guide pegs national software engineer ranges around $109K-$175K while AI/ML engineering runs $134K-$193K — roughly a $25K floor premium. At senior level the gap widens: AI staff engineers earn about 18.7% more than their non-AI peers per Levels.fyi data. That premium is the market pricing the scarcity you can see in the listing index numbers.&lt;/p&gt;

&lt;h3&gt;
  
  
  The agentic stack, concretely
&lt;/h3&gt;

&lt;p&gt;"Agentic workflow" sounds like a buzzword until you list the actual layers a hiring manager means. The demand spike is for people who can wire these five pieces together:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model access&lt;/strong&gt; — an API (or a self-hosted model) behind a thin service layer, with cost and rate limits understood&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context management&lt;/strong&gt; — a retrieval layer (embeddings + a vector store, or a well-structured prompt cache) because raw context windows don't scale to real products&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability wiring&lt;/strong&gt; — letting the model invoke your own backend functions: search, CRUD, external APIs. This is where most "AI features" actually live&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails&lt;/strong&gt; — validation of model output before it touches a database or a customer: schemas, rejection rules, retry policies&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation&lt;/strong&gt; — a fixed scored dataset so you can say "v2 answers correctly 92% of the time vs. 84% for v1" instead of "it feels better"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's the encouraging part: four of those five layers are ordinary backend engineering wearing a new costume. If you already build APIs, you know 80% of the mechanics — the job posting just needs to see them applied to the model layer. That is exactly where the 30-40% screening gap lives and exactly what Move 1 is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three-move playbook
&lt;/h2&gt;

&lt;p&gt;If the data is the map, here's the route. None of this requires a return to grad school or a five-year plan. All three moves are available this quarter.&lt;/p&gt;

&lt;h3&gt;
  
  
  Move 1: Add a shipping layer to your stack, not a new language
&lt;/h3&gt;

&lt;p&gt;The market isn't demanding that every engineer become a research scientist. It's demanding that engineers who touch AI show deployment hygiene. Pick the framework you already use, and add one production artifact you can prove:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A model or agent service deployed with a real endpoint, error handling, and a health check&lt;/li&gt;
&lt;li&gt;A retrieval or memory layer with observable trace logs&lt;/li&gt;
&lt;li&gt;An evaluation harness that scores your system's outputs on a fixed dataset&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One deployed system with evidence beats five tutorials on a resume, because it matches what hiring managers are actually screening for. The "30-40% gap" is closing distance on exactly this axis.&lt;/p&gt;

&lt;h3&gt;
  
  
  Move 2: Choose your applicant pool deliberately
&lt;/h3&gt;

&lt;p&gt;You can't control the 6-10× applicant flood on fully-remote generalist roles, so don't line up in that crowd. Two levers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Skill-level competition:&lt;/strong&gt; hybrid roles get a fraction of the applications remote does. If you can commute to a tech hub even two days a week, the same role drops you into a much smaller pool — and hybrid listings are 18% of the market, more than double the remote share.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Title-level competition:&lt;/strong&gt; "Data analyst," "ML engineer," and "AI engineer" titles are growing 420%, 159%, and at record rates respectively. The same Python skills applied to analytics or ML engineering land in a growing category instead of a shrinking one. Same fundamentals, different shelf.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The takeaway isn't "abandon remote" — it's that remote is the premium you pay for with either seniority or specialization. Junior generalist + fully remote is the worst combination in the 2026 market. Pick one axis to upgrade.&lt;/p&gt;

&lt;h3&gt;
  
  
  Move 3: Track the signal, not the narrative
&lt;/h3&gt;

&lt;p&gt;Market narratives are slow and wrong. Job-post data is fast and honest. A few things worth watching weekly instead of reading headlines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Posting volume by role:&lt;/strong&gt; when your target title's posting index rises, application decisions favor you; when it falls, competition tightens even without layoff news&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The remote share:&lt;/strong&gt; if fully-remote share drops from 8% to 6%, the flood on the remaining 6% grows proportionally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The AI-skill share of postings:&lt;/strong&gt; 275K+ postings referencing AI skills in January 2026 is a floor that's been rising for two years — it tells you whether the "add an AI layer" signal is still compounding&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Salary dispersion at your level:&lt;/strong&gt; Levels.fyi-style data shows where the premium is real versus where job titles just renamed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You don't need special access to any of this. CompTIA publishes its Tech Jobs Report monthly, Indeed and Stanford publish their analyses openly, and the sources update on schedule. Reading the data weekly takes twenty minutes and gives you an edge that most applicants — who react to headlines months late — simply don't have.&lt;/p&gt;

&lt;p&gt;Here's what one of those twenty-minute reads actually looks like. CompTIA's monthly report lands: your target title's postings are up, but the fully-remote share ticked down another point. Indeed's monthly index refreshes: generalist backend is still underwater, but ML engineer demand crossed 160. Levels.fyi shows the AI premium holding at senior while mid-senior generalist pay is flat for the fourth straight quarter. One actionable conclusion falls out: the niche you've been orbiting is getting more remote-scarce and more AI-flavored at the same time — so the hybrid role you dismissed last month deserves a second look. That's not a headline you'd have read anywhere. It's a decision you made from data in twenty minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;This is a US-heavy picture. If you're outside the US, the direction of travel — AI skills up, general engineering flat, remote scarce and competitive — appears consistently across every global labor-market source, but the magnitudes differ by region. Check your own market's data before acting on the specifics.&lt;/p&gt;

&lt;p&gt;Also: no stat in this article should read as "abandon general engineering." Markets rotate, and the rotation that emptied general engineering in 2024-2026 is the same one that emptied the COBOL pool in the 90s and the sysadmin pool in the 2010s — the work didn't vanish, it moved up the stack. The engineers who win these rotations are the ones who notice the move while it's still cheap, not after the peak.&lt;/p&gt;

&lt;p&gt;The two markets will keep running in parallel for a while. Pick which one you're competing in — deliberately, with data, this quarter.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I track this kind of hiring signal every week and send the most useful patterns — new remote roles, demand shifts, and what employers actually ask for — in the Remote Signal newsletter. No spam, one concise email a week. If you're navigating the 2026 market, it's worth your inbox space.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>career</category>
      <category>ai</category>
      <category>remote</category>
      <category>python</category>
    </item>
    <item>
      <title>Remote Tech Hiring Publish Test</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Wed, 09 Sep 2026 00:56:02 +0000</pubDate>
      <link>https://dev.to/primesieve/remote-tech-hiring-publish-test-37h7</link>
      <guid>https://dev.to/primesieve/remote-tech-hiring-publish-test-37h7</guid>
      <description>&lt;h1&gt;
  
  
  Test
&lt;/h1&gt;

&lt;p&gt;body&lt;/p&gt;

</description>
      <category>career</category>
      <category>ai</category>
    </item>
    <item>
      <title>How I built a Hacker News scraper that pulled 1,000 posts in 10 seconds</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Fri, 04 Sep 2026 13:36:49 +0000</pubDate>
      <link>https://dev.to/primesieve/how-i-built-a-hacker-news-scraper-that-pulled-1000-posts-in-10-seconds-gi3</link>
      <guid>https://dev.to/primesieve/how-i-built-a-hacker-news-scraper-that-pulled-1000-posts-in-10-seconds-gi3</guid>
      <description>&lt;h1&gt;
  
  
  How I built a Hacker News scraper that pulled 1,000 posts in 10 seconds
&lt;/h1&gt;

&lt;p&gt;Most HN scrapers I've seen over-engineer the problem. Proxy rotation, headless browsers, captcha solving — all to pull data that's already exposed through a free JSON API. Here's the version that takes 10 seconds and costs nothing&lt;/p&gt;

&lt;h2&gt;
  
  
  The API nobody charges for
&lt;/h2&gt;

&lt;p&gt;Hacker News uses Algolia to power search, and Algolia exposes a public endpoint with no key required&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://hn.algolia.com/api/v1/search_by_date?tags=story&amp;amp;hitsPerPage=200&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That's it. No auth. No rate limit. Returns clean JSON with title, URL, points, comments, author, timestamp — everything you need to analyze HN&lt;/p&gt;

&lt;h2&gt;
  
  
  The fetch (no Apify, no SDK, just urllib
&lt;/h2&gt;

&lt;p&gt;`import json&lt;br&gt;
import urllib.request&lt;br&gt;
import time&lt;/p&gt;

&lt;p&gt;stories = []&lt;br&gt;
for page in range(10):&lt;br&gt;
    url = f"&lt;a href="https://hn.algolia.com/api/v1/search_by_date?tags=story&amp;amp;hitsPerPage=100&amp;amp;page=%7Bpage%7D" rel="noopener noreferrer"&gt;https://hn.algolia.com/api/v1/search_by_date?tags=story&amp;amp;hitsPerPage=100&amp;amp;page={page}&lt;/a&gt;"&lt;br&gt;
    req = urllib.request.Request(url, headers={"User-Agent": "hn-research/1.0"})&lt;br&gt;
    with urllib.request.urlopen(req) as r:&lt;br&gt;
        data = json.loads(r.read())&lt;br&gt;
        stories.extend(data.get("hits", []))&lt;br&gt;
    time.sleep(0.3)  # be polite&lt;/p&gt;

&lt;p&gt;print(f"Pulled {len(stories)} stories")`&lt;/p&gt;

&lt;p&gt;Total runtime on my VPS: about 12 seconds for 1,000 posts. The 0.3s sleep is optional but polite — Algolia serves everyone from the same backend&lt;/p&gt;

&lt;h2&gt;
  
  
  What the data actually says
&lt;/h2&gt;

&lt;p&gt;I ran the script for the last 7 days and bucketed by points&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Median points: 1&lt;/strong&gt;. Half of all HN stories get exactly one upvote&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Only 1.5% reach 50 points&lt;/strong&gt; — that's "visible" by HN's standards&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Zero posts crossed 500 points&lt;/strong&gt; in the entire 7-day window&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The top story of the week landed at 104 points (a Jane Street reverse engineering writeup). The HN virality curve is steep. If you're posting there hoping for traction, you're competing in a 1.5% funnel&lt;/p&gt;

&lt;h2&gt;
  
  
  When the data was posted (UTC
&lt;/h2&gt;

&lt;p&gt;06:00 UTC    5 posts&lt;br&gt;
07:00 UTC   35 posts&lt;br&gt;
08:00 UTC   28 posts&lt;br&gt;
09:00 UTC   27 posts&lt;br&gt;
10:00 UTC   40 posts&lt;br&gt;
11:00 UTC   46 posts&lt;br&gt;
12:00 UTC   19 posts&lt;br&gt;
(other hours: ~0)&lt;/p&gt;

&lt;p&gt;Almost everything posts between 07:00 and 12:00 UTC (US morning). Counterintuitively, the &lt;em&gt;best&lt;/em&gt; window to post is 06:00–07:00 UTC — you catch the morning traffic spike with less competition&lt;/p&gt;

&lt;h2&gt;
  
  
  Domain analysis
&lt;/h2&gt;

&lt;p&gt;Top linked domains across 1,000 posts&lt;/p&gt;

&lt;p&gt;twitter.com      45&lt;br&gt;
github.com       29&lt;br&gt;
openai.com       18&lt;br&gt;
nytimes.com      13&lt;br&gt;
anthropic.com    13&lt;/p&gt;

&lt;p&gt;AI-related domains are over-represented. HN readers actively look for AI news, so if you're posting in that space, you get structural tailwind&lt;/p&gt;

&lt;h2&gt;
  
  
  Packaged as an Apify actor
&lt;/h2&gt;

&lt;p&gt;I wrapped the script as an Apify actor so I can re-run it weekly without rewriting the boilerplate. The actor charges per result (PAY_PER_EVENT), so cost scales with what you actually pull&lt;/p&gt;

&lt;p&gt;`// src/main.js — Apify actor (free Algolia endpoint)&lt;br&gt;
import { Actor } from 'apify';&lt;/p&gt;

&lt;p&gt;await Actor.main(async () =&amp;gt; {&lt;br&gt;
  const input = await Actor.getInput() || {};&lt;br&gt;
  const queries = input.queries || [];&lt;br&gt;
  const tags = input.tags || 'story';&lt;br&gt;
  const maxResults = Math.min(500, input.maxResults || 100);&lt;/p&gt;

&lt;p&gt;for (const q of queries) {&lt;br&gt;
    const url = &lt;code&gt;https://hn.algolia.com/api/v1/search?query=${encodeURIComponent(q)}&amp;amp;tags=${tags}&amp;amp;hitsPerPage=${Math.min(100, maxResults)}&lt;/code&gt;;&lt;br&gt;
    const res = await fetch(url, { headers: { accept: 'application/json' } });&lt;br&gt;
    const data = await res.json();&lt;br&gt;
    for (const hit of data.hits || []) {&lt;br&gt;
      await Actor.pushData({&lt;br&gt;
        title: hit.title || hit.story_title,&lt;br&gt;
        url: hit.url || hit.story_url,&lt;br&gt;
        author: hit.author,&lt;br&gt;
        points: hit.points,&lt;br&gt;
        comments: hit.num_comments,&lt;br&gt;
        createdAt: hit.created_at,&lt;br&gt;
      });&lt;br&gt;
      try { await Actor.charge({ eventName: 'result', count: 1 }); } catch {}&lt;br&gt;
    }&lt;br&gt;
  }&lt;br&gt;
});`&lt;/p&gt;

&lt;h2&gt;
  
  
  The full pipeline
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pull&lt;/strong&gt; — fetch 1,000 fresh stories (10 seconds&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Score&lt;/strong&gt; — bucket by points, comments, posting hour (Python pandas, 30 seconds&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Analyze&lt;/strong&gt; — extract the patterns above (5 minutes of thinking&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Write up&lt;/strong&gt; — turn the data into a Medium article (the long-form version&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Total time: about an hour from API call to published article. Total cost: free (Algolia is free, the actor charges per-result on Apify&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;You don't need a scraping service to do competitive research on public forums. HN, Reddit, GitHub, Wikipedia, Stack Overflow — all of them have free public APIs or fetch-clean endpoints. The bottleneck isn't data access, it's analysis&lt;/p&gt;

&lt;p&gt;The same approach works for any forum where posts are timestamped. The trick is: pick a 7-day window, pull everything, score by engagement, and look for the patterns. You don't need ML — a 50-line script does it&lt;/p&gt;

&lt;p&gt;This is part of a broader weekly data routine I run — HN virality, GitHub trends, and (separately) the remote job market. If you like this style of low-cost data analysis, I publish a weekly report on remote hiring signals every Wednesday&lt;/p&gt;

</description>
      <category>python</category>
      <category>data</category>
      <category>scraping</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Remote Engineering Hiring Signal: What 250+ Scraped Jobs Reveal This Week</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Wed, 02 Sep 2026 14:54:51 +0000</pubDate>
      <link>https://dev.to/primesieve/remote-engineering-hiring-signal-what-250-scraped-jobs-reveal-this-week-44hm</link>
      <guid>https://dev.to/primesieve/remote-engineering-hiring-signal-what-250-scraped-jobs-reveal-this-week-44hm</guid>
      <description>&lt;h1&gt;
  
  
  Remote Engineering Hiring Signal: What 250+ Scraped Jobs Reveal This Week
&lt;/h1&gt;

&lt;p&gt;Every week, our automated scrapers index over 250 remote-first engineering job boards, filtering out ghost listings, low-effort recruiter aggregators, and stale postings to find out who is actually hiring senior developers.&lt;/p&gt;

&lt;p&gt;Here is a quick snapshot of this week's data, featured roles, and engineering trends.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 This Week's Market Signal
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Total Jobs Scanned:&lt;/strong&gt; 256&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High-Signal Listings (Grade A/B):&lt;/strong&gt; 217&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Top In-Demand Stacks:&lt;/strong&gt; Python, PostgreSQL, React, Node.js, Go, FastAPI, Docker&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Query Breakdown
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mobile developer remote:&lt;/strong&gt; 55 roles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend developer remote:&lt;/strong&gt; 52 roles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DevOps engineer remote:&lt;/strong&gt; 52 roles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI / ML engineer remote:&lt;/strong&gt; 50 roles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full stack engineer remote:&lt;/strong&gt; 47 roles&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🏆 Featured Roles (Sample from This Week)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Senior Backend Engineer&lt;/strong&gt; — Nametag (Seattle / Remote)

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Tech:&lt;/em&gt; Python, PostgreSQL, REST&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Signal:&lt;/em&gt; Identity verification platform scaling backend infrastructure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full Stack Developer&lt;/strong&gt; — Softermii (Los Angeles / Remote)

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Tech:&lt;/em&gt; React, Node.js, PostgreSQL&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Signal:&lt;/em&gt; Fintech solutions expanding remote engineering team.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Senior GoLang Developer&lt;/strong&gt; — ImagineX (Atlanta / Remote)

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Tech:&lt;/em&gt; Go, PostgreSQL, Docker&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Signal:&lt;/em&gt; High-performance backend microservices.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  🛠️ Tool of the Week: Donsetch
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Donsetch&lt;/strong&gt; — Keyless web fetch/search/crawl for AI agents. Auto bot-wall bypass, clean markdown output. Perfect for building data pipelines without proxy costs. &lt;a href="https://github.com/anthropics/donsetch" rel="noopener noreferrer"&gt;GitHub →&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Get the Full 15-Role Briefing Every Wednesday
&lt;/h3&gt;

&lt;p&gt;This is just a small sample of our weekly briefing. We deliver 5 deep-dive featured roles with direct apply links, plus 10 quick hits and a cold pitch template straight to your inbox every week.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;&lt;a href="https://remotesignal.substack.com" rel="noopener noreferrer"&gt;Subscribe to Remote Signal&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>career</category>
      <category>remote</category>
      <category>hiring</category>
      <category>webdev</category>
    </item>
    <item>
      <title>I Built a Wikipedia Scraper That Runs on Apify Without Proxies</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Fri, 28 Aug 2026 10:18:11 +0000</pubDate>
      <link>https://dev.to/primesieve/i-built-a-wikipedia-scraper-that-runs-on-apify-without-proxies-525l</link>
      <guid>https://dev.to/primesieve/i-built-a-wikipedia-scraper-that-runs-on-apify-without-proxies-525l</guid>
      <description>&lt;h1&gt;
  
  
  I Built a Wikipedia Scraper That Runs on Apify Without Proxies
&lt;/h1&gt;

&lt;p&gt;Most Wikipedia scrapers on Apify use Playwright or headless browsers. I built one that hits the MediaWiki API directly — 200 responses, zero proxy costs, 54 lines of code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Wikipedia Data?
&lt;/h2&gt;

&lt;p&gt;Wikipedia is the largest open knowledge base on the internet. 60M+ articles across 300+ languages. Researchers, NLP teams, knowledge graph builders, and content aggregators all need structured access to it.&lt;/p&gt;

&lt;p&gt;The problem: most tools overcomplicate it. They spin up headless browsers to render pages, manage proxy pools to avoid rate limits, and charge you for the overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Actor Does
&lt;/h2&gt;

&lt;p&gt;Search any keyword across any Wikipedia language. Get back structured results: title, page ID, word count, snippet, timestamp, and direct URL. Paginate up to 500 results per keyword with automatic offset handling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;keywords&lt;/code&gt; — array of search terms&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;lang&lt;/code&gt; — language code (default: &lt;code&gt;en&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;maxResults&lt;/code&gt; — up to 500 per keyword&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Output per result:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"keyword"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"quantum computing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"lang"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"en"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Quantum computing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"pageid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;22967&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"size"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;85432&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"wordcount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;12400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"snippet"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Quantum computing is a type of computation..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2024-12-15T10:30:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://en.wikipedia.org/wiki/Quantum_computing"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The No-Proxy Advantage
&lt;/h2&gt;

&lt;p&gt;Wikipedia's MediaWiki API is open. No authentication required. No rate limit tricks needed if you respect their guidelines.&lt;/p&gt;

&lt;p&gt;This means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero proxy costs&lt;/strong&gt; — runs clean on Apify's AWS infrastructure&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No browser overhead&lt;/strong&gt; — plain HTTP GET, ~250ms per request&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliable&lt;/strong&gt; — no CAPTCHA challenges, no IP blocks, no flaky browser sessions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most competing actors charge $3-5 per 1k results and burn proxy credits. This one runs at the Apify free tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use Cases
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NLP training data&lt;/strong&gt; — bulk collect article metadata for text classification&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge graph construction&lt;/strong&gt; — page IDs + titles + timestamps for entity linking&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content gap analysis&lt;/strong&gt; — compare coverage across languages&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Research automation&lt;/strong&gt; — systematic literature discovery on any topic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SEO research&lt;/strong&gt; — find what Wikipedia covers about your niche&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Running It
&lt;/h2&gt;

&lt;p&gt;Search for "Wikipedia Search Scraper" on Apify Store, or use the API directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.apify.com/v2/acts/primesievecoder~wikipedia-search-scraper/runs"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"keywords": ["quantum computing", "machine learning"], "lang": "en", "maxResults": 100}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Adding full-text extract mode (prop=extracts) for teams that need article body content, not just search metadata. Stay tuned.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by &lt;a href="https://apify.com/primesievecoder" rel="noopener noreferrer"&gt;Prime Sieve&lt;/a&gt; — scraping tools that just work. No proxies, no headless browsers, no drama.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>scraping</category>
      <category>apify</category>
      <category>wikipedia</category>
      <category>buildinpublic</category>
    </item>
    <item>
      <title>Reverse-Engineering YouTube's InnerTube Search API (WEB Client + Continuations)</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Mon, 24 Aug 2026 12:07:08 +0000</pubDate>
      <link>https://dev.to/primesieve/reverse-engineering-youtubes-innertube-search-api-web-client-continuations-480m</link>
      <guid>https://dev.to/primesieve/reverse-engineering-youtubes-innertube-search-api-web-client-continuations-480m</guid>
      <description>&lt;h1&gt;
  
  
  Reverse-Engineering YouTube's InnerTube Search API (WEB Client + Continuations)
&lt;/h1&gt;

&lt;p&gt;YouTube search in a browser loads HTML. YouTube search as a client sends JSON.&lt;/p&gt;

&lt;p&gt;Same query, two paths. The HTML path costs a browser, a consent click, and 2 GB of RAM. The JSON path costs one &lt;code&gt;fetch&lt;/code&gt;. I picked the second.&lt;/p&gt;

&lt;p&gt;This is the shape of YouTube's InnerTube search from the WEB client side — what it sends, what it returns, how it pages — and where I stopped so this stays a tour, not a recipe.&lt;/p&gt;

&lt;p&gt;The actor is live: &lt;strong&gt;primesieve/youtube-search-scraper&lt;/strong&gt; — keywords in, normalized videos out, 20 per page, fetch-clean.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the browser actually does
&lt;/h2&gt;

&lt;p&gt;Open YouTube, type &lt;code&gt;golang tutorial&lt;/code&gt;, watch network. One POST fires. JSON goes out with &lt;code&gt;query&lt;/code&gt; and a &lt;code&gt;context.client&lt;/code&gt; block. JSON comes back with a tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;contents
└─ twoColumnSearchResultsRenderer
   └─ primaryContents
      └─ sectionListRenderer
         └─ contents[]
            ├─ itemSectionRenderer
            │  └─ contents[]  → videoRenderer
            └─ continuationItemRenderer → token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus a second bucket for paginated results:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;onResponseReceivedCommands[]
└─ appendContinuationItemsAction
   └─ continuationItems[]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The leaves are &lt;code&gt;videoRenderer&lt;/code&gt; objects — title runs, owner text, view count text, length text, publish text, thumbnail array, videoId. Everything you need is in the leaf. The rest is shelf furniture.&lt;/p&gt;

&lt;p&gt;I will not dump the request body here. The client name, version string, and header shape rotate. Copy-pasting them is how you build a fork that breaks next Tuesday. The stable contract is the shape of the tree and the fact that it pages by token.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why WEB client
&lt;/h2&gt;

&lt;p&gt;InnerTube has many clients — WEB, MWEB, ANDROID, IOS, TV and more. Each negotiates a slightly different shape and policy. I picked &lt;code&gt;WEB&lt;/code&gt; with &lt;code&gt;DESKTOP&lt;/code&gt; platform because:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It answers 200 without a login or consent wall.&lt;/li&gt;
&lt;li&gt;It returns 20 videos per page — the same count the HTML shows.&lt;/li&gt;
&lt;li&gt;It returns &lt;code&gt;visitorData&lt;/code&gt; you can grab from the homepage in one fetch and reuse. Optional, but responses are more stable with it.&lt;/li&gt;
&lt;li&gt;It pages by a single opaque continuation string you send back in the next POST. No cursor math.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Other clients change the shelf mix or require different tokens. WEB is the boring one that answers politely. Boring wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continuations, not pages
&lt;/h2&gt;

&lt;p&gt;There is no &lt;code&gt;?page=2&lt;/code&gt;. There is a token at the bottom of page one. Send it back as &lt;code&gt;continuation&lt;/code&gt; in the next body and you get page two. Repeat.&lt;/p&gt;

&lt;p&gt;The walk looks like this at a high level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// pseudocode — shape only, not the real selector or token pick&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;maxPages&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;maxResults&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;json&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;postSearch&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;continuation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;videos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;walkTree&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;          &lt;span class="c1"&gt;// visit sectionListRenderer → itemSectionRenderer → videoRenderer&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;pickContinuation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;    &lt;span class="c1"&gt;// prefer the token near the search apiUrl&lt;/span&gt;
  &lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(...&lt;/span&gt;&lt;span class="nx"&gt;videos&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;videos&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;350&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Shorts shelves sit in the same parent list as regular shelves. A parser that assumes &lt;code&gt;contents[0]&lt;/code&gt; is always videos misses them. Walk the full list, check each section's type, visit the &lt;code&gt;itemSectionRenderer.contents&lt;/code&gt; bag if present, recurse into &lt;code&gt;sectionListRenderer.contents&lt;/code&gt; when nested. That is the whole traversal.&lt;/p&gt;

&lt;p&gt;Continuation tokens are long opaque strings. Don't decode them. Don't trim them. Forward them verbatim.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I left out on purpose
&lt;/h2&gt;

&lt;p&gt;No &lt;code&gt;INNERTUBE_API_KEY&lt;/code&gt;, no client version string, no visitorData regex, no header map, no exact JSON path for the token. Those change. I have a notes file of failure modes — one line each: what broke, how the site changed, what I would do differently. Most are boring. A few become README warnings. Publishing the exact payload just makes a copy that rots and a GitHub issue I have to answer with "yeah, that changed Tuesday."&lt;/p&gt;

&lt;p&gt;What matters to you is the interface — yours, not mine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"keywords"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"lofi hip hop"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"maxResults"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"maxPages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"keyword"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"lofi hip hop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"videoId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5qap5aO4i9A"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"channel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"views"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"duration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"published"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"thumbnail"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://i.ytimg.com/vi/5qap5aO4i9A/hqdefault.jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://www.youtube.com/watch?v=5qap5aO4i9A"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same five fields in, same eight fields out, keyword after keyword. My parse ladder is my maintenance burden.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fetch-clean is the check that mattered
&lt;/h2&gt;

&lt;p&gt;Before writing anything I ran one fetch with a desktop user-agent against the homepage. 200. Then one POST against the search endpoint. 200. No proxy. That two-line check killed the browser branch before it started.&lt;/p&gt;

&lt;p&gt;The bar for fetch-clean is low and the savings are high: no Playwright, no Puppeteer, no 2 GB browser, no consent dialog to click, no scroll loop to babysit. Apify SDK for input/dataset/pay-per-event, native fetch for HTTP, Node 20, 512 MB, 600s. One file in &lt;code&gt;src/main.js&lt;/code&gt;. The boring stack that never wakes you at 2am.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three small lessons from shipping it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. visitorData is optional but cheap.&lt;/strong&gt; Fetch the homepage once, grab the token if the HTML has one, fall back to a default if not. If the homepage fetch flakes, don't fail the run. Small stability win, not a dependency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Never assume flat contents.&lt;/strong&gt; The tree mixes video sections and shorts sections and continuation sentinels in one array. Recurse. Check the key that is present — &lt;code&gt;itemSectionRenderer&lt;/code&gt;, &lt;code&gt;continuationItemRenderer&lt;/code&gt;, &lt;code&gt;sectionListRenderer&lt;/code&gt; — and handle each. Flat-map misses shelves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Empty + no next means stop.&lt;/strong&gt; Zero videos and no continuation is end-of-results, not an error. Zero videos with a continuation is a layout change — log the key set and stop that keyword instead of looping forever. A clear stop beats a silent spin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The pricing is boring on purpose: &lt;strong&gt;$0.80 per 1,000 videos&lt;/strong&gt;, pay-per-event, no tiers.&lt;/p&gt;

&lt;p&gt;Browser-based YouTube actors cluster around $0.50–$5 per 1k because they pay the browser tax. This one doesn't. 1k = $0.80, 10k = $8. Quote it without a table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://apify.com/primesieve/youtube-search-scraper" rel="noopener noreferrer"&gt;https://apify.com/primesieve/youtube-search-scraper&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Start with one keyword and &lt;code&gt;maxResults: 50&lt;/code&gt;. Push to 500 when you trust the shape. Same schema. Same price. Boring on purpose.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm Prime Sieve — I build small tools that do one thing honestly. I write about what broke, not just what shipped. More at &lt;a href="https://github.com/primesievecoder" rel="noopener noreferrer"&gt;apify.com/Prime-Sieve&lt;/a&gt; and &lt;a href="https://github.com/primesievecoder" rel="noopener noreferrer"&gt;github.com/primesievecoder&lt;/a&gt;. Thanks for trying it. If it breaks, tell me. It will break.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>youtube</category>
      <category>javascript</category>
      <category>apify</category>
    </item>
    <item>
      <title>I built an OLX scraper for 24 countries — the boring version that actually ships</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Sat, 22 Aug 2026 09:49:04 +0000</pubDate>
      <link>https://dev.to/primesieve/i-built-an-olx-scraper-for-24-countries-the-boring-version-that-actually-ships-432p</link>
      <guid>https://dev.to/primesieve/i-built-an-olx-scraper-for-24-countries-the-boring-version-that-actually-ships-432p</guid>
      <description>&lt;h1&gt;
  
  
  I built an OLX scraper for 24 countries — the boring version that actually ships
&lt;/h1&gt;

&lt;p&gt;OLX runs classifieds in about two dozen countries. Same brand, different domains, different anti-bot setups. Everyone scraping it does one country at a time.&lt;/p&gt;

&lt;p&gt;I got tired of forking.&lt;/p&gt;

&lt;p&gt;So I put 24 countries behind one input. &lt;code&gt;country: "id"&lt;/code&gt; or &lt;code&gt;country: "pl"&lt;/code&gt; or &lt;code&gt;country: "br"&lt;/code&gt; — same schema out. It's live on Apify as &lt;code&gt;primesieve/olx-global-scraper&lt;/code&gt;. One file. No browser. Here is the boring part that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Input:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"country"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"keywords"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"iphone 13"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"maxResults"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"maxPages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"proxyConfiguration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"useApifyProxy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"apifyProxyGroups"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"RESIDENTIAL"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;country&lt;/code&gt; — two-letter code (&lt;code&gt;id&lt;/code&gt;, &lt;code&gt;pl&lt;/code&gt;, &lt;code&gt;in&lt;/code&gt;, &lt;code&gt;br&lt;/code&gt;, &lt;code&gt;ua&lt;/code&gt;, &lt;code&gt;pt&lt;/code&gt;, &lt;code&gt;ro&lt;/code&gt;, &lt;code&gt;bg&lt;/code&gt;, &lt;code&gt;kz&lt;/code&gt;, &lt;code&gt;uz&lt;/code&gt;, &lt;code&gt;pk&lt;/code&gt;, &lt;code&gt;za&lt;/code&gt;, &lt;code&gt;ng&lt;/code&gt;, &lt;code&gt;ke&lt;/code&gt;, &lt;code&gt;eg&lt;/code&gt;, &lt;code&gt;lb&lt;/code&gt;, &lt;code&gt;ph&lt;/code&gt;, &lt;code&gt;co&lt;/code&gt;, &lt;code&gt;ar&lt;/code&gt;, &lt;code&gt;pe&lt;/code&gt;, &lt;code&gt;ec&lt;/code&gt;, &lt;code&gt;gt&lt;/code&gt;, &lt;code&gt;az&lt;/code&gt;, &lt;code&gt;ma&lt;/code&gt;). Default &lt;code&gt;id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;keywords&lt;/code&gt; — one or more search terms. Each runs sequentially.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;maxResults&lt;/code&gt; / &lt;code&gt;maxPages&lt;/code&gt; — caps. Defaults 50 / 3, max 1000 / 30.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;proxyConfiguration&lt;/code&gt; — optional for Indonesia, required for the other 23.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Output — same shape every country:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"listingId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"123456789"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"iPhone 13 128GB mulus"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"price"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;6500000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"priceText"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Rp 6.500.000"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"currency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"IDR"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"city"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jakarta Selatan"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"location"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Tebet, Jakarta Selatan, DKI Jakarta"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"images"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"https://...jpg"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"thumbnailUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://...jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"listingUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://www.olx.co.id/item/123456789"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"country"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"id"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Title, price (numeric plus display text), currency, location, images, URL. No seller PII beyond what the listing page shows. No tricks.&lt;/p&gt;

&lt;p&gt;Try: &lt;a href="https://apify.com/primesieve/olx-global-scraper" rel="noopener noreferrer"&gt;https://apify.com/primesieve/olx-global-scraper&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring stack
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// no playwright, no puppeteer&lt;/span&gt;
&lt;span class="c1"&gt;// apify + fetch + cheerio. That's it.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scraper is one file. Apify SDK for input, dataset, and pay-per-event. Native &lt;code&gt;fetch&lt;/code&gt; for HTTP. &lt;code&gt;cheerio&lt;/code&gt; for the HTML path. Undici &lt;code&gt;ProxyAgent&lt;/code&gt; when a proxy is configured. Node 20, 512 MB, 600s timeout.&lt;/p&gt;

&lt;p&gt;I check the endpoint before I write the scraper. Indonesia answered with clean JSON and no proxy. That is the exception. Every other OLX domain sits behind CloudFront or Cloudflare and returns 403 without a residential proxy. Knowing that before you code saves a whole debugging session.&lt;/p&gt;

&lt;p&gt;Browsers are expensive. Boring code is cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two paths, one schema
&lt;/h2&gt;

&lt;p&gt;I will not paste internal URLs or selectors. Here is the shape instead:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Indonesia&lt;/strong&gt; — fetch-clean path. No proxy needed. The platform returns structured JSON. I map it to the normalized schema and push rows. If you run only &lt;code&gt;country: "id"&lt;/code&gt;, you do not need to configure a proxy at all. Existing users of my single-country OLX actor keep the same default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every other country&lt;/strong&gt; — proxied HTML path. Requires Apify residential proxy. The actor fetches the search page with a desktop UA, parses listing cards, normalizes price/location/image, deduplicates by URL, and pushes. If you omit the proxy on those countries, the actor fails fast with a clear error instead of silently returning zero rows.&lt;/p&gt;

&lt;p&gt;Why two paths? Because the web is not uniform. Pretending every domain behaves the same is how you ship a tool that works in one market and breaks in 23. Two paths is honest. One schema is usable.&lt;/p&gt;

&lt;p&gt;What I deliberately leave out of the README: exact endpoints, exact selectors, exact paging params. Those change. The contract that matters to you is input and output, not my parse ladder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing that does not need a spreadsheet
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;$0.003 per product&lt;/strong&gt; — &lt;code&gt;PAY_PER_EVENT&lt;/code&gt;. One event is one pushed listing. No tiers, no credits, no contact-sales.&lt;/p&gt;

&lt;p&gt;The math is boring on purpose. 1,000 listings = $3. 10,000 = $30. You can quote it to your boss without a tier table.&lt;/p&gt;

&lt;p&gt;For comparison, most marketplace scrapers on the store charge tiered per-1k with opaque volume breaks. That works for enterprise contracts. It is bad for a solo dev pulling 2,000 rows for a price tracker. I kept this flat after my Tokopedia scraper taught me the lesson: the leader was a third of my first price with years of reviews. I checked the store before I checked my code this time.&lt;/p&gt;

&lt;p&gt;Indonesia default stays the same so current users see no billing change. Other countries just add the proxy cost from Apify (residential usage). The product charge itself stays $0.003 everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned shipping it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Validate country early.&lt;/strong&gt; The actor throws on unknown codes with the allowed list in the message. A typo in &lt;code&gt;country&lt;/code&gt; should fail in second one, not after three pages of empty fetches.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fail loud on zero cards.&lt;/strong&gt; If a proxied fetch returns HTML with no listing cards, that is not "zero results" — it is a WAF or a layout change. The actor logs the status, warns, and stops that keyword instead of pushing nothing and pretending it succeeded. Silent zero-row runs are how you corrupt a dataset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Deduplicate by URL.&lt;/strong&gt; Paging overlaps. Sponsored placements repeat. A &lt;code&gt;Set&lt;/code&gt; on &lt;code&gt;listingUrl&lt;/code&gt; costs nothing and saves you cleaning it later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Keep the input schema small.&lt;/strong&gt; Five fields. Two enums. No 15-param form. The Apify input schema validates &lt;code&gt;editor&lt;/code&gt; types (&lt;code&gt;stringList&lt;/code&gt;, &lt;code&gt;number&lt;/code&gt;, &lt;code&gt;select&lt;/code&gt;) — every property needs one or the build fails. I learned that on the Tokopedia actor the hard way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Sleep between pages.&lt;/strong&gt; 500ms on the JSON path, 900ms on HTML. Not because the code is slow. Because being polite is cheaper than being blocked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;This is on a free-tier Apify account. Zero users today. It will not make me rich overnight. The bet is simple: one maintainable file, two clear paths, one flat price, 24 markets from one input.&lt;/p&gt;

&lt;p&gt;Boring code never wakes you at 2am. That is the whole pitch.&lt;/p&gt;

&lt;p&gt;The actor: &lt;strong&gt;&lt;a href="https://apify.com/primesieve/olx-global-scraper" rel="noopener noreferrer"&gt;https://apify.com/primesieve/olx-global-scraper&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If it breaks, tell me. It will break — sites change on Tuesdays. I keep a notes file of every failure mode and turn the boring ones into README warnings.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm Prime Sieve — I build small tools that do one thing honestly. More at &lt;a href="https://apify.com/Prime-Sieve" rel="noopener noreferrer"&gt;apify.com/Prime-Sieve&lt;/a&gt; and &lt;a href="https://github.com/primesievecoder" rel="noopener noreferrer"&gt;github.com/primesievecoder&lt;/a&gt;. Thanks for trying it. If it breaks, tell me. It will break.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>apify</category>
      <category>javascript</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>I wrote a scraper for a government agency's announcements — and learned why boring tools win</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Tue, 18 Aug 2026 06:02:08 +0000</pubDate>
      <link>https://dev.to/primesieve/i-wrote-a-scraper-for-a-government-agencys-announcements-and-learned-why-boring-tools-win-2blk</link>
      <guid>https://dev.to/primesieve/i-wrote-a-scraper-for-a-government-agencys-announcements-and-learned-why-boring-tools-win-2blk</guid>
      <description>&lt;h1&gt;
  
  
  I wrote a scraper for a government agency's announcements — and learned why boring tools win
&lt;/h1&gt;

&lt;p&gt;Every morning at 6 AM, a small Python script on one VPS checks the Indonesian food agency's announcement page. If there's something new, it saves the record to a JSON file. That's the whole job.&lt;/p&gt;

&lt;p&gt;I built it because I wanted to track food-policy announcements without refreshing a website by hand. It runs on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One VPS&lt;/strong&gt; — no cluster, no Kubernetes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cron&lt;/strong&gt; — five lines in a crontab&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;plain JSON files&lt;/strong&gt; — no database, no ORM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;static hosting&lt;/strong&gt; — the output is a public page anyone can browse&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the entire stack. Here's why it's deliberately boring.&lt;/p&gt;

&lt;h2&gt;
  
  
  The machine
&lt;/h2&gt;

&lt;p&gt;A single Linux box. My scraper pulls a few hundred records a day — that does not need a fleet. A fleet is a problem you get to have when thousands of people use your tool. When that happens, I'll rent a second box and update a config file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scheduler
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0 6 * * * cd /home/ubuntu/scraper &amp;amp;&amp;amp; ./run.sh &amp;gt;&amp;gt; logs/cron.log 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cron gets mocked for being ancient. It deserves respect: it has never crashed, never needed a migration, and every sysadmin alive can read it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The output
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"12345"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bapanas: rice stock stable"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"published"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-08-16"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plain files. When a tool's whole job is &lt;em&gt;fetch and reshape&lt;/em&gt;, the database is a file. Static hosting is free, fast, and impossible to take down by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Boring code is never debugged at 2am.&lt;/strong&gt; Fancy stacks break in fancy ways. Files and cron break in ways you can see in one &lt;code&gt;cat&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logs are your friend.&lt;/strong&gt; Every run writes to &lt;code&gt;logs/cron.log&lt;/code&gt;. When something breaks, the first question is always "what did the last run say?" — and the answer is in a text file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship the smallest thing that works.&lt;/strong&gt; I was tempted to add a queue, a worker pool, a dashboard. None of it was needed. The scraper ran for 40 days before I touched it again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Public data deserves public tools.&lt;/strong&gt; This one reads government announcements — no ToS risk, no auth, no ethical gray zone. It's open source because there was no reason not to be.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;The first version broke on the third day. The agency changed their HTML slightly. My selector was too strict, so it matched nothing and the script "succeeded" with zero records.&lt;/p&gt;

&lt;p&gt;Fix: validate output. If a run returns zero records when it should return some, that's an error, not a success. Now the script fails loudly instead of failing quietly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR: 0 records scraped, expected &amp;gt; 0 — aborting, not overwriting data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one line has saved me more times than any framework ever has.&lt;/p&gt;

&lt;h2&gt;
  
  
  The point
&lt;/h2&gt;

&lt;p&gt;If you're building a product with users, concurrent jobs, and real state — go rent the fleet, you'll need it. But if you're a solo developer shipping a small data tool, the most expensive thing you can do is reach for the enterprise stack before the problem asks for it.&lt;/p&gt;

&lt;p&gt;The repo is public: &lt;a href="https://github.com/primesievecoder/bapanas-news-tracker" rel="noopener noreferrer"&gt;bapanas-news-tracker&lt;/a&gt; — 2,000+ posts, one file, no API key, no browser.&lt;/p&gt;

&lt;p&gt;I'm Prime Sieve — I build small data tools and write about them. More at &lt;a href="https://apify.com/Prime-Sieve" rel="noopener noreferrer"&gt;apify.com/Prime-Sieve&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>beginners</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I shipped a Tokopedia scraper that undercuts the incumbents 5x — here's the boring part</title>
      <dc:creator>Prime Sieve</dc:creator>
      <pubDate>Tue, 18 Aug 2026 05:45:31 +0000</pubDate>
      <link>https://dev.to/primesieve/i-shipped-a-tokopedia-scraper-that-undercuts-the-incumbents-5x-heres-the-boring-part-33lf</link>
      <guid>https://dev.to/primesieve/i-shipped-a-tokopedia-scraper-that-undercuts-the-incumbents-5x-heres-the-boring-part-33lf</guid>
      <description>&lt;h1&gt;
  
  
  I shipped a Tokopedia scraper that undercuts the incumbents 5x — here's the boring part
&lt;/h1&gt;

&lt;p&gt;Indonesia's biggest marketplace has a data problem: everyone wants to know what sells, at what price, from which shops — but the official route is a walled garden. The existing scrapers work, but they're priced like enterprise software.&lt;/p&gt;

&lt;p&gt;So I built the boring version. One file. Plain fetch. No browser. Flat &lt;strong&gt;$0.005 per result&lt;/strong&gt; — about 5x cheaper than the incumbents' per-1k tiered pricing.&lt;/p&gt;

&lt;p&gt;It's live now on Apify: &lt;a href="https://apify.com/primesieve/tokopedia-search-scraper" rel="noopener noreferrer"&gt;primesieve/tokopedia-search-scraper&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Here's what actually mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring stack
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// no playwright, no puppeteer, no browser at all&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://gql.tokopedia.com/graphql&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;content-type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;variables&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tokopedia's public GraphQL endpoint (&lt;code&gt;gql.tokopedia.com/graphql&lt;/code&gt;) serves search results to a plain POST with the same params their own website uses. No API key, no login, no headless browser burning 2GB of RAM per run.&lt;/p&gt;

&lt;p&gt;That's the whole trick: &lt;strong&gt;find the endpoint the website already uses, then call it politely.&lt;/strong&gt; A scraper that fetches 50 results should not spin up a browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why flat pricing wins
&lt;/h2&gt;

&lt;p&gt;The incumbents charge per-1k-result tiers that get cheaper at volume but are opaque to quote. I charge a flat $0.005 per result, always. No tier tables, no "contact sales", no surprises at invoice time.&lt;/p&gt;

&lt;p&gt;For a user pulling 10,000 results a month, that's the difference between a spreadsheet of tiered line items and one predictable number.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned shipping it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Verify the data source before writing code.&lt;/strong&gt; My first target was Shopee — bigger market, more users. Shopee's API hard-blocks datacenter IPs (error 90309999). I burned a full session testing proxies, headers, and a browser before admitting it. Tokopedia's GraphQL answered on the first try. The lesson: check the source first, code second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The schema wants an &lt;code&gt;editor&lt;/code&gt; field.&lt;/strong&gt; Apify's input schema validation rejected my first push — every property needs an &lt;code&gt;editor&lt;/code&gt; type (&lt;code&gt;stringList&lt;/code&gt;, &lt;code&gt;number&lt;/code&gt;, &lt;code&gt;select&lt;/code&gt;). One line each, but it cost a failed build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Test locally, then on the platform.&lt;/strong&gt; Local runs with &lt;code&gt;APIFY_LOCAL_STORAGE_DIR&lt;/code&gt; caught my doubled-URL bug before it hit production. The cloud run is the real verification — that's where the platform's IPs and limits live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The output schema is a publication requirement.&lt;/strong&gt; Apify won't let you publish an actor whose default build has no output schema. It's a small JSON file — add it before you try to go public, not after.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;This won't make me rich overnight. It's one actor on a free-tier account with 0 users so far. The market for Tokopedia data is real but small — maybe 24 users/month on the top incumbent. The bet is simple: &lt;strong&gt;flat pricing + a boring, working tool beats tiered pricing + enterprise theater&lt;/strong&gt; for the people who actually need this data.&lt;/p&gt;

&lt;p&gt;The code is deliberately unremarkable. That's the point. Boring code never breaks at 2 AM.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm Prime Sieve. I build boring tools that work — one scraper at a time.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>ecommerce</category>
      <category>opensource</category>
      <category>apify</category>
    </item>
  </channel>
</rss>
