<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nitish Mane</title>
    <description>The latest articles on DEV Community by Nitish Mane (@nitish_mane_2c399ba312e53).</description>
    <link>https://dev.to/nitish_mane_2c399ba312e53</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3941301%2F492cc1e7-f3a4-4fc5-bc07-32ff7f30d98e.png</url>
      <title>DEV Community: Nitish Mane</title>
      <link>https://dev.to/nitish_mane_2c399ba312e53</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nitish_mane_2c399ba312e53"/>
    <language>en</language>
    <item>
      <title>I Built Scrapers That Heal Themselves (and Won an NVIDIA DGX Spark)</title>
      <dc:creator>Nitish Mane</dc:creator>
      <pubDate>Fri, 09 Oct 2026 04:24:11 +0000</pubDate>
      <link>https://dev.to/nitish_mane_2c399ba312e53/i-built-scrapers-that-heal-themselves-and-won-an-nvidia-dgx-spark-k13</link>
      <guid>https://dev.to/nitish_mane_2c399ba312e53/i-built-scrapers-that-heal-themselves-and-won-an-nvidia-dgx-spark-k13</guid>
      <description>&lt;h1&gt;
  
  
  I Built Scrapers That Heal Themselves (and Won an NVIDIA DGX Spark)
&lt;/h1&gt;

&lt;p&gt;Scraper Factory took first place at the WeMakeDevs Zero Downtime Hackathon in San Francisco. On the surface it is a small product: it scrapes about twelve days of TV guide data for Philo, YouTube TV and Sling, ranks the top ten upcoming shows and the top ten upcoming movies, and gives you one click to record each title on Philo.&lt;/p&gt;

&lt;p&gt;The interesting part is underneath. I wanted scrapers that notice when they break, attempt their own repair, and only accept that repair once it is proven correct. Three tools carried the project, each doing what it is best at: Bright Data for scraping and self-healing, Port as the control plane, and SigNoz for the audit trail. This post walks through how they fit together.&lt;/p&gt;

&lt;p&gt;Live webapp: &lt;a href="https://scraper-factory-top20.vercel.app" rel="noopener noreferrer"&gt;scraper-factory-top20.vercel.app&lt;/a&gt;&lt;br&gt;
Source code: &lt;a href="https://github.com/Nitishmane/scraper-factory" rel="noopener noreferrer"&gt;github.com/Nitishmane/scraper-factory&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: scrapers rot without anyone noticing
&lt;/h2&gt;

&lt;p&gt;Every scraper I have built eventually fails. A site ships a redesign, a selector stops matching, and the data either stops arriving or, worse, keeps arriving slightly wrong. Times land in the wrong column, titles swap with descriptions, rows go missing. You usually find out a week later when something downstream looks odd.&lt;/p&gt;

&lt;p&gt;Most teams treat scrapers as scripts. Scripts don't have owners, health states or an on-call rotation. I wanted each scraper treated like any other production service: state you can see, automation that reacts to that state, and a clear record of every repair.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bright Data: scrapers that repair themselves
&lt;/h2&gt;

&lt;p&gt;The usual approach to scraping is structural: you find that the airtime lives in &lt;code&gt;div.card &amp;gt; span:nth-child(3)&lt;/code&gt; and write that down. That selector says where the data sits, not what it means. Rename a CSS class and it breaks. Reorder the cards and it breaks. The worse case is when it doesn't fail loudly and keeps returning data that is slightly wrong.&lt;/p&gt;

&lt;p&gt;Bright Data's Scraper Studio builds a collector from a natural language description of the data instead of selectors. That changes what a scraper is: a short paragraph about meaning. A description of meaning survives a redesign far better than a selector does. And the prompt becomes the thing you heal from: when a scraper drifts, the &lt;code&gt;scraper heal&lt;/code&gt; command reasons over that same description plus a plain-language report of what went wrong.&lt;/p&gt;

&lt;p&gt;The guide data comes from a third-party site where every channel has its own page: roughly 455 program cards covering about twelve days. Because every channel page has the same structure, one shared collector serves all three providers (137 Philo channels, 175 YouTube TV channels, 119 Sling channels). A second collector reads Philo's browse pages so the ranked list can link straight to a record button.&lt;/p&gt;

&lt;p&gt;The repair loop works like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The factory calls &lt;code&gt;scraper heal&lt;/code&gt; with the collector ID and a symptom description written by a triage agent.&lt;/li&gt;
&lt;li&gt;Bright Data reasons over the original prompt and returns a proposal in an &lt;code&gt;awaiting_approval&lt;/code&gt; state, with a preview of what the repaired scraper would extract.&lt;/li&gt;
&lt;li&gt;A deterministic verifier (plain Python, no model) diffs that preview against a hand-checked expected output for a frozen fixture.&lt;/li&gt;
&lt;li&gt;If the diff passes, the factory approves the proposal. If it fails, it rejects it, marks the scraper broken, and sends it to a person.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;awaiting_approval&lt;/code&gt; state is what makes this safe. The model proposes, my code checks, and only then does anything change.&lt;/p&gt;

&lt;p&gt;One practical tip: generate the collector against a small, clean sample. I trimmed a real channel page down to 30 program cards and created the collector from that. That same trimmed page later became my frozen test fixture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Port: the control plane
&lt;/h2&gt;

&lt;p&gt;Scrapers needed to be entities with state, not scripts. Port's data model made that possible: seven blueprints model the system (data_source, scraper, scrape_run, heal_event, ranked_title, requirement, factory_deployment). Each scraper has a health property that moves between healthy, drifting and broken. Every run is a scrape_run entity. Every repair attempt is a heal_event recording its trigger, the drift description and the outcome.&lt;/p&gt;

&lt;p&gt;Port drives the daily loop with no person involved:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A scheduled workflow refreshes the Philo catalog at 07:30 UTC.&lt;/li&gt;
&lt;li&gt;A second scheduled workflow starts the guide scrape at 08:00 UTC.&lt;/li&gt;
&lt;li&gt;When a scrape_run lands with status success, a workflow publishes the ranked titles back to the Context Lake.&lt;/li&gt;
&lt;li&gt;When a run lands in error or drift, a heal workflow fires.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Four AI agents sit next to the catalog. Triage reads run and heal history and writes the drift description used for repair (approval required). A Heal Explainer describes in plain words why a repair was approved or rejected. An On-Call Assistant answers operator questions. A Catalog Coverage agent reports how many ranked titles have a direct Philo record link. The agent's job stays narrow on purpose: Triage writes the symptom report but never approves its own fix. That decision stays with plain code.&lt;/p&gt;

&lt;p&gt;The 20 ranked_title entities are the product itself. The public Next.js webapp has no database of its own; it reads those entities straight from Port's Context Lake. Publishing is just an upsert.&lt;/p&gt;

&lt;p&gt;Everything was provisioned through the API by one idempotent bootstrap script, so the whole setup could be rebuilt at any time.&lt;/p&gt;

&lt;h2&gt;
  
  
  SigNoz: the audit trail
&lt;/h2&gt;

&lt;p&gt;Any system that changes itself needs a clear record of what it did and why. SigNoz carried every trace, metric and log, and its alerting acted as the safety net for the repair loop.&lt;/p&gt;

&lt;p&gt;A scrape run produces one trace with spans for each step: fetch per channel page, extract, validate, and when a repair runs, heal-reprompt and verify spans. Because the web layer is auto-instrumented, even Port's automation calling the heal endpoint shows up as a span. Logs carry trace and span IDs, so I could jump from a log line to the exact span it came from.&lt;/p&gt;

&lt;p&gt;Detection has two paths calling the same idempotent heal endpoint. The fast path: the inline detector marks the scraper as drifting in Port within seconds. The safety net: a SigNoz alert fires when drift is detected over a five-minute window and calls the heal endpoint through a webhook about five minutes later. Whichever arrives first does the work. A second rule fires when data stops arriving altogether, which is the only signal that catches a pipeline that died before it could run its own detector. A pipeline that never runs never reports drift either.&lt;/p&gt;

&lt;p&gt;The dashboard (overview stats, run failures, control-plane activity, Bright Data budget) was created by a script and embedded in Port's operator view, so the on-call screen has everything in one place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unhappy path, on purpose
&lt;/h2&gt;

&lt;p&gt;The clearest proof came from a fault I injected during the hackathon. A scrape run landed with an error, the heal workflow fired, the Triage agent wrote up the symptom, and the repair endpoint ran a real cycle against Bright Data. The proposal came back with a preview of 2 rows where the fixture expected 30. My verifier rejected it. Port recorded the heal event as rolled back, set the scraper to broken, and left it for a person. Nothing bad reached the product. The next clean scheduled run moved health back to healthy on its own.&lt;/p&gt;

&lt;p&gt;Opening the SigNoz trace told the whole story: the fetch spans with their row counts, the heal span holding the symptom description, and the verify span whose log read "row count 2, expected 30." No guessing, no digging through separate log files.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would pass on
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Write prompts about meaning, not layout. The prompt is what the healer works from.&lt;/li&gt;
&lt;li&gt;Capture the raw string next to every parsed value. Raw present but parsed null turned out to be one of my most reliable drift signals.&lt;/li&gt;
&lt;li&gt;Treat a heal preview as a candidate. Check it against known good output before approving it.&lt;/li&gt;
&lt;li&gt;Model the unusual things. Treating scrapers, runs and repairs as entities is what made every other feature useful.&lt;/li&gt;
&lt;li&gt;Keep the agent's job narrow. Triage writes the symptom report; it does not approve its own fix.&lt;/li&gt;
&lt;li&gt;Alert on a metric, not a log line. And add an alert for silence.&lt;/li&gt;
&lt;li&gt;Keep the setup in scripts. Re-running one idempotent bootstrap beats clicking through settings at 2 AM during a hackathon.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  About the project
&lt;/h2&gt;

&lt;p&gt;Scraper Factory was built for the WeMakeDevs Zero Downtime Hackathon, where it won first place and an NVIDIA DGX Spark. Thanks to Kunal Kushwaha and the WeMakeDevs team for putting on a great event, and to Port, Bright Data and SigNoz for the tools that made it possible.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>webdev</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
