<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chris Healey</title>
    <description>The latest articles on DEV Community by Chris Healey (@chriscompiles).</description>
    <link>https://dev.to/chriscompiles</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4022219%2F69055c92-03e0-4c2e-9bb1-4ebce33d7de7.png</url>
      <title>DEV Community: Chris Healey</title>
      <link>https://dev.to/chriscompiles</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chriscompiles"/>
    <language>en</language>
    <item>
      <title>A month of PulseWatch: billing, better setup and deliberately broken scripts</title>
      <dc:creator>Chris Healey</dc:creator>
      <pubDate>Wed, 07 Oct 2026 09:26:05 +0000</pubDate>
      <link>https://dev.to/chriscompiles/a-month-of-pulsewatch-billing-better-setup-and-deliberately-broken-scripts-39m</link>
      <guid>https://dev.to/chriscompiles/a-month-of-pulsewatch-billing-better-setup-and-deliberately-broken-scripts-39m</guid>
      <description>&lt;p&gt;On 1 September, I wrote &lt;a href="https://dev.to/chriscompiles/who-watches-the-watchdog-the-boring-work-behind-a-monitoring-saas-7n8"&gt;Who watches the watchdog?&lt;/a&gt; about the less glamorous work behind PulseWatch: monitoring its own watchdog, explaining scheduler delays and making sure the support link reached a real person.&lt;/p&gt;

&lt;p&gt;I’ve been quiet here since then. The project has been rather less quiet.&lt;/p&gt;

&lt;p&gt;Over the past month, I’ve worked on the interface, billing, terms and privacy pages, and a simpler way to connect an existing script. Yesterday, I deliberately failed, stalled and stopped disposable jobs to check that the alerts actually reached my inbox.&lt;/p&gt;

&lt;p&gt;That last part was particularly satisfying. A monitoring product ought to be able to demonstrate what happens when something goes wrong.&lt;/p&gt;

&lt;p&gt;For anyone new to the project: PulseWatch monitors unattended scripts and scheduled jobs from outside the system running them. Your job sends a &lt;code&gt;/start&lt;/code&gt;, then a &lt;code&gt;/success&lt;/code&gt; or &lt;code&gt;/fail&lt;/code&gt;. A separate watchdog notices when expected work goes missing or a started run takes too long.&lt;/p&gt;

&lt;p&gt;The original motivation still holds: a script that never starts cannot report its own absence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making the interface useful when something breaks
&lt;/h3&gt;

&lt;p&gt;One of the bigger September changes was bringing the landing page, dashboard, monitor details, documentation and admin screens into a consistent interface.&lt;/p&gt;

&lt;p&gt;The useful part was deciding what belonged at the top.&lt;/p&gt;

&lt;p&gt;When an account already has monitors, their health should come before the form for creating another one. When a job fails, the monitor page should show the current problem and recent run evidence before asking you to think about configuration.&lt;/p&gt;

&lt;p&gt;That’s obvious when written down. It wasn’t how every screen was arranged.&lt;/p&gt;

&lt;p&gt;Private ping URLs are also masked by default, with reveal and copy controls. They’re credentials: anyone holding one can send signals to that monitor.&lt;/p&gt;

&lt;p&gt;The same pass included browser-form security work, server-side validation and checks at narrow mobile widths. Long error messages and tokens are very good at finding the weak spots in a layout.&lt;/p&gt;

&lt;p&gt;It was a useful reminder that a UI refresh can improve how someone investigates a failure, as well as how the product looks in a screenshot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Billing became a proper piece of application logic
&lt;/h3&gt;

&lt;p&gt;The early payment setup involved Stripe links and manually changing someone’s plan.&lt;/p&gt;

&lt;p&gt;That was manageable at a tiny scale, but it left too much dependent on me matching a payment to an account and remembering what should happen next.&lt;/p&gt;

&lt;p&gt;The newer flow links Checkout to the signed-in PulseWatch account. Confirmed payment is what grants paid access; returning to a success page is not enough.&lt;/p&gt;

&lt;p&gt;I’ve also worked through renewal handling and made the purchase pages clearer about the monthly price, tax treatment, when paid access begins and how to stop renewal. Terms, privacy and support pages are now published.&lt;/p&gt;

&lt;p&gt;In the sandbox, I verified an initial purchase and a successful renewal, including the extension of the account’s paid-through date.&lt;/p&gt;

&lt;p&gt;There is still billing verification to finish. The failed-renewal simulation hit a restriction in the Managed Payments setup, so that particular lifecycle test remains outstanding. Getting one successful payment through doesn’t prove every subscription edge case.&lt;/p&gt;

&lt;p&gt;I also investigated what initially looked like a payment bug. An account with a manually assigned plan was being stopped from starting a conflicting purchase, exactly as intended.&lt;/p&gt;

&lt;p&gt;The awkward part was the explanation. I improved the wording so the user could understand why they’d been stopped.&lt;/p&gt;

&lt;p&gt;A fairly representative solo-builder afternoon: investigate a possible bug, find the guard working, then discover that the message still needs fixing.&lt;/p&gt;

&lt;h3&gt;
  
  
  A free tool to get the first script connected
&lt;/h3&gt;

&lt;p&gt;The newest addition is &lt;a href="https://pulsewatch.ai/tools/monitor-my-script" rel="noopener noreferrer"&gt;Monitor my script&lt;/a&gt;, a free Python/Bash wrapper generator.&lt;/p&gt;

&lt;p&gt;It creates the start and outcome handling around an existing job. The Python version uses the standard library; the Bash version runs your command and preserves its exit status.&lt;/p&gt;

&lt;p&gt;The private ping URL comes from an environment variable.&lt;/p&gt;

&lt;p&gt;The important constraint is that monitoring should preserve the job’s result. If your script raises an exception, a failed monitoring request must not replace it. If your command exits with an error, the wrapper must return that error.&lt;/p&gt;

&lt;p&gt;The generated pings have five-second timeouts and guarded failure handling.&lt;/p&gt;

&lt;p&gt;This came from thinking about the first integration. Someone arriving with a working script shouldn’t have to piece together several examples before they can try monitoring it.&lt;/p&gt;

&lt;p&gt;I’ve also published a &lt;a href="https://pulsewatch.ai/guides/monitor-python-bash-script" rel="noopener noreferrer"&gt;Python and Bash setup guide&lt;/a&gt;, covering the wrapper, schedule settings, the first successful run and checking failure and recovery.&lt;/p&gt;

&lt;p&gt;Writing those steps down is its own product review. It exposes every place where you’ve assumed the reader already understands something.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then I tried breaking it
&lt;/h3&gt;

&lt;p&gt;Yesterday’s checks used disposable monitors and jobs to exercise the generated code against the live service.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;What I checked&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;An ordinary Python exception or a Bash command exiting with code 7&lt;/td&gt;
&lt;td&gt;Failure was recorded, the original job outcome was preserved, and failure and recovery emails arrived.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A started job left unfinished beyond its maximum runtime&lt;/td&gt;
&lt;td&gt;STUCK was detected, followed by recovery and both emails.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A job stopped after establishing a successful baseline&lt;/td&gt;
&lt;td&gt;MISSING was detected, followed by recovery and both emails.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I checked the dashboard and the actual iCloud inbox. An email provider accepting a request is useful evidence, but seeing the intended message arrive completes another part of the chain.&lt;/p&gt;

&lt;p&gt;The missing-run test needed its initial schedule setup correcting before it exercised the intended scenario. Worth recording: the test setup needs scrutiny too.&lt;/p&gt;

&lt;p&gt;Those checks didn’t introduce new failure states. They verified existing behaviour through the new integration path.&lt;/p&gt;

&lt;p&gt;They also reinforced two details the guide now makes explicit: finish a successful run to establish the baseline, and expect alerts on the next applicable watchdog check. The check cadence depends on the plan.&lt;/p&gt;

&lt;h3&gt;
  
  
  What does “success” actually mean?
&lt;/h3&gt;

&lt;p&gt;Feedback here has also pushed me to think harder about a job that runs, finishes and produces nothing useful.&lt;/p&gt;

&lt;p&gt;An export can exit cleanly while returning zero records. A report can complete without generating the file anyone needed.&lt;/p&gt;

&lt;p&gt;PulseWatch receives the outcome your integration reports. It doesn’t inspect arbitrary business output and decide whether the work was correct.&lt;/p&gt;

&lt;p&gt;I’ve extended the &lt;a href="https://pulsewatch.ai/docs" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; with an example of checking useful output before sending &lt;code&gt;/success&lt;/code&gt;. For a daily export where an empty result is a failure, a small local rule can raise an exception and take the existing &lt;code&gt;/fail&lt;/code&gt; path.&lt;/p&gt;

&lt;p&gt;The rule has to fit the job. Zero records might be perfectly valid for an import that only processes new entries.&lt;/p&gt;

&lt;p&gt;That seems a useful next step to explore with real users: help them define a meaningful success condition, then make the signal reflect it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keeping the priorities proportionate
&lt;/h3&gt;

&lt;p&gt;One correction to the previous post: I had put automated database backups next on the list. After reviewing the paid Render Postgres recovery provision, I accepted its existing point-in-time recovery as proportionate for the current stage and deferred additional backup work.&lt;/p&gt;

&lt;p&gt;I haven’t completed a separate restore drill. That distinction matters.&lt;/p&gt;

&lt;p&gt;There is still plenty to do, including finishing billing verification and getting more consequential workflows connected.&lt;/p&gt;

&lt;p&gt;The engineering progress is tangible. The commercial question still needs evidence from people who keep using it.&lt;/p&gt;

&lt;p&gt;For now, there’s a clearer interface, a more complete buying path, and a shorter route from an existing script to a verified first run. You can try it with two free monitors and no credit card.&lt;/p&gt;

&lt;p&gt;If you maintain scheduled jobs, have you ever had one finish “successfully” and still leave you with nothing useful?&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>python</category>
      <category>devops</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Who watches the watchdog? The boring work behind a monitoring SaaS</title>
      <dc:creator>Chris Healey</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:36:48 +0000</pubDate>
      <link>https://dev.to/chriscompiles/who-watches-the-watchdog-the-boring-work-behind-a-monitoring-saas-7n8</link>
      <guid>https://dev.to/chriscompiles/who-watches-the-watchdog-the-boring-work-behind-a-monitoring-saas-7n8</guid>
      <description>&lt;p&gt;My last post about PulseWatch ended with a satisfying discovery: the monitoring bug I was investigating wasn’t actually a bug.&lt;/p&gt;

&lt;p&gt;A GitHub Actions workflow configured to run every 15 minutes was sometimes going missing for well over an hour. PulseWatch noticed the absence, raised a MISSING alert and reported recovery when GitHub eventually ran the workflow again.&lt;/p&gt;

&lt;p&gt;The smoke alarm wasn’t broken. There really was smoke.&lt;/p&gt;

&lt;p&gt;That was useful validation for the product. It also exposed several less comfortable questions about the product itself.&lt;/p&gt;

&lt;p&gt;Who watches the PulseWatch watchdog?&lt;/p&gt;

&lt;p&gt;How should users configure grace periods when schedulers are unreliable?&lt;/p&gt;

&lt;p&gt;And if someone needs help, can they actually contact me?&lt;/p&gt;

&lt;p&gt;So, for the past couple of weeks, I haven’t been building a clever new feature. I’ve been working through the unglamorous jobs that make a monitoring product less fragile.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The watchdog now has its own watchdog&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;PulseWatch has a separate scheduled service that checks every monitor and decides whether it is OK, MISSING, FAILED or STUCK.&lt;/p&gt;

&lt;p&gt;That separation is fundamental to how the product works. A monitored script cannot report that it never started, so something outside the script has to notice its absence.&lt;/p&gt;

&lt;p&gt;But there was an awkward flaw in my architecture:&lt;/p&gt;

&lt;p&gt;If the PulseWatch watchdog stopped running, nobody would receive an alert — including me.&lt;/p&gt;

&lt;p&gt;The web application could still be online. Jobs could continue sending pings. Run history could continue accumulating in the database.&lt;/p&gt;

&lt;p&gt;But the process responsible for evaluating those runs and sending alerts could be dead.&lt;/p&gt;

&lt;p&gt;For a monitoring product, that’s a fairly serious blind spot.&lt;/p&gt;

&lt;p&gt;I’ve now instrumented the watchdog using Healthchecks.io, which runs outside the PulseWatch infrastructure.&lt;/p&gt;

&lt;p&gt;Each watchdog cycle sends:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a start signal when checking begins&lt;/li&gt;
&lt;li&gt;a success signal after the full cycle completes&lt;/li&gt;
&lt;li&gt;a failure signal if the cycle exits with an unhandled exception&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Healthchecks also knows how frequently the watchdog should run, so it can raise an alert if the process never starts at all.&lt;/p&gt;

&lt;p&gt;The integration is deliberately best-effort and non-blocking. If Healthchecks is unavailable, that must not stop PulseWatch from checking its users’ monitors. Monitoring the monitor should never break the thing being monitored.&lt;/p&gt;

&lt;p&gt;It creates a simple chain:&lt;/p&gt;

&lt;p&gt;User's job → PulseWatch watchdog → external watchdog&lt;/p&gt;

&lt;p&gt;Nothing is perfectly self-monitoring, but the failure domains are now separated. PulseWatch is no longer responsible for noticing the death of its own alerting process.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The scheduler incident became documentation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The GitHub Actions incident also showed me that PulseWatch’s configuration needed much better explanation.&lt;/p&gt;

&lt;p&gt;A monitor has three important time settings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Expected interval: how often the job should complete&lt;/li&gt;
&lt;li&gt;Grace: additional time allowed before a missing run triggers an alert&lt;/li&gt;
&lt;li&gt;Max runtime: how long an active run may remain unfinished before it is considered stuck&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These sound obvious until a scheduler configured for 15-minute intervals produces gaps of 90 minutes or more.&lt;/p&gt;

&lt;p&gt;A 15-minute schedule does not mean a 15-minute grace period is safe.&lt;/p&gt;

&lt;p&gt;GitHub explicitly documents that scheduled workflows can be delayed during periods of high load and that, under sufficiently high load, some queued jobs may be dropped.&lt;/p&gt;

&lt;p&gt;The right grace period therefore depends on the behaviour you actually observe, not just the cron expression you wrote.&lt;/p&gt;

&lt;p&gt;That lesson is now captured in a new public PulseWatch documentation page.&lt;/p&gt;

&lt;p&gt;It includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a 60-second getting-started guide&lt;/li&gt;
&lt;li&gt;the /start → /success or /fail lifecycle&lt;/li&gt;
&lt;li&gt;dependency-free Python integration&lt;/li&gt;
&lt;li&gt;a Bash/curl example&lt;/li&gt;
&lt;li&gt;expected interval, grace and max-runtime guidance&lt;/li&gt;
&lt;li&gt;a concrete warning about scheduler jitter&lt;/li&gt;
&lt;li&gt;definitions of MISSING, STUCK, FAILED, flapping and NEVER_RUN&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also made the examples defensive.&lt;/p&gt;

&lt;p&gt;Monitoring requests use short timeouts and do not replace the original job’s error. The Bash example preserves the job’s actual exit code. Failure messages are URL-encoded. The example ping URLs are deliberately fake so nobody accidentally publishes a real monitor token.&lt;/p&gt;

&lt;p&gt;The page is linked from the main navigation and footer, and the public route now has its own focused test.&lt;/p&gt;

&lt;p&gt;None of this is particularly exciting engineering. It may be more valuable than another feature.&lt;/p&gt;

&lt;p&gt;A monitoring tool that is configured incorrectly can behave exactly as designed and still be useless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. “Priority support” now has an actual support channel&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;PulseWatch’s paid Indie tier listed priority support.&lt;/p&gt;

&lt;p&gt;Until recently, there was no obvious way to request it.&lt;/p&gt;

&lt;p&gt;That is the sort of gap that survives when you’re concentrating on the application itself: the pricing page makes a promise, but the operational plumbing behind the promise doesn’t exist.&lt;/p&gt;

&lt;p&gt;There is now a live support link and a working domain email address at &lt;a href="mailto:chris@pulsewatch.ai"&gt;chris@pulsewatch.ai&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I tested the complete loop rather than stopping once outgoing mail worked:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send from the PulseWatch address&lt;/li&gt;
&lt;li&gt;Receive it in an external inbox&lt;/li&gt;
&lt;li&gt;Reply to it&lt;/li&gt;
&lt;li&gt;Confirm the reply arrives in the correct PulseWatch inbox&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Again, boring work. But if someone is deciding whether to trust a tiny monitoring service run by one person, “can I reach that person when something breaks?” is not a minor detail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The pattern I keep finding&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first version of PulseWatch was mostly about capability:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can it receive pings?&lt;/li&gt;
&lt;li&gt;Can it detect a missing job?&lt;/li&gt;
&lt;li&gt;Can it send an alert?&lt;/li&gt;
&lt;li&gt;Can it show run history?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The work after launch has increasingly been about trust:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Will its own watchdog failure be noticed?&lt;/li&gt;
&lt;li&gt;Can a user configure it without creating false alerts?&lt;/li&gt;
&lt;li&gt;Are the examples safe to copy?&lt;/li&gt;
&lt;li&gt;Is there a real person behind the support link?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those jobs don’t produce impressive launch screenshots. They don’t make the feature list much longer. They do make the product less likely to fail in an embarrassing way.&lt;/p&gt;

&lt;p&gt;There are still gaps.&lt;/p&gt;

&lt;p&gt;Automated database backups, including proving that a restore actually works, are next. Terms of Service and Privacy Policy pages also need shipping. More ambitious work, including detecting abnormal behaviour in unattended and AI-driven workflows, comes after that.&lt;/p&gt;

&lt;p&gt;For now, I think a monitoring product has to earn the right to become clever by first becoming dependable.&lt;/p&gt;

&lt;p&gt;If you were deciding whether to trust a tiny monitoring service with your unattended jobs, what evidence would you want to see next?&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>devops</category>
      <category>python</category>
      <category>showdev</category>
    </item>
    <item>
      <title>The monitoring bug that turned out not to be a bug</title>
      <dc:creator>Chris Healey</dc:creator>
      <pubDate>Thu, 13 Aug 2026 15:21:41 +0000</pubDate>
      <link>https://dev.to/chriscompiles/the-monitoring-bug-that-turned-out-not-to-be-a-bug-103e</link>
      <guid>https://dev.to/chriscompiles/the-monitoring-bug-that-turned-out-not-to-be-a-bug-103e</guid>
      <description>&lt;p&gt;I spent the last couple of days debugging what looked like a fairly nasty problem in PulseWatch, the dead-man’s-switch monitoring tool I’ve been building.&lt;/p&gt;

&lt;p&gt;My inbox was filling up with alerts:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MISSING → OK → MISSING → OK&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Overnight, I received 29 of them.&lt;/p&gt;

&lt;p&gt;Not ideal behaviour for a product whose job is to reduce monitoring noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I thought was happening
&lt;/h2&gt;

&lt;p&gt;PulseWatch monitors unattended jobs from the outside.&lt;/p&gt;

&lt;p&gt;A job sends a simple HTTP request when it starts and another when it completes. PulseWatch stores those runs and a server-side watchdog checks whether a successful run has arrived within the expected interval plus a configurable grace period.&lt;/p&gt;

&lt;p&gt;If not, the monitor becomes &lt;code&gt;MISSING&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When a new successful run arrives, it becomes &lt;code&gt;OK&lt;/code&gt; again.&lt;/p&gt;

&lt;p&gt;The canary I use to test PulseWatch is itself a scheduled GitHub Actions workflow. It is configured to run every 15 minutes. At first I assumed my alert state logic was broken.&lt;/p&gt;

&lt;p&gt;I increased the PulseWatch threshold to an hour plus a 30-minute grace period but the flip-flopping continued.&lt;/p&gt;

&lt;p&gt;Typical alerts looked like this:&lt;/p&gt;

&lt;p&gt;05:10 — MISSING — last success 1h34m ago&lt;br&gt;
05:30 — OK — new successful run&lt;br&gt;
07:00 — MISSING — last success 1h32m ago&lt;br&gt;
07:05 — OK — new successful run&lt;br&gt;
08:35 — MISSING — last success 1h33m ago&lt;br&gt;
08:50 — OK — new successful run&lt;/p&gt;

&lt;p&gt;That looked to me to be suspiciously precise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Following the state transitions
&lt;/h2&gt;

&lt;p&gt;So, what did I do? I got Codex to work tracing the relevant code.&lt;/p&gt;

&lt;p&gt;PulseWatch declares a monitor missing when:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;now - latest_success &amp;gt; expected_interval + grace_period&lt;br&gt;
&lt;/code&gt;&lt;br&gt;
With the monitor configured for 60 minutes plus 30 minutes of grace, that gives a 90-minute boundary, so those MISSING alerts were correct.&lt;/p&gt;

&lt;p&gt;More importantly, the watchdog cannot recover a monitor by itself. An &lt;code&gt;OK&lt;/code&gt; recovery requires PulseWatch to receive a new &lt;code&gt;/success&lt;/code&gt; request.&lt;/p&gt;

&lt;p&gt;So the state machine wasn't oscillating, something really was sending successful pings after each MISSING alert and that moved the investigation upstream.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I looked at GitHub Actions
&lt;/h2&gt;

&lt;p&gt;The canary workflow is configured to run every 15 minutes.&lt;/p&gt;

&lt;p&gt;Its actual run history looked more like this:&lt;/p&gt;

&lt;p&gt;00:05&lt;br&gt;
00:53&lt;br&gt;
03:35&lt;br&gt;
05:27&lt;br&gt;
07:00&lt;br&gt;
08:45&lt;br&gt;
09:55&lt;br&gt;
11:16&lt;br&gt;
12:12&lt;br&gt;
13:01&lt;br&gt;
14:36&lt;/p&gt;

&lt;p&gt;One gap was &lt;strong&gt;2 hours 42 minutes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Another was &lt;strong&gt;1 hour 52 minutes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Another was &lt;strong&gt;1 hour 45 minutes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The interesting part was that the executions GitHub &lt;em&gt;did&lt;/em&gt; run were green. The canary itself generally completed in around 10–20 seconds.&lt;/p&gt;

&lt;p&gt;There weren't corresponding failed 15-minute runs filling those gaps. The scheduled executions simply weren't happening reliably.&lt;/p&gt;

&lt;p&gt;GitHub documents that scheduled Actions can be delayed and, under sufficiently high load, some queued jobs may be dropped.&lt;/p&gt;

&lt;h2&gt;
  
  
  The smoke alarm wasn't broken
&lt;/h2&gt;

&lt;p&gt;This completely changed my interpretation of the incident. PulseWatch wasn't generating false positives, it was detecting a genuine failure of the system it was monitoring.&lt;/p&gt;

&lt;p&gt;Take the 07:00 example; the previous GitHub Actions execution had occurred at approximately 05:27. Nothing successfully ran for the next 90 minutes and so PulseWatch crossed its configured threshold and raised &lt;code&gt;MISSING&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;GitHub eventually executed the workflow again at 07:00 which is when the canary sent its success ping...PulseWatch reported recovery exactly as designed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The more interesting monitoring lesson
&lt;/h2&gt;

&lt;p&gt;This highlighted a failure mode that's easy to overlook, that is logs are excellent at telling you what happened &lt;strong&gt;inside a process that ran&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;GitHub Actions was similarly quite capable of showing me green ticks for the workflows it executed successfully.&lt;/p&gt;

&lt;p&gt;Neither helps much with the more awkward question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What about the job that never ran at all?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is no exception. There is no failed execution. There may be no application log. Nothing inside the job can report the failure because the job never started.&lt;/p&gt;

&lt;p&gt;That's precisely where an external dead-man's switch becomes useful.&lt;/p&gt;

&lt;p&gt;It doesn't ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Did the job report an error?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Did I hear from the job when I expected to?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those turn out to be quite different questions.&lt;/p&gt;

&lt;p&gt;I thought I’d found a bug in PulseWatch. Instead, I’d found the exact failure PulseWatch was built to catch.&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>devops</category>
      <category>testing</category>
    </item>
    <item>
      <title>Two Bugs, Two Strangers, One Week: What Shipping Early Actually Buys You</title>
      <dc:creator>Chris Healey</dc:creator>
      <pubDate>Fri, 17 Jul 2026 09:29:10 +0000</pubDate>
      <link>https://dev.to/chriscompiles/two-bugs-two-strangers-one-week-what-shipping-early-actually-buys-you-16im</link>
      <guid>https://dev.to/chriscompiles/two-bugs-two-strangers-one-week-what-shipping-early-actually-buys-you-16im</guid>
      <description>&lt;p&gt;A week ago I put a rough, honestly-a-bit-thin version of PulseWatch in front of real people for the first time. Within days, two different strangers — independently, unprompted — found two real gaps in it. Neither was catastrophic. Both were exactly the kind of thing you only find by watching someone else use the thing you built.&lt;/p&gt;

&lt;p&gt;This is the story of both, and the fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug one: the run that never ends
&lt;/h2&gt;

&lt;p&gt;This first bug came from a friend testing it on a real script. His question was simple:&lt;/p&gt;

&lt;p&gt;"What happens if &lt;code&gt;start&lt;/code&gt; fires twice before &lt;code&gt;end&lt;/code&gt;?"&lt;/p&gt;

&lt;p&gt;Good question. At the time: nothing good. Here's why.&lt;/p&gt;

&lt;p&gt;PulseWatch works on two pings — a job calls &lt;code&gt;/start&lt;/code&gt; when it begins and &lt;code&gt;/success&lt;/code&gt; (or &lt;code&gt;/fail&lt;/code&gt;) when it's done. The server tracks whichever run is currently "open" for a monitor. The bug: if a job's process restarts mid-run — a crash-and-retry, a redeploy that catches it mid-flight, a scheduler firing twice — you get a second &lt;code&gt;/start&lt;/code&gt; before the first run ever closes. The old run just sits there, open forever, an orphan with no ending. Worse, because the watchdog was still waiting on that run's expected finish time, it could fire a false "still running" alert for a run that was, for all practical purposes, dead and abandoned.&lt;/p&gt;

&lt;p&gt;The fix is a small rule with an outsized effect: a new &lt;code&gt;/start&lt;/code&gt; supersedes whatever run is currently open. The old run gets marked superseded — a terminal, non-alerting status — and a fresh run begins clean. The watchdog was updated to treat superseded as a dead end: nothing to wait on, nothing to alert about, and it never shows up in a user's run history. It's not a failure and it's not a success. It's just "this run doesn't matter anymore, a newer one replaced it."&lt;/p&gt;

&lt;p&gt;The logic, roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle_start&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;open_run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_open_run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;open_run&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;open_run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;superseded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;open_run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;finished_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;new_run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;running&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;started_at&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;new_run&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple once you see it. Invisible until someone actually restarts a job mid-flight and asks the right question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug two: the failure that isn't a failure, exactly
&lt;/h2&gt;

&lt;p&gt;The second came from a comment on my first post here on Dev.to, from someone whose day job is infrastructure monitoring. He'd read the piece on catching jobs that die silently and asked the harder follow-up:&lt;/p&gt;

&lt;p&gt;"How does PulseWatch deal with jobs that aren't dead, just flaky, before alert fatigue kicks in?"&lt;/p&gt;

&lt;p&gt;That's a different failure mode entirely, and at the time I only had half an answer.&lt;/p&gt;

&lt;p&gt;PulseWatch already alerts on state transitions, not conditions — one email when a monitor goes down, one when it recovers, not a repeat every check cycle. That solves the "job has been down for six hours, stop emailing me about it" problem. What it didn't solve: a job that's actually unstable — fails, recovers, fails, recovers, each transition genuinely real. Alert on every transition, and you get exactly what it sounds like: fail, recover, fail, recover, five emails in twenty minutes, and you mute the whole channel by Thursday. Technically correct alerting, practically useless.&lt;/p&gt;

&lt;p&gt;The naive fix — just suppress alerts if a monitor's been unstable recently — has an obvious failure mode of its own: it delays the one alert you actually wanted immediately, buried under a threshold meant for noise. That tradeoff is exactly what I flagged as unsolved in my reply to him at the time.&lt;/p&gt;

&lt;p&gt;Here's where I landed. Rather than alerting on every up/down transition, PulseWatch now watches a rolling window and counts transitions within it. Cross a threshold — four or more state changes within two hours, by default — and the monitor enters a distinct flapping state. Two things change once it's flapping:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;One alert fires on entry — "this monitor is unstable," not "this monitor is down," a meaningfully different signal.&lt;/li&gt;
&lt;li&gt;Ordinary up/down alerts are suppressed while flapping continues, and one more alert fires when it settles back below the flap threshold.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So a genuinely flaky job gets exactly two emails — "it's flapping" and "it's settled" — no matter how many times it actually flipped in between. A job that's just plain down still gets its immediate down alert, because that's a single transition, not a pattern — the threshold only kicks in once instability itself becomes the story.&lt;/p&gt;

&lt;p&gt;The state machine, simplified:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;FLAP_WINDOW_MINUTES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;120&lt;/span&gt;   &lt;span class="c1"&gt;# rolling window to look back over
&lt;/span&gt;&lt;span class="n"&gt;FLAP_THRESHOLD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;          &lt;span class="c1"&gt;# this many state changes in the window = flapping
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_currently_flapping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Count how many times status changed across recent runs.
    Flapping if it crosses the threshold within the window.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;runs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recent_terminal_runs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;minutes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;FLAP_WINDOW_MINUTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# oldest → newest
&lt;/span&gt;    &lt;span class="n"&gt;changes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:])&lt;/span&gt;
                  &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;changes&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;FLAP_THRESHOLD&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_monitor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;currently&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;is_currently_flapping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;currently&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_flapping&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_flapping&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="nf"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unstable — flapping between states&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;currently&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_flapping&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_flapping&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
        &lt;span class="nf"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stabilised&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_flapping&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# normal single-transition alerting, unchanged
&lt;/span&gt;        &lt;span class="nf"&gt;handle_normal_transition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The one new piece of state this needed — a boolean is_flapping column — became, almost by accident, the first schema change to go through a proper migration pipeline I'd just set up on the live database. Small feature, useful excuse to prove the plumbing works before something bigger needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual point
&lt;/h2&gt;

&lt;p&gt;Neither of these bugs was findable by staring at the code longer. The double-start case only shows up when a real process actually restarts mid-run. The flapping case only matters once you've watched enough alert emails pile up to feel the fatigue yourself. Both needed someone else's script, someone else's failure pattern, someone else's patience to ask "but what about—".&lt;/p&gt;

&lt;p&gt;That's the actual case for shipping something rough and putting it in front of people early, not the version of the advice that's become a cliché. It's not just "get feedback." It's that some categories of bug are structurally invisible to the person who wrote the code, because the code matches their own mental model of how it'll be used — and the whole value of a stranger is that their mental model is different.&lt;/p&gt;

&lt;p&gt;Both fixes are live now. If you're running unattended jobs — cron, scheduled scripts, agent pipelines — and any of this sounds familiar, PulseWatch is free to try, and I'd genuinely like to know what breaks next.&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>python</category>
      <category>showdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>How I Built a Dead-Man's Switch for My AI Trading Pipeline in Python</title>
      <dc:creator>Chris Healey</dc:creator>
      <pubDate>Thu, 09 Jul 2026 06:58:39 +0000</pubDate>
      <link>https://dev.to/chriscompiles/how-i-built-a-dead-mans-switch-for-my-ai-trading-pipeline-in-python-2ndm</link>
      <guid>https://dev.to/chriscompiles/how-i-built-a-dead-mans-switch-for-my-ai-trading-pipeline-in-python-2ndm</guid>
      <description>&lt;h2&gt;
  
  
  The morning I noticed my trading bot hadn't traded in three days.
&lt;/h2&gt;

&lt;p&gt;I run a small trading bot. Every morning before the market opens it pulls signals, weighs them, and decides whether to act. It had been humming along for weeks, so I'd mostly stopped watching it.&lt;/p&gt;

&lt;p&gt;Then one morning I opened the dashboard out of idle curiosity and something was off. No trades. Not "it decided not to trade", just &lt;em&gt;crickets&lt;/em&gt;. I scrolled back. No trades the day before either. Or the day before that.&lt;/p&gt;

&lt;p&gt;Three days. The bot had silently stopped running and I had no idea.&lt;/p&gt;

&lt;p&gt;Here's the part that stuck with me: there was no error to find. No red text, no stack trace, no alert in my inbox. The overnight process that was supposed to kick everything off had just… not kicked off. A scheduling hiccup on my machine, most likely. The code was fine. The code simply never ran.&lt;/p&gt;

&lt;p&gt;And every single tool I had for catching problems was completely blind to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why everything I had was looking the wrong way
&lt;/h2&gt;

&lt;p&gt;Think about how we normally catch failures. We wrap risky things in &lt;code&gt;try/except&lt;/code&gt;. We log errors. We wire up something to shout when an exception gets thrown:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;run_morning_routine&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Morning routine failed: %s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;send_alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is good practice and you should do it. But notice the assumption baked into every line: &lt;strong&gt;it assumes the code ran.&lt;/strong&gt; The &lt;code&gt;try&lt;/code&gt; block has to execute for the &lt;code&gt;except&lt;/code&gt; to ever fire. The logger has to be reached for it to log. The alert has to be triggered by a process that is, by definition, alive.&lt;/p&gt;

&lt;p&gt;None of that helps you when the process never starts.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The server rebooted and the scheduler didn't come back up.&lt;/li&gt;
&lt;li&gt;The cron entry got wiped in a config change.&lt;/li&gt;
&lt;li&gt;The disk filled and the job couldn't even launch.&lt;/li&gt;
&lt;li&gt;A dependency upstream hung and your job was never invoked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In every one of these cases there is no exception, because there is no running code to throw one. There's no log line, because nothing got far enough to write it. Your monitoring is sitting there patiently waiting for a signal that will never come, and interpreting the silence as "all good."&lt;/p&gt;

&lt;p&gt;That's the trap. &lt;strong&gt;Absence of an error is not the same as success.&lt;/strong&gt; For unattended jobs like cron tasks, scheduled scripts, overnight pipelines, AI agents running on a timer etc. the most dangerous failure is the one that produces no output at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inverting the model: a dead-man's switch
&lt;/h2&gt;

&lt;p&gt;The fix is to stop waiting for bad news and start &lt;em&gt;expecting good news&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;A dead-man's switch (the name comes from the pedal that stops a train if the driver goes unresponsive) flips the logic around. Instead of your code reporting &lt;strong&gt;when it fails&lt;/strong&gt;, it reports &lt;strong&gt;when it succeeds&lt;/strong&gt; and something &lt;em&gt;external&lt;/em&gt; watches for that report to arrive. If the expected check-in doesn't show up on time, the watcher raises the alarm.&lt;/p&gt;

&lt;p&gt;The crucial word is &lt;em&gt;external&lt;/em&gt;. The thing doing the watching has to live somewhere your job doesn't. If the watchdog runs on the same machine as the job, then when that machine dies, the watchdog dies with it and a dead watchdog can't tell you anything is wrong. You need something running elsewhere whose entire purpose is to notice an absence.&lt;/p&gt;

&lt;p&gt;Concretely, three pieces:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Your job sends a ping when it starts, and another when it finishes.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A server somewhere records those pings and knows roughly when to expect the next one.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;If the expected ping doesn't arrive inside a grace window, the server alerts you.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The silence becomes the signal. You just need something that's actually listening for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building it in Python
&lt;/h2&gt;

&lt;p&gt;Let me show the shape of it. This is deliberately minimal. The point is the pattern, not a framework.&lt;/p&gt;

&lt;h3&gt;
  
  
  The client side (what your job adds)
&lt;/h3&gt;

&lt;p&gt;The job needs to announce itself at the start and confirm at the end. Pure standard library, no new dependencies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;

&lt;span class="n"&gt;PING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://your-watchdog.example/ping/&amp;lt;your-token&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;PING&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Monitoring must never crash the job it monitors.
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;

&lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;run_morning_routine&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things worth calling out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The ping is wrapped so it can &lt;strong&gt;never take down the real job&lt;/strong&gt;. Monitoring that crashes the thing it's monitoring is worse than no monitoring.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;start&lt;/code&gt; and &lt;code&gt;success&lt;/code&gt; bracket the work, so the watcher can tell "never started" apart from "started but never finished" (a hang) apart from "finished cleanly."&lt;/li&gt;
&lt;li&gt;It's a plain HTTP GET, so this works from any language. A shell job can do the same with &lt;code&gt;curl&lt;/code&gt;:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; https://your-watchdog.example/ping/&amp;lt;token&amp;gt;/start
./run_backup.sh &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; https://your-watchdog.example/ping/&amp;lt;token&amp;gt;/success &lt;span class="se"&gt;\&lt;/span&gt;
                &lt;span class="o"&gt;||&lt;/span&gt; curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; https://your-watchdog.example/ping/&amp;lt;token&amp;gt;/fail
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The receiving side
&lt;/h3&gt;

&lt;p&gt;On the server, receiving a ping is just recording that it happened and when:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timezone&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;flask&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Flask&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Flask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/ping/&amp;lt;token&amp;gt;/&amp;lt;event&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;receive_ping&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;monitor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;first_or_404&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;last_event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;
    &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;last_ping_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;utc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;last_finished_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;last_ping_at&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing clever here, and that's the point. The receiver's only job is to remember the last time it heard from you.&lt;/p&gt;

&lt;h3&gt;
  
  
  The watchdog — the part that actually matters
&lt;/h3&gt;

&lt;p&gt;The interesting logic is the part that runs &lt;em&gt;on its own schedule&lt;/em&gt;, independent of any job, and asks: is anyone overdue?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timedelta&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_all&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;utc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;Monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;deadline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;last_ping_at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;minutes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;period_minutes&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grace_minutes&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;is_late&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;deadline&lt;/span&gt;
        &lt;span class="n"&gt;new_status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DOWN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_late&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

        &lt;span class="c1"&gt;# Only act on a *change* of state.
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;new_status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;new_status&lt;/span&gt;
            &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;new_status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DOWN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="nf"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Your job hasn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t checked in — it may have stopped.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="nf"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;monitor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Your job is reporting again — recovered.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two design decisions here earned their keep, and both are about &lt;em&gt;not being annoying&lt;/em&gt;:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. State-transition alerting.&lt;/strong&gt; You alert on the &lt;em&gt;change&lt;/em&gt;, not on the condition. A job that's been down for six hours should generate exactly one "it's down" email, not one every time the watchdog loops. And exactly one "it's back" email when it recovers. Alert on the edge, not the level. This is the single biggest thing standing between "useful monitoring" and "inbox noise you'll mute within a day."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A grace window.&lt;/strong&gt; Jobs are never perfectly punctual; a run that usually takes ten minutes occasionally takes twenty. The &lt;code&gt;grace_minutes&lt;/code&gt; buffer stops those normal wobbles from firing false alarms. A decent rule of thumb is to set the grace to roughly 20–30% of the expected run time, with a sensible floor of a few minutes.&lt;/p&gt;

&lt;p&gt;Run &lt;code&gt;check_all()&lt;/code&gt; on its own timer - its own cron entry, a scheduled worker, whatever just somewhere separate from the jobs it's watching. That separation is the whole game. The watchdog has to outlive the thing it watches.&lt;/p&gt;

&lt;h2&gt;
  
  
  From "fix for me" to "thing other people can use"
&lt;/h2&gt;

&lt;p&gt;I built the version above for exactly one job: my trading bot. It worked. The next morning the bot failed to start again, and this time my phone buzzed instead of me finding out three days later by accident.&lt;/p&gt;

&lt;p&gt;But once I had it, I started pointing it at other things almost reflexively. A nightly database backup which, it turned out, had quietly failed a week earlier and I'd never noticed. A scraper on a schedule. A content pipeline. Every unattended job I owned had the same blind spot, and I'd just been getting lucky that none of them had bitten me yet.&lt;/p&gt;

&lt;p&gt;That's when it stopped feeling like a personal script and started feeling like something other solo devs probably needed too. Especially the growing pile of us running AI agents and LLM pipelines on timers, where "did the agent actually run this morning?" is a real and slightly unnerving question. So I cleaned it up, put a proper dashboard on it, and turned it into a small product called &lt;a href="https://pulsewatch.ai" rel="noopener noreferrer"&gt;PulseWatch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The whole idea is what's above: your job sends a start and a finish ping, a server watches for them, and you get one email when something goes quiet and one when it comes back. There's a free tier with no card required if you want to point it at a job and see it work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one thing to take away
&lt;/h2&gt;

&lt;p&gt;Even if you never use a tool for this, even if you go build the twenty lines yourself, internalise the core idea, because it changes how you think about reliability:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your monitoring is only watching for the failures it can see. The failure that produces no output is invisible to everything that waits for output.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Go look at your most important unattended job right now and ask one question: &lt;em&gt;if this silently stopped running tonight, what would tell me?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If the honest answer is "I'd notice eventually, when something downstream breaks" then that's the gap. And the fix is genuinely about twenty lines of code away.&lt;/p&gt;

</description>
      <category>python</category>
      <category>devops</category>
      <category>showdev</category>
      <category>buildinpublic</category>
    </item>
  </channel>
</rss>
