On 1 September, I wrote Who watches the watchdog? about the less glamorous work behind PulseWatch: monitoring its own watchdog, explaining scheduler delays and making sure the support link reached a real person.
I’ve been quiet here since then. The project has been rather less quiet.
Over the past month, I’ve worked on the interface, billing, terms and privacy pages, and a simpler way to connect an existing script. Yesterday, I deliberately failed, stalled and stopped disposable jobs to check that the alerts actually reached my inbox.
That last part was particularly satisfying. A monitoring product ought to be able to demonstrate what happens when something goes wrong.
For anyone new to the project: PulseWatch monitors unattended scripts and scheduled jobs from outside the system running them. Your job sends a /start, then a /success or /fail. A separate watchdog notices when expected work goes missing or a started run takes too long.
The original motivation still holds: a script that never starts cannot report its own absence.
Making the interface useful when something breaks
One of the bigger September changes was bringing the landing page, dashboard, monitor details, documentation and admin screens into a consistent interface.
The useful part was deciding what belonged at the top.
When an account already has monitors, their health should come before the form for creating another one. When a job fails, the monitor page should show the current problem and recent run evidence before asking you to think about configuration.
That’s obvious when written down. It wasn’t how every screen was arranged.
Private ping URLs are also masked by default, with reveal and copy controls. They’re credentials: anyone holding one can send signals to that monitor.
The same pass included browser-form security work, server-side validation and checks at narrow mobile widths. Long error messages and tokens are very good at finding the weak spots in a layout.
It was a useful reminder that a UI refresh can improve how someone investigates a failure, as well as how the product looks in a screenshot.
Billing became a proper piece of application logic
The early payment setup involved Stripe links and manually changing someone’s plan.
That was manageable at a tiny scale, but it left too much dependent on me matching a payment to an account and remembering what should happen next.
The newer flow links Checkout to the signed-in PulseWatch account. Confirmed payment is what grants paid access; returning to a success page is not enough.
I’ve also worked through renewal handling and made the purchase pages clearer about the monthly price, tax treatment, when paid access begins and how to stop renewal. Terms, privacy and support pages are now published.
In the sandbox, I verified an initial purchase and a successful renewal, including the extension of the account’s paid-through date.
There is still billing verification to finish. The failed-renewal simulation hit a restriction in the Managed Payments setup, so that particular lifecycle test remains outstanding. Getting one successful payment through doesn’t prove every subscription edge case.
I also investigated what initially looked like a payment bug. An account with a manually assigned plan was being stopped from starting a conflicting purchase, exactly as intended.
The awkward part was the explanation. I improved the wording so the user could understand why they’d been stopped.
A fairly representative solo-builder afternoon: investigate a possible bug, find the guard working, then discover that the message still needs fixing.
A free tool to get the first script connected
The newest addition is Monitor my script, a free Python/Bash wrapper generator.
It creates the start and outcome handling around an existing job. The Python version uses the standard library; the Bash version runs your command and preserves its exit status.
The private ping URL comes from an environment variable.
The important constraint is that monitoring should preserve the job’s result. If your script raises an exception, a failed monitoring request must not replace it. If your command exits with an error, the wrapper must return that error.
The generated pings have five-second timeouts and guarded failure handling.
This came from thinking about the first integration. Someone arriving with a working script shouldn’t have to piece together several examples before they can try monitoring it.
I’ve also published a Python and Bash setup guide, covering the wrapper, schedule settings, the first successful run and checking failure and recovery.
Writing those steps down is its own product review. It exposes every place where you’ve assumed the reader already understands something.
Then I tried breaking it
Yesterday’s checks used disposable monitors and jobs to exercise the generated code against the live service.
| Test | What I checked |
|---|---|
| An ordinary Python exception or a Bash command exiting with code 7 | Failure was recorded, the original job outcome was preserved, and failure and recovery emails arrived. |
| A started job left unfinished beyond its maximum runtime | STUCK was detected, followed by recovery and both emails. |
| A job stopped after establishing a successful baseline | MISSING was detected, followed by recovery and both emails. |
I checked the dashboard and the actual iCloud inbox. An email provider accepting a request is useful evidence, but seeing the intended message arrive completes another part of the chain.
The missing-run test needed its initial schedule setup correcting before it exercised the intended scenario. Worth recording: the test setup needs scrutiny too.
Those checks didn’t introduce new failure states. They verified existing behaviour through the new integration path.
They also reinforced two details the guide now makes explicit: finish a successful run to establish the baseline, and expect alerts on the next applicable watchdog check. The check cadence depends on the plan.
What does “success” actually mean?
Feedback here has also pushed me to think harder about a job that runs, finishes and produces nothing useful.
An export can exit cleanly while returning zero records. A report can complete without generating the file anyone needed.
PulseWatch receives the outcome your integration reports. It doesn’t inspect arbitrary business output and decide whether the work was correct.
I’ve extended the documentation with an example of checking useful output before sending /success. For a daily export where an empty result is a failure, a small local rule can raise an exception and take the existing /fail path.
The rule has to fit the job. Zero records might be perfectly valid for an import that only processes new entries.
That seems a useful next step to explore with real users: help them define a meaningful success condition, then make the signal reflect it.
Keeping the priorities proportionate
One correction to the previous post: I had put automated database backups next on the list. After reviewing the paid Render Postgres recovery provision, I accepted its existing point-in-time recovery as proportionate for the current stage and deferred additional backup work.
I haven’t completed a separate restore drill. That distinction matters.
There is still plenty to do, including finishing billing verification and getting more consequential workflows connected.
The engineering progress is tangible. The commercial question still needs evidence from people who keep using it.
For now, there’s a clearer interface, a more complete buying path, and a shorter route from an existing script to a verified first run. You can try it with two free monitors and no credit card.
If you maintain scheduled jobs, have you ever had one finish “successfully” and still leave you with nothing useful?
Top comments (0)