A script that runs at 6 a.m. every day has an odd problem. You only hear from it when something goes wrong, and sometimes not even then.
No news could mean everything worked. It could also mean the script crashed, the schedule stopped, or the output was quietly wrong for a week.
This is the code companion to my guide on how to monitor your automations when something breaks, which covers the ideas. Here we build them for a Python script, using one small helper file and a GitHub Actions workflow. Everything is free to run.
The four things worth having
Each one answers a different question:
- Logs tell you what happened.
- Failure alerts tell you when something broke.
- Result checks catch runs that finish without an error but produce the wrong result.
- A heartbeat tells you when the script never ran at all.
You don't need all four for every script. Logs and a failure alert cover most small jobs. The other two are for scripts where a wrong result, or a missing run, would cost you something.
Part of the skill is knowing which scripts deserve this effort in the first place. My post on when not to automate is a useful gut check before you add monitoring to everything.
The example job
To keep this realistic, let's use a script that syncs new sign-ups from a form to a spreadsheet every morning.
It works fine at first. Then someone updates the form, the email field comes through differently, and the script keeps running without complaint. It just saves rows with no email address in them.
No error, no crash and a green checkmark. A script like this is exactly where a result check earns its place.
Step 1: Add the helper file
Create a file called watch.py:
"""Small helpers for scripts that run without anyone watching."""
import json
import logging
import os
import sys
import time
import traceback
import urllib.request
from datetime import datetime, timezone
WEBHOOK = os.environ.get("DISCORD_WEBHOOK", "")
HEARTBEAT_URL = os.environ.get("HEARTBEAT_URL", "").rstrip("/")
logging.basicConfig(level=logging.INFO, stream=sys.stdout,
format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger("job")
class ResultCheck(Exception):
"""Raised when the script ran without errors, but the result looks wrong."""
def check(condition, message):
if not condition:
raise ResultCheck(message)
def _send(url, data=None):
headers = {"User-Agent": "watched-job/1.0"}
if data is not None:
data = json.dumps(data).encode()
headers["Content-Type"] = "application/json"
request = urllib.request.Request(url, data=data, headers=headers)
urllib.request.urlopen(request, timeout=15).close()
def alert(text):
"""Tell a person. A failed alert is logged, but never hides the real problem."""
if not WEBHOOK:
return
try:
_send(WEBHOOK, {"content": text[:1900]})
except Exception as error:
log.warning("Could not send the alert: %s", error)
def heartbeat(suffix=""):
"""Ping a heartbeat service, for example Healthchecks.io."""
if not HEARTBEAT_URL:
return
for attempt in range(3):
try:
_send(HEARTBEAT_URL + suffix)
return
except Exception as error:
log.warning("Could not send the heartbeat (attempt %d): %s", attempt + 1, error)
time.sleep(1)
def run_link():
"""A link to this run, when running in GitHub Actions."""
if not os.environ.get("GITHUB_RUN_ID"):
return ""
return "{}/{}/actions/runs/{}".format(os.environ.get("GITHUB_SERVER_URL", "https://github.com"),
os.environ["GITHUB_REPOSITORY"], os.environ["GITHUB_RUN_ID"])
def run(job, name):
"""Run job() with start and finish logging, and make sure any failure is loud."""
started = time.time()
log.info("%s started", name)
heartbeat("/start")
try:
result = job()
except Exception as error:
crashed = not isinstance(error, ResultCheck)
title = f"{name} crashed" if crashed else f"{name}: the result looks wrong"
detail = traceback.format_exc() if crashed else str(error)
log.error("%s\n%s", title, detail)
if os.environ.get("GITHUB_ACTIONS"):
print(f"::error::{title}") # shows up as a red note on the run page
when = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
alert(f"**{title}**\nWhen: {when}\n{detail[-1400:]}\n{run_link()}".strip())
heartbeat("/fail")
sys.exit(1)
log.info("%s finished in %.1f seconds: %s", name, time.time() - started, result)
if os.environ.get("GITHUB_STEP_SUMMARY"): # a short summary on the run page
with open(os.environ["GITHUB_STEP_SUMMARY"], "a") as summary:
summary.write(f"### {name}\n{result}\n")
heartbeat()
return result
The helper has five small parts:
-
logwrites timestamped lines to the output. GitHub Actions keeps that output for each run. -
check()raises a "result looks wrong" error when a condition you care about isn't true. -
alert()sends a Discord message. If the alert itself fails, that gets logged, and it never hides the original problem. -
heartbeat()pings a monitoring service, with a few retries. -
run()wraps your job. It logs the start and finish, catches crashes and failed checks, sends an alert, and ends the script with an error code.
That last point matters. A non-zero exit code is how a script says "this failed". Without it, a script that catches its own error and carries on looks like a success.
Step 2: Write the job
Create job.py:
import csv
import json
import os
import urllib.request
from watch import check, log, run
def job():
with urllib.request.urlopen(os.environ["SOURCE_URL"], timeout=20) as response:
signups = json.load(response)
log.info("Fetched %d sign-ups", len(signups))
if not signups:
return "no new sign-ups today" # a quiet day is normal, not a failure
missing = [s for s in signups if not s.get("email")]
check(not missing, f"{len(missing)} of {len(signups)} sign-ups have no email address. "
"Did the form or the API change?")
with open("signups.csv", "w", newline="") as file: # stand-in for your real work
writer = csv.DictWriter(file, fieldnames=["name", "email"], extrasaction="ignore")
writer.writeheader()
writer.writerows(signups)
return f"saved {len(signups)} sign-ups"
run(job, "signup-sync")
Swap the middle for your own work, like updating a spreadsheet or sending an email. The pattern stays the same:
- Put the work in a function.
- Log the numbers that matter.
- Add a check for what "wrong" looks like.
- Hand the function to
run().
Notice the empty case. A quiet day with zero sign-ups is normal, so the script returns a message and carries on. If you made that a failure, you'd get false alarms, and people stop reading alerts that cry wolf.
What to put in a log
A log is a note to your future self at 7 a.m., trying to work out what happened. Write lines that help.
A healthy run looks like this:
2026-10-02 06:23:11,429 INFO signup-sync started
2026-10-02 06:23:12,031 INFO Fetched 14 sign-ups
2026-10-02 06:23:12,402 INFO signup-sync finished in 1.0 seconds: saved 14 sign-ups
Some guidelines:
- Log counts, not just "done". "Fetched 14 sign-ups" is useful, and you'll notice when it suddenly says 0.
- Log decisions. If the script skips something, retries something or ignores duplicates, say so and say how many.
-
Use the levels the way they're meant.
INFOis normal,WARNINGis odd but handled, andERRORmeans a person should look. - Keep secrets and personal data out. Never log passwords, tokens or full webhook addresses. Log counts or IDs instead of email addresses. Logs are easy to share by accident.
Step 3: Add the workflow
Create .github/workflows/signup-sync.yml:
name: Sign-up sync
on:
schedule:
- cron: '23 6 * * *' # every day at 06:23 UTC
workflow_dispatch: # adds a "Run workflow" button for testing
permissions:
contents: read
jobs:
sync:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4 # use the latest major version
- name: Run the job
env:
SOURCE_URL: https://example.com/api/signups # change this
DISCORD_WEBHOOK: ${{ secrets.DISCORD_WEBHOOK }}
HEARTBEAT_URL: ${{ secrets.HEARTBEAT_URL }}
run: python3 job.py
When the script exits with an error, GitHub marks the run as failed. With default notification settings, you also get an email.
Two small extras come from the helper. A failed run shows a red note at the top of the run page, and a successful one adds a one-line summary, such as "saved 14 sign-ups". You can see the result without opening the logs.
The odd start time avoids the top of the hour, when GitHub is busiest. The timeout stops a hung script from running for hours.
Step 4: Make the alert useful
Add a DISCORD_WEBHOOK secret in your repository (Settings, then Secrets and variables, then Actions). A webhook is a private address that lets a script post into a Discord channel. If the idea is new, I explain it in what a webhook is.
When something breaks, the message looks like this:
signup-sync: the result looks wrong
When: 2026-10-02 06:23 UTC
2 of 3 sign-ups have no email address. Did the form or the API change?
https://github.com/yourname/yourrepo/actions/runs/123456789
That's a deliberate shape. A good alert answers three questions: what broke, why, and when. The link on the last line takes you straight to the failed run.
An alert that only says "job failed" sends you digging. One that says what it expected and what it found often tells you the fix.
Step 5: Add a heartbeat
Everything so far only works if the script runs. What if it doesn't?
In a public repository, GitHub switches off scheduled workflows after 60 days with no repository activity. Schedules can also be disabled by mistake. In both cases there's no failed run, and no email, because nothing ran.
A heartbeat catches this. The script pings a monitoring service each time it finishes. If a ping doesn't arrive on time, the service alerts you. It's sometimes called a dead man's switch.
Healthchecks.io is one service that does this, and it has a free plan (check their pricing page for current limits). To set it up:
- Create an account and add a new check
- Give it a clear name, like "signup-sync"
- Set the period to how often the job should run (one day for ours)
- Set the grace time to how late you'll tolerate. An hour or two is sensible, because GitHub can delay scheduled runs
- Choose where its alerts should go
- Copy the check's ping URL and save it as a repository secret named
HEARTBEAT_URL
The helper already does the rest. It sends three kinds of signal:
- Start, when the job begins
- Success, when it finishes properly
- Fail, when it crashes or a result check fails
If a job starts but never finishes within the grace time, the service treats that as a failure too. That covers scripts that hang.
Treat the ping URL like a password, since anyone who has it can send pings to your check. The pings carry no data from your script, only the address.
If what you need to watch is a website rather than a script, I covered that in my Python uptime monitor guide, and there's a ready-made free Python uptime monitor template.
Step 6: Break it on purpose
Don't wait for a real failure to find out whether your alerts work. Run the workflow with the "Run workflow" button and try three things:
-
Point
SOURCE_URLat an address that doesn't exist. You should get a "crashed" alert, with the error in it. - Return data with a missing email. You should get the "result looks wrong" alert.
- Disable the workflow in the Actions tab and wait. After the period and grace time pass, the heartbeat service should alert you. That's the "never ran" case, and nothing else would catch it.
Then put everything back. If any of the three stays silent, you've learned something worth knowing.
Choosing good result checks
Result checks work best when they're boring. A few that suit most scripts:
- Something came back when it should have. If the source normally has data every day, a total of zero is suspicious.
- Required fields are there. This is the sign-up example. If a field you depend on is blank, stop.
- The count is in a normal range. If you usually process 20 to 40 records, 0 and 500 are both worth a look.
- The data is fresh. If the newest record is three days old, the source may have stopped updating.
These matter because the thing you depend on can change without warning. An API can start returning a different format, and your script keeps running as if nothing happened.
Pick checks that wouldn't fire on a normal quiet day. If a check fails often for harmless reasons, you'll learn to ignore it.
What can go wrong
Too many alerts. If every small hiccup pings you, you'll mute the channel. Alert only on things you'd act on, and let minor ones sit in the log.
Checks that are too strict. A rule that fires every weekend gets disabled. Tune it until a failure always means something.
The alert channel is broken. If someone deletes the webhook, your alerts go nowhere. The script logs a warning, and the heartbeat service still has its own way of reaching you. That's one reason to have both.
Secrets in logs. Print counts and IDs, not tokens or personal details. Look through a few real logs after the first runs.
A script that hides its own errors. If you wrap code in try/except and carry on, the run looks fine. Let errors reach run(), or call check().
Quick answers
Why not just rely on GitHub's failure emails?
They only cover runs that fail. They don't catch a run that finishes with the wrong result, or a schedule that stopped, and they're easy to miss in a busy inbox.
Do I need all four parts?
No. Start with logs and a failure alert. Add a result check where wrong data would cost you, and a heartbeat for scheduled scripts you depend on.
Can I use Slack or Telegram instead of Discord?
Yes. Change the alert() function to match what that service expects. The rest stays the same.
Does it work outside GitHub Actions?
Yes. The helper works anywhere Python runs, such as cron on a server. The run link, red note and summary only appear inside GitHub Actions.
Does it cost anything?
The workflow is free on a public repository and uses a few minutes a month on a private one. Healthchecks.io has a free plan, but check its current limits.
The short version
A script nobody watches needs to speak up for itself. Logs show what happened, alerts say when it broke, result checks catch wrong answers, and a heartbeat notices when it goes quiet.
Start small. Add the helper, wrap one important script and break it on purpose to see what happens.
You can find more practical automation guides on Procwire.
Top comments (0)