DEV Community

mattleeee
mattleeee

Posted on Originally published at hkcode.dpdns.org

Health Checks for a Fleet of Scheduled Tasks

Nine unattended scheduled tasks is a fleet, whether you call it that or not. Once you pass three or four, checking taskschd.msc by hand stops being a habit and starts being a liability — you only look when something is visibly broken, which is exactly the failure mode you were trying to avoid by scheduling the work in the first place.

This post covers a small audit script that enumerates a set of tasks by name pattern, reads their state and last result, compares them against an expected schedule, and pushes an alert when something is wrong. It runs as a scheduled task itself, because a monitoring script that depends on you remembering to run it is not monitoring.

What the Task Scheduler actually exposes

The Task Scheduler COM API (Schedule.Service) is the authoritative source. The schtasks.exe text output is fine for humans and terrible for parsing — column widths shift, and localized Windows installs translate the field names. Go through COM.

import win32com.client

TASK_ENUM_HIDDEN = 1

def connect():
    service = win32com.client.Dispatch("Schedule.Service")
    service.Connect()
    return service

def iter_tasks(service, folder_path="\\"):
    folder = service.GetFolder(folder_path)
    for task in folder.GetTasks(TASK_ENUM_HIDDEN):
        yield task
Enter fullscreen mode Exit fullscreen mode

Each IRegisteredTask gives you a handful of fields worth reading:

Field Type What it tells you
Name str Task name, without folder path
Path str Full path including folder, e.g. \Quant\Quant_Factor
State int 0 Unknown, 1 Disabled, 2 Queued, 3 Ready, 4 Running
Enabled bool Separate from state; a disabled task can still report Ready in some edge cases
LastRunTime datetime When the last run started
LastTaskResult int Exit code of the last completed run
NextRunTime datetime Next scheduled fire, 1899-12-30 if none

Two traps here. First, State and Enabled are not redundant. A task can be enabled but sitting in Disabled state if its trigger was disabled, or disabled at the task level while the state enum lags. Check both.

Second, LastTaskResult is only meaningful after a run completes. While a task is running, LastTaskResult still holds the previous run's value. If you alert on a non-zero result without checking State == 4, you will page yourself about a task that is mid-run and about to succeed.

The 1899-12-30 sentinel for "no next run" is a COM date artifact — it's the zero of the OLE automation date format. Treat anything before 1980 as "none".

Defining what "healthy" means

The script needs a table of expectations, because "task exists and ran once" is not a health check. Each task has a cadence, and a task that last ran three days ago when it should run hourly is broken even if its exit code was zero.

from dataclasses import dataclass
from datetime import timedelta

@dataclass(frozen=True)
class Expectation:
    pattern: str            # fnmatch pattern against task Path
    max_age: timedelta      # how stale LastRunTime may be
    allow_disabled: bool = False
    description: str = ""

EXPECTED = [
    Expectation("\\Quant\\Quant_*",        timedelta(hours=2),  description="quant pipeline"),
    Expectation("\\TeleAgent_*",           timedelta(minutes=30), description="telegram agent"),
    Expectation("\\Quant\\Quant_Nightly",  timedelta(hours=26), description="nightly batch"),
]
Enter fullscreen mode Exit fullscreen mode

max_age should be generous enough to absorb a single missed cycle plus jitter, but not so generous that a task can silently skip twice. For an hourly task, two hours is a reasonable ceiling. For a nightly job, 26 hours covers a late start without hiding a full day of silence.

The allow_disabled flag matters for tasks you intentionally park. Without it, every disabled task shows up as a finding, and you learn to ignore the alerts — which defeats the point.

Classifying each task

The classification logic is where most naive audits go wrong. There are four distinct failure modes and they need different messages.

import fnmatch
from datetime import datetime, timezone

STATE_UNKNOWN, STATE_DISABLED, STATE_QUEUED, STATE_READY, STATE_RUNNING = range(5)

def classify(task, expectation, now):
    """Return (severity, reason) where severity is 'ok', 'warn' or 'fail'."""
    state = task.State
    enabled = task.Enabled
    last_result = task.LastTaskResult
    last_run = task.LastRunTime

    if not enabled or state == STATE_DISABLED:
        if expectation.allow_disabled:
            return "ok", "disabled by design"
        return "fail", "task is disabled"

    if state == STATE_RUNNING:
        # Don't judge LastTaskResult while a run is in flight.
        return "ok", "currently running"

    if last_result != 0:
        return "fail", f"last result {last_result:#010x}"

    if last_run is None or last_run.year < 1980:
        return "fail", "never ran"

    # COM returns naive local datetimes; compare in local time.
    age = now - last_run
    if age > expectation.max_age:
        hours = age.total_seconds() / 3600
        return "fail", f"stale, last run {hours:.1f}h ago"

    return "ok", f"ok, last run {age.total_seconds()/60:.0f}m ago"
Enter fullscreen mode Exit fullscreen mode

The last_result non-zero check catches the common case: the task fired, the script inside exited with code 1, and nobody noticed because the scheduler doesn't care. 0x80070002 (file not found) and 0x80070005 (access denied) are the two you will see most often when someone moves a script or changes a service account.

Note the running-state short-circuit. A long-running task that takes 90 minutes will otherwise trip the staleness check on every audit pass, because LastRunTime only updates when a run starts. If your task runs longer than your max_age, you need to either raise max_age or check State == 4 before evaluating staleness — which is what the code above does.

Walking the folder tree

Tasks live in folders, and GetTasks only returns the tasks directly in the folder you asked for. If your tasks are organized under \Quant\ and \, you need to recurse.

def walk(service, folder_path="\\"):
    folder = service.GetFolder(folder_path)
    for task in folder.GetTasks(TASK_ENUM_HIDDEN):
        yield task
    for sub in folder.GetFolders(0):
        yield from walk(service, sub.Path)
Enter fullscreen mode Exit fullscreen mode

GetFolders(0) returns immediate subfolders only; the recursion handles the rest. The 0 flag is a reserved bitmask that must be zero — passing anything else raises.

Matching tasks to expectations

Use fnmatch on the full Path, not the bare Name. Two folders can hold tasks with the same name, and matching on Name alone will silently pick the wrong one.

def audit(service, expectations, now):
    findings = []
    seen = set()

    for task in walk(service):
        for exp in expectations:
            if fnmatch.fnmatch(task.Path, exp.pattern):
                severity, reason = classify(task, exp, now)
                findings.append({
                    "path": task.Path,
                    "severity": severity,
                    "reason": reason,
                    "result": task.LastTaskResult,
                    "last_run": task.LastRunTime,
                })
                seen.add(exp.pattern)
                break

    # Expectations that matched nothing are themselves a finding.
    for exp in expectations:
        if exp.pattern not in seen:
            findings.append({
                "path": exp.pattern,
                "severity": "fail",
                "reason": "no matching task found",
                "result": None,
                "last_run": None,
            })

    return findings
Enter fullscreen mode Exit fullscreen mode

The unmatched-expectation check is the one people skip, and it's the most valuable. A task that was deleted, renamed, or moved to a different folder disappears from the enumeration entirely — a naive audit reports "all clear" because there is nothing to complain about. The expectation table is what makes absence detectable.

The alert

PushPlus takes a token and a message body over HTTP POST and delivers to WeChat. Keep the alert short: a subject line that says how many tasks failed, and a body that lists only the failures. Nobody reads a wall of green checkmarks.

import requests

PUSHPLUS_TOKEN = "your-token-here"

def format_alert(findings):
    bad = [f for f in findings if f["severity"] != "ok"]
    if not bad:
        return None
    lines = [f"{len(bad)} task(s) need attention", ""]
    for f in bad:
        lines.append(f"{f['path']}")
        lines.append(f"  {f['reason']}")
    return "\n".join(lines)

def push(message):
    resp = requests.post(
        "https://www.pushplus.plus/send",
        json={
            "token": PUSHPLUS_TOKEN,
            "title": "Task fleet alert",
            "content": message,
            "template": "txt",
        },
        timeout=15,
    )
    resp.raise_for_status()
Enter fullscreen mode Exit fullscreen mode

Send only when there is something to say. An alert channel that fires every hour with "all fine" trains you to swipe past it, and then the one real alert gets swiped too.

Wiring it together

def main():
    service = connect()
    now = datetime.now()
    findings = audit(service, EXPECTED, now)

    for f in findings:
        print(f"{f['severity']:4} {f['path']:40} {f['reason']}")

    message = format_alert(findings)
    if message:
        push(message)

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run it manually once, fix whatever it reports, then register it as a scheduled task. A 15-minute cadence is fine — the audit is cheap, and a shorter interval means you catch a failure closer to when it happened.

Why the audit must itself be scheduled

The obvious objection: "I'll just run this when I think about it." That is the same failure mode the audit exists to catch. If you remember to check, you don't need the check; if you forget, the check never runs. The only configuration where the audit has value is one where it runs without you.

That means the audit task is itself in the fleet, and it needs to be in the expectation table — a task that monitors nine tasks and is monitored by nobody is a single point of failure. Point an expectation at it with a max_age slightly larger than its own interval:

Expectation("\\Quant\\Audit_TaskFleet", timedelta(minutes=45), description="this audit"),
Enter fullscreen mode Exit fullscreen mode

If the audit task dies, you stop receiving alerts — and the way you find out is that you stop receiving alerts, which is exactly the signal you would miss. There is no fully self-healing answer here, but you can reduce the blast radius: have the audit write a heartbeat row to a table or a file, and check that heartbeat from a second, independent mechanism (a cron job on a different host, a cloud function, anything that doesn't share the same scheduler). Two schedulers watching each other is more resilient than one scheduler watching itself.

What to watch over time

Three signals are worth tracking beyond pass/fail:

  • Run duration drift. A task that used to finish in 40 seconds and now takes 12 minutes is heading for a timeout. Log LastRunTime deltas and alert on a sustained increase, not a single spike.
  • Result code distribution. A task that alternates between 0 and 0x80070002 is flapping. Count non-zero results over a rolling window rather than checking only the most recent.
  • Stale-but-zero. The nastiest case: exit code 0, last run within tolerance, but the task is doing nothing because its input file is empty or its upstream dependency silently stopped producing. The scheduler can't see inside the process. If it matters, have the task itself emit a completion marker with a row count or a timestamp, and audit that marker instead of trusting the exit code.

The audit script is the first layer. It catches the failures the scheduler can see. The second layer — output validation — catches the ones it can't, and it belongs inside the task, not in the audit.

More notes like this ship every week on this site.


Daily Picks

The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-06 12:36:33

Long

Pair Signal Price Take Profit Stop Loss R/R
SKYAI $0.0416 $0.0433 $0.0406 1:1.6
RE $0.4991 $0.5181 $0.4866 1:1.5

2 picks selected. Scanner runs every 15 minutes.

Top comments (0)