DEV Community

mattleeee
mattleeee

Posted on Originally published at hkcode.dpdns.org

Why I Replaced Daemon Processes with Windows Task Scheduler

Corporate Windows machines are hostile territory for long-running processes. You write a daemon, it runs fine on your dev box, then you deploy it to a managed laptop and it dies at some random point in the night with no traceback and no Windows Event Log entry. I spent a few weeks chasing this before giving up on the daemon pattern entirely and moving everything to Windows Task Scheduler. Here's what happened and how the migration looks in practice.

The Symptom: A Daemon That Dies Quietly

The first version of my job runner was a classic Python daemon: a parent process that spawned child workers, a watchdog thread that restarted workers if they crashed, and a log file that recorded every restart. On my own machine it ran for weeks.

On the corporate laptop, the log looked like this:

2024-03-11 09:14:02 INFO  worker started (pid 18432)
2024-03-11 09:14:02 INFO  worker pid 18432 exited with code 0
2024-03-11 09:14:02 INFO  restarting worker...
2024-03-11 09:14:02 INFO  worker started (pid 19004)
2024-03-11 09:14:02 INFO  worker pid 19004 exited with code 0
2024-03-11 09:14:02 INFO  restarting worker...
Enter fullscreen mode Exit fullscreen mode

A tight restart loop with no error. The child processes were being created and immediately killed. No Python exception, because the kill happened at the OS level before the interpreter could raise anything.

The culprit was Sangfor aTrust, the endpoint security suite that ships with a lot of enterprise environments. It hooks process creation and, depending on policy, blocks or silently terminates child processes that don't match an allowlist. A daemon that spawns workers looks exactly like the kind of behavior a security product is built to stop. The watchdog made it worse: every restart attempt was another blocked CreateProcess call, which is why the loop was so tight.

You can confirm this pattern quickly. Run a trivial spawner and see whether the child ever gets to execute:

import subprocess
import sys

# If this prints nothing, the child never ran.
# That is a security-policy block, not a Python bug.
proc = subprocess.Popen(
    [sys.executable, "-c", "print('child alive')"],
    stdout=subprocess.PIPE,
    stderr=subprocess.PIPE,
)
out, err = proc.communicate(timeout=10)
print("returncode:", proc.returncode)
print("stdout:", out)
print("stderr:", err)
Enter fullscreen mode Exit fullscreen mode

If returncode is non-zero and both streams are empty, the child was killed before it could write anything. That's the signature.

Why Task Scheduler Survives Where Daemons Don't

The key insight is that Task Scheduler is not a user process. It's a Windows service (Schedule) running in session 0, and the processes it launches are created by the service host, not by your Python interpreter. Sangfor aTrust and similar suites generally allow this path because:

  1. The parent is a trusted system service. Process-creation hooks are usually scoped by parent identity. svchost.exe running the Schedule service is on the allowlist; python.exe spawning python.exe is not.
  2. Task Scheduler is a documented management surface. Enterprise security products are designed not to break the OS's own scheduling. Blocking it would break Windows Update, backup agents, and half the management stack. So it stays open.
  3. No persistent parent process. There's nothing long-running to kill. Each task run is a short-lived process, launched, executed, and reaped by the service. The "daemon" is the scheduler itself, which you don't have to keep alive.

That last point is the real architectural win. Instead of building a process that stays up and supervises itself, you hand the supervision to the OS. The OS is very good at it and the security suite leaves it alone.

Mapping Jobs to Tasks

Once I accepted the model, the migration was mechanical. Each recurring job becomes one scheduled task with four things defined: a trigger, a working directory, a timeout, and a result convention.

Trigger

Triggers replace your while True: sleep(...) loop. For a job that runs every five minutes, use a daily trigger with a repetition interval. For a job that runs once at 06:00 on weekdays, use a weekly trigger with the day set. The scheduler handles the clock; you don't need time.sleep at all.

Working Directory

This is the most common source of "it works in my terminal but not in the task." Task Scheduler defaults the working directory to C:\Windows\System32. Any relative path in your script breaks. Always set an explicit working directory, and make your script resolve its own paths rather than relying on os.getcwd():

from pathlib import Path

# Anchor everything to the script location, not the process cwd.
BASE_DIR = Path(__file__).resolve().parent
LOG_DIR = BASE_DIR / "logs"
STATE_DIR = BASE_DIR / "state"
LOG_DIR.mkdir(exist_ok=True)
STATE_DIR.mkdir(exist_ok=True)
Enter fullscreen mode Exit fullscreen mode

Timeout

Task Scheduler has a per-task execution time limit. Set it. A hung job that holds a database connection or a lock file forever is worse than a failed job. For a job expected to run in under two minutes, set the limit to ten. The scheduler terminates the process and records a failure, which you can see.

Result Convention

The scheduler records the process exit code. Use it. A script that always exits 0 is invisible to the scheduler's history. Exit non-zero on failure so the last-run result is meaningful:

import sys
import logging

logging.basicConfig(
    filename=str(LOG_DIR / "job.log"),
    level=logging.INFO,
    format="%(asctime)s %(levelname)s %(message)s",
)

def main() -> int:
    try:
        run_job()
    except Exception:
        logging.exception("job failed")
        return 1
    return 0

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Creating Tasks from Python

You can register tasks by hand in the GUI, but that doesn't scale and it isn't reviewable. I keep a single setup_tasks.py that registers every job idempotently. It shells out to schtasks.exe, which is present on every Windows install and needs no extra dependency.

import subprocess
from pathlib import Path

PYTHON = r"C:\Python311\python.exe"
BASE = Path(r"C:\jobs")

TASKS = [
    {
        "name": "QuantJob_Fetch",
        "script": BASE / "fetch.py",
        "schedule": "MINUTE /MO 5",
        "timeout_min": 10,
    },
    {
        "name": "QuantJob_Report",
        "script": BASE / "report.py",
        "schedule": "DAILY /ST 06:00",
        "timeout_min": 30,
    },
]

def register(task: dict) -> None:
    cmd = [
        "schtasks", "/Create",
        "/TN", task["name"],
        "/TR", f'"{PYTHON}" "{task["script"]}"',
        "/SC", *task["schedule"].split(),
        "/RL", "LIMITED",          # run with normal user rights
        "/F",                       # overwrite if it exists
    ]
    subprocess.run(cmd, check=True, capture_output=True, text=True)

    # Set working directory and execution time limit separately.
    # schtasks has no direct flag for cwd, so use XML for full control.
    set_task_xml(task)

def set_task_xml(task: dict) -> None:
    # Export, patch, re-import. This is the reliable way to set
    # WorkingDirectory and ExecutionTimeLimit.
    name = task["name"]
    xml = subprocess.run(
        ["schtasks", "/Query", "/TN", name, "/XML"],
        check=True, capture_output=True, text=True,
    ).stdout

    xml = xml.replace(
        "<WorkingDirectory></WorkingDirectory>",
        f"<WorkingDirectory>{BASE}</WorkingDirectory>",
    )
    xml = xml.replace(
        "<ExecutionTimeLimit>PT72H</ExecutionTimeLimit>",
        f"<ExecutionTimeLimit>PT{task['timeout_min']}M</ExecutionTimeLimit>",
    )
    # StartWhenAvailable: if the machine was off at trigger time,
    # run the task as soon as it comes back.
    if "<StartWhenAvailable>false</StartWhenAvailable>" in xml:
        xml = xml.replace(
            "<StartWhenAvailable>false</StartWhenAvailable>",
            "<StartWhenAvailable>true</StartWhenAvailable>",
        )

    subprocess.run(
        ["schtasks", "/Create", "/TN", name, "/XML", "-", "/F"],
        input=xml, check=True, capture_output=True, text=True,
    )

if __name__ == "__main__":
    for t in TASKS:
        register(t)
        print(f"registered {t['name']}")
Enter fullscreen mode Exit fullscreen mode

The XML round-trip is the part people skip and then wonder why their working directory is System32. schtasks doesn't expose /WorkingDirectory or /ExecutionTimeLimit as command-line flags, so you export the task, edit the XML, and re-import. It's ugly but it's stable and it's scriptable.

Operational Wins You Get for Free

The migration wasn't just about surviving the security suite. Three things got better immediately.

Reboot survival

A daemon dies on reboot and you need a service wrapper, a startup shortcut, or a login script to bring it back. A scheduled task just runs on its trigger. If the trigger is a repetition interval, the task resumes on the next interval after boot. If you need it to catch up, <StartWhenAvailable>true</StartWhenAvailable> makes the scheduler run a missed task as soon as the machine is back. That single flag replaced an entire startup-recovery code path in my old daemon.

Built-in retry

Task Scheduler has a native retry policy. If a task fails, you can configure it to retry every N minutes for M attempts. That's the watchdog pattern, implemented by the OS, and it doesn't spawn child processes from your interpreter so the security suite has nothing to object to. You set it in the task's Settings tab, or in XML:

<RestartOnFailure>
  <Interval>PT5M</Interval>
  <Count>3</Count>
</RestartOnFailure>
Enter fullscreen mode Exit fullscreen mode

Last-run result codes

Every task records its last run time and result code. You can query them without any custom instrumentation:

import subprocess
import csv
from io import StringIO

def last_results() -> list[dict]:
    out = subprocess.run(
        ["schtasks", "/Query", "/FO", "CSV", "/V"],
        check=True, capture_output=True, text=True,
    ).stdout
    rows = list(csv.DictReader(StringIO(out)))
    return [
        {
            "name": r["TaskName"],
            "status": r["Status"],
            "last_run": r["Last Run Time"],
            "result": r["Last Result"],
        }
        for r in rows
        if r["TaskName"].startswith("\\QuantJob_")
    ]

for row in last_results():
    print(row)
Enter fullscreen mode Exit fullscreen mode

Last Result is the process exit code. 0 means success, anything else means the job failed and you can look at your log file. That's a free health dashboard for every job on the machine. I pipe this into a small status page and I haven't written a single line of process-monitoring code since.

What I'd Do Differently

Two things. First, I'd set explicit working directories from day one instead of debugging FileNotFoundError for an afternoon. Second, I'd design every job to be idempotent and short-lived. The scheduler model rewards jobs that start, do one thing, and exit. If your job needs to hold state between runs, write it to a file or a database rather than keeping it in memory. The process is gone when it's done; that's the point.

The daemon pattern isn't wrong in general. It's wrong on a locked-down corporate machine where you don't control the process-creation policy. Task Scheduler is the platform's own answer to "run this on a schedule," and in that environment, using the platform's answer is the shortest path to something that actually stays running.

More notes like this ship every week on this site.


Daily Picks

The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-06 12:36:33

Long

Pair Signal Price Take Profit Stop Loss R/R
SKYAI $0.0416 $0.0433 $0.0406 1:1.6
RE $0.4991 $0.5181 $0.4866 1:1.5

2 picks selected. Scanner runs every 15 minutes.

Top comments (0)