A market-signal job ran at HH:00. A data-refresh job ran at HH:00. Both touched the same JSON files and both wanted the CPU at the same instant. The result was predictable: PermissionError on Windows file locks, occasional half-written JSON that the signal job then parsed as empty, and a scheduler log full of overlapping runs.
The fix wasn't a faster machine or a rewrite. It was a timetable. This post walks through how I inventoried the triggers, decided on offsets, handled the shared-state file locks, and ended up with a single page that shows the whole fleet.
Why top-of-hour schedules collide
Cron-style schedulers are seductive because 0 * * * * reads like "do the thing." It doesn't say "do the thing when nothing else is doing the thing." Every job author independently picks round numbers, and round numbers cluster. If nine jobs all live on 0 * * * * or 0 4 * * *, they share a trigger boundary whether they need to or not.
Two failure modes show up immediately:
- CPU contention. Two CPU-bound jobs start together, each takes roughly twice as long as it would alone, and the overlap window grows.
-
File-lock collisions. On Windows, opening a file for write while another process holds it raises
PermissionErroror silently truncates depending on the flags. Two writers to the same JSON file is a data-corruption bug waiting for a slow disk.
The signal job and the refresh job hit both. The signal job read signals.json and wrote positions.json. The refresh job wrote signals.json. When they ran together, the signal job sometimes read an empty file mid-write and produced an empty position list — which then propagated downstream.
Step 1: Inventory every trigger
You can't stagger what you haven't listed. I dumped every scheduled entry into one place. On Windows, the Task Scheduler is scriptable through schtasks or the COM interface. The COM path gives structured data without parsing text.
import win32com.client
from datetime import datetime
def list_scheduled_tasks(folder_path="\\"):
"""Walk a Task Scheduler folder and return trigger summaries."""
scheduler = win32com.client.Dispatch("Schedule.Service")
scheduler.Connect()
root = scheduler.GetFolder(folder_path)
tasks = root.GetTasks(1) # 1 = include hidden
rows = []
for task in tasks:
definition = task.Definition
for trigger in definition.Triggers:
rows.append({
"name": task.Name,
"enabled": task.Enabled,
"type": trigger.Type,
"start": getattr(trigger, "StartBoundary", None),
"repetition": getattr(trigger, "Repetition", None),
})
return rows
for row in list_scheduled_tasks():
print(f"{row['name']:<40} {row['start']}")
That gives you names and start boundaries. It doesn't tell you what the jobs do — which is the part that matters for spacing. So I added a second pass: for each task, record the action (the script path) and a manual tag for CPU-bound vs I/O-bound.
TAGS = {
"MarketSignalHourly": {"kind": "cpu", "reads": ["signals.json"],
"writes": ["positions.json"]},
"DataRefreshHourly": {"kind": "io", "reads": [],
"writes": ["signals.json"]},
"NightlyMaintenance": {"kind": "cpu", "reads": ["positions.json"],
"writes": ["positions.json", "audit.log"]},
"BlogPublish": {"kind": "io", "reads": ["posts/*.md"],
"writes": ["site/"]},
"BlogPublishDE": {"kind": "io", "reads": ["posts/de/*.md"],
"writes": ["site/de/"]},
}
The reads/writes sets are the important column. Two jobs that share a write target must never overlap, regardless of how cheap they are.
Step 2: Rules of thumb for spacing
Once you have the inventory, spacing is mostly mechanical. Three rules cover almost everything:
CPU-bound jobs need a full estimated runtime of separation
If a job takes 12 minutes of CPU on a busy box, don't schedule the next CPU-bound job 5 minutes later. Measure the p95 runtime, not the median. The overlap you're trying to avoid is the tail, not the typical case.
I/O-bound jobs can share a slot if they touch different files
Two jobs that read disjoint file sets and write to disjoint outputs can run together. Disk throughput is usually fine for a handful of readers. The moment they write the same path, they can't share.
Never let a writer and a reader of the same file overlap
This is the hard rule. Even a fast reader can catch a partial write. Either serialize them or make the write atomic (see the next section).
Applying these to the fleet:
| Job | Kind | Runtime (p95) | Constraint |
|---|---|---|---|
| MarketSignalHourly | cpu | 14 min | writes positions.json
|
| DataRefreshHourly | io | 6 min | writes signals.json
|
| NightlyMaintenance | cpu | 40 min | writes positions.json, audit.log
|
| BlogPublish | io | 3 min | writes site/
|
| BlogPublishDE | io | 3 min | writes site/de/
|
The signal job and the refresh job both ran at HH:00. The signal job reads signals.json; the refresh job writes it. That's the writer/reader overlap, and it's the one that corrupted data. The fix: move the refresh job off the top of the hour so it finishes before the signal job starts, or at least so the write window doesn't intersect the read window.
I moved the refresh to HH:45 — 15 minutes before the signal job. The refresh takes 6 minutes at p95, so it's done by HH:51, leaving a 9-minute buffer. That buffer absorbs the tail.
Step 3: The offset schedule
Here's the timetable that emerged. The signal job stays at HH:00 because downstream consumers expect it there. Everything else moves around it.
HH:00 MarketSignalHourly (cpu, 14 min)
HH:45 DataRefreshHourly (io, 6 min) -> completes before next HH:00
04:30 NightlyMaintenance (cpu, 40 min) -> done by 05:10
09:00 BlogPublish (io, 3 min)
10:30 BlogPublishDE (io, 3 min)
The daily jobs sit in a quiet window. 04:30 is after the overnight batch and before the morning signal jobs ramp up. 09:00 and 10:30 are separated by 90 minutes, which is overkill for two 3-minute jobs, but the gap gives room for retries and for the English site build to finish before the German one starts — they share a site/ parent directory, so a partial build on one side can confuse a deploy that scans the whole tree.
The HH:45 offset is the one that matters most. It's not a round number, which is the point. Round numbers cluster; odd offsets don't.
Step 4: Handle the shared state file
Staggering reduces overlap but doesn't eliminate it. A job that overruns its window will still collide with the next one. The durable fix is atomic writes plus a lock file.
On Windows, os.replace is atomic on the same volume. Write to a temp file, then replace the target. Readers either see the old file or the new one, never a partial write.
import json
import os
import tempfile
from pathlib import Path
def atomic_write_json(path: Path, payload: dict) -> None:
"""Write JSON so readers never observe a partial file."""
path = Path(path)
fd, tmp_name = tempfile.mkstemp(
dir=path.parent, prefix=path.name + ".", suffix=".tmp"
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as fh:
json.dump(payload, fh)
fh.flush()
os.fsync(fh.fileno())
os.replace(tmp_name, path) # atomic on same volume
except BaseException:
os.unlink(tmp_name)
raise
That handles readers. For writers, add a lock so two processes can't both decide to write. A simple lock file with an exclusive create works on Windows and POSIX:
import time
from contextlib import contextmanager
@contextmanager
def file_lock(path: Path, timeout: float = 30.0):
"""Exclusive lock via atomic create. Windows-safe."""
lock = Path(str(path) + ".lock")
deadline = time.monotonic() + timeout
while True:
try:
fd = os.open(lock, os.O_CREAT | os.O_EXCL | os.O_WRONLY)
os.close(fd)
break
except FileExistsError:
if time.monotonic() > deadline:
raise TimeoutError(f"lock held: {lock}")
time.sleep(0.5)
try:
yield
finally:
lock.unlink(missing_ok=True)
Then any job that writes signals.json wraps its write:
with file_lock(SIGNALS_PATH):
data = load_json(SIGNALS_PATH)
data.update(new_signals)
atomic_write_json(SIGNALS_PATH, data)
The lock serializes writers; the atomic replace protects readers. Together they make the schedule forgiving of small overruns. Staggering is still the first line of defense — locks are the second.
One caveat: a stale lock file from a crashed process will block everything. Either add a PID check or a timestamp inside the lock file and expire it. For a small fleet, a timestamp is enough:
def lock_is_stale(lock: Path, max_age: float = 300.0) -> bool:
try:
age = time.time() - lock.stat().st_mtime
except FileNotFoundError:
return False
return age > max_age
Check lock_is_stale before raising TimeoutError, and unlink if it's old.
Step 5: One page for the whole fleet
The inventory script and the tags table are useful, but they drift. The thing that actually kept the schedule sane was a single generated page — a timetable that lists every job, its trigger, its kind, and its file dependencies. I regenerate it from the source of truth (the task definitions plus the tags dict) and commit it.
def render_timetable(tasks, tags) -> str:
lines = ["| Job | Trigger | Kind | Reads | Writes |", "|---|---|---|---|---|"]
for task in sorted(tasks, key=lambda t: t["start"]):
tag = tags.get(task["name"], {})
lines.append(
f"| {task['name']} | {task['start']} | {tag.get('kind', '?')} "
f"| {', '.join(tag.get('reads', [])) or '-'} "
f"| {', '.join(tag.get('writes', [])) or '-'} |"
)
return "\n".join(lines)
The value isn't the table itself — it's that adding a job forces you to fill in the reads/writes columns. That's the moment you notice two jobs want the same file at the same time. The timetable makes the collision visible before it happens, instead of after a corrupted JSON file shows up in the logs.
What changed
After the rebalance, the PermissionError entries stopped. The signal job's parse failures went to zero because signals.json is now written by a job that finishes 9 minutes before the signal job reads it, and even if it overruns, the atomic write means the signal job reads a complete file. The nightly maintenance window moved off the top of the hour, so it no longer competes with the first morning signal run.
None of this required new hardware. It required looking at the schedule as a system instead of a list of independent jobs, measuring the tail runtimes, and picking offsets that don't cluster on round numbers.
More notes like this ship every week on this site.
Daily Picks
The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-06 12:36:33
Long
| Pair | Signal | Price | Take Profit | Stop Loss | R/R |
|---|---|---|---|---|---|
| SKYAI | $0.0416 | $0.0433 | $0.0406 | 1:1.6 | |
| RE | $0.4991 | $0.5181 | $0.4866 | 1:1.5 |
2 picks selected. Scanner runs every 15 minutes.
Top comments (0)