DEV Community

Hugo Jose
Hugo Jose

Posted on Originally published at hugoj0s3.dev

JobMaster achieving horizontal scaling for .NET jobs while maintaining a centralised audit log

The Problem

People usually like me like to use a job scheduler for background work because it keeps an audit trail. You can see what happened to job XYZ and reschedule it by hand if you need to.

The problem is it doesn't scale well. Every worker ends up fighting for the same table, and you get lock contention.

The other option is a message broker. That scales. But you lose the useful parts: retries, recurring schedules, and the audit trail.

Brainstorming the Solution

With that issue, I started to think of a way to figure out this dilemma and have a solution that we can audit a job and, at the same time, have high performance. The first thing that came to mind was: why not separate the auditing database from the execution pool? Why do not we divide the jobs into partitions and have a particular worker that owns this partition? Then I started to think about it. It fixes the lock contention problem, but we still need to store the job in a central table. Also, if I schedule a job for a few days, will I provision it into a partition and wait that long? At this point, I was jogging and thinking about how to architect this.

The Bucket System Solution

I started thinking about storing the job in a partition for execution, while asynchronously recording it in the audit database.

So here's what I landed on: there's an 'Agent' layer that handles execution and moving jobs around, and a Master DB that keeps track of the audit history and coordinates everything. I call it an 'Agent DB,' but honestly, it doesn't even have to be a database. It could be a message broker like NATS, or something else entirely. You can run a bunch of Agent connections in the same cluster if you want. The Master DB is where the long-term memory lives. The Agent layer is just for jobs that are about to run right now.

The bucket system lets us write to the Master DB asynchronously, which helps with throughput, since the Agent database only holds short-lived jobs. But then I started wondering: what if I schedule a job for next year? I can't just let it sit in a bucket for twelve months, right?

So here's what I came up with: always write the job into a bucket first, then save it to the Master DB in the background. That way, nothing has to wait for a slow database round trip. If the job is scheduled way in the future, it gets kicked out of the bucket once it's safely in the Master DB, and just hangs out there until it's almost time to run. When the execution time is close, it moves back into a bucket so a worker can pick it up. Buckets aren't a long-term home for jobs. They're just the front door and the last stop before execution. Everything else lives in the Master DB.

Sure, there's still some lock contention on the Master DB, but it's a different beast now. Moving jobs back into buckets as their execution time gets close happens in big batches, not one job at a time fighting for the same rows. And this only matters for jobs scheduled way ahead. Most jobs just zip through a bucket and never need to take the scenic route. So the Master DB only gets busy if you're scheduling lots of jobs far in advance, not just because you have a ton of jobs overall.

The Architecture

Let me get straight to the point and walk you through the architecture.

Save Process

Every job is written into a bucket immediately, before checking when it's actually due to execute. The TransientThreshold configuration then decides whether the job stays in that bucket for near-term execution or gets moved to wait on the Master DB instead. By default, it's 10 minutes, but you can tweak it to suit your workload.

Assign to Buckets Mechanism

When a job's execution time is almost here, a Coordinator grabs it from the Master DB and puts it into the right bucket based on its priority and which worker lane it belongs to. Each worker has its own set of buckets, one for each priority. So if a job goes to Worker A's Critical bucket, it doesn't have to fight with jobs on Worker B or C. And once a job is assigned, it's exclusive. No two workers can grab the same job.

If a job fails, it doesn't just get retried where it is. Instead, it goes back to the Master DB, waits there, and then gets sent to a bucket again when it's time for another try, up to whatever retry limit you set. No matter what happens, whether the job succeeds or fails, the result goes back to the Master DB. That way, you always have the full history in one place, no matter which worker actually ran the job.

Try It

JobMaster is open source and still in alpha. The integration tests are pretty solid, but I'd love to see it survive some long, production-like runs before I call it truly battle-tested.

If you want to look under the hood, the code is on GitHub: https://github.com/hugoj0s3/jobmaster-net

Docs: https://docs.jobmaster.hugoj0s3.dev/

Dashboard sandbox: https://sandbox.jobmaster.hugoj0s3.dev/jm-dashboard

I'd really appreciate any feedback, especially if you've hit the same scaling wall with Hangfire or Quartz.NET and found your own creative way around it.

Final Thoughts

JobMaster is still a work in progress, born out of real scaling headaches and my wish for better visibility into distributed jobs. The bucket system is what I'm trying right now, but I'm sure there are other ways to solve this problem.

If you're working on similar problems or just want to help make large-scale job processing less mysterious and more reliable, I'd love to hear from you. Questions, feedback, and contributions are all welcome.

Top comments (2)

Collapse
 
iqtechsolutions profile image
Ivan Rossouw •

The boundary I’d make explicit is the acknowledgement between the bucket and the Master DB. If the bucket is ephemeral and the caller receives success before the audit record is durable, a crash can leave accepted or even executed work without an authoritative record; retries can create the inverse risk of duplicate dispatch. What durable write permits the client acknowledgement, which idempotency key survives rehydration, and how is bucket/master disagreement reconciled? A failure-injection test that kills the agent between those writes would reveal more than a steady-state throughput benchmark.

Collapse
 
hugo_jose_9 profile image
Hugo Jose • • Edited

First, thanks for the Question; I really appreciate that!

The client ack is the agent write (PendingSave), not the Master write. That’s the tradeoff for keeping the audit store off the scheduling hot path.

The identity that survives rehydration is the job GUID. Same ID hitting Master again is treated as AlreadyExists (insert-first / upsert on drain). There is no separate client idempotency key. A caller retry of the schedule call creates a new Guid and a second job.

What happens on crash:

  • Worker dies, agent storage still there: another worker assumes the bucket. Drain runners take PendingSave jobs and upsert them to Master as HeldOnMaster. Master becomes the source of truth again, then the job will be assigned to another bucket for execution.

  • Agent storage is gone forever: that accepted job is lost. Same class of risk as a broker that acked and then lost.
    but at least we will have log that agent storage gone.

  • Master is down after the agent accepted: the job stays PendingSave and the polling runner retries until Master is writable. It should not execute as a normal job until that persist happens. If Master never comes back, it is neither audited nor executed.

Two buckets claiming the same job should fail the version check when processing starts. Rare, but possible. That attempt shows up in the execution log. Still at-least-once, not exactly-once.

So: Guid is the reconcile key between bucket and Master. If Master already has that id, the write conflicts (duplicate or version mismatch). We log it and ignore. The job is already there; this copy isn't new. Still at-least-once, not exactly-once.

I already have scenario tests that kill workers and check jobs are not lost after bucket assumption. What’s missing now is a clear failure-injection case: kill the agent between the PendingSave ack and the Master persist, then show the drain/reconcile outcome. I’ll make that explicit.