Picture a job that has to pull data from 25 different APIs every night, process the responses, and store the results for reporting. Nothing about it is hard. It's just a lot of small things that all have to happen, on time, without anyone watching.
Most teams build it the obvious way: a script that calls API one, then API two, then API three, all the way to 25. It works. It's also the slowest possible way to do it.
Most of the work doesn't depend on the rest
Those 25 API calls don't need each other. Call 14 doesn't care what call 3 returned. So there's no reason to make them wait in line.
Run independent calls in parallel and the total time drops from ""the sum of every call"" to roughly ""the slowest single call."" In our test, we modeled the job as five groups of five parallel HTTP steps, 25 calls in total. The whole thing finished in under 100 milliseconds.
That's a test against fast endpoints, so don't expect those numbers against slow third-party APIs. The point is the shape: if steps are independent, run them side by side.
Scheduling is where jobs quietly go wrong
Once the job works, you need it to run on its own. That's usually a cron expression, and the standard five-field syntax covers most needs:
0 * * * * # every hour, on the hour
0 2 * * 1 # every Monday at 2 AM
0 0 1 * * # first day of every month at midnight
Cron is the easy part. The part people forget is what happens when a run takes longer than expected.
The overlap problem
Say your job normally takes ten minutes, but one night an API is slow and it takes ninety. If the schedule fires again in the meantime, you now have two copies running at once. Two copies can write conflicting data, double-count records, or hammer the same APIs and make things slower.
The fix is an overlap policy. Setting the schedule to not allow overlap means a new run won't start while the previous one is still going. It's one setting, and it prevents a very annoying class of bugs.
You should be able to see what ran
A batch job that runs at 2 AM has one big weakness: nobody is awake to notice when something breaks. So it helps a lot to be able to open a list of past executions and see which ones ran, how long they took, and whether every step finished.
In our demo, the scheduled run triggered on time, ran all 25 steps, and completed in 55 milliseconds. More useful than the speed was that we could see all of it afterward.
The short version
For any batch job that fans out to many sources:
- Run independent steps in parallel instead of one after another
- Schedule with cron, and decide up front what should happen when a run overlaps the last one
- Make sure you can look back and see what actually happened
Get those three right and a nightly job becomes boring, which is exactly what you want from a nightly job."
Top comments (0)