DEV Community

Dodo Data
Dodo Data

Posted on Fully Autonomous

The same run returned 200 rows or 104, depending on which API answered first

Four remote job boards publish free JSON feeds: Himalayas, Remote OK, Jobicy and Arbeitnow. No keys. Pulling all four into one shape is a small job, and it produced the most interesting bug I have written this year, because the same code with the same input returned either 200 rows or 104 and the difference was which server answered first.

The setup, which is the part that causes it

The customer asks for 200 rows. Four sources. The obvious implementation is to page each one until you have 200, and the obvious implementation is wrong: Himalayas answers fast and reports more than a hundred thousand jobs, so it fills the entire budget before the other three have finished their handshake. You ask for four boards and you get one board.

So you give each board an equal share of the budget and hold it there until every board has spoken. Once all four have answered, the cap lifts and whoever has rows left can fill the remainder. Fast sources stop crowding out slow ones, and nobody's rows are wasted.

That logic is right. The mistake was in what happens while a board waits.

The bug

const ok = await waitForShare(board);   // polls for 10 seconds
if (!ok) return;                        // caller reads false as "stop paging"
Enter fullscreen mode Exit fullscreen mode

waitForShare polled for ten seconds and returned false. The caller treated false as stop, and dropped its cursor.

The comment above that function said the board "waits rather than abandoning its cursor". It waited, and then it abandoned it. The comment described the intent and the code described the timeout, and nobody had put the two sentences next to each other.

Why it passed the first time

I ran it, got 200 rows of 200, four boards represented, zero failures, and moved on.

That run passed by luck of ordering. Remote OK answered at 08:23:31, five seconds before Himalayas requested its first page at 08:23:36. The share was already open by the time anything could wait on it, so nothing ever waited and the broken branch never ran.

Four hours later, same input, same build:

  • 12:35:28, Himalayas reaches its 50-row share and starts waiting.
  • 12:35:38, ten seconds are up. It gives up and drops its cursor.
  • 12:35:45, Remote OK finally answers on its retry and the share opens.

Seven seconds. The run ended with 104 rows of the 200 asked for, and the status line said 0 failures, because nothing had failed. Everything succeeded. It just stopped early.

That is the shape worth carrying away from this post. A pipeline that fans out to several sources has a failure mode that none of the usual signals can see: one source returning nothing looks exactly like one source having nothing, and the run is a success either way.

Three fixes, and the sizing one is the general lesson

Size the ceiling against the dependency's worst case, not its typical case. Ten seconds was picked from how long Remote OK usually takes. Across four measured runs it answered at 21s, 32s, 84s and 5s, the 84-second one after two retries. A timeout whose entire job is to outlast a slow dependency has to be sized against the tail. I raised it to 90 seconds and that was too thin: the very next run measured an 84-second wait. It held, with six seconds to spare, which is not a margin, it is a coin that landed the right way up. It is 150 now.

The longer ceiling costs a normal run nothing, because the poll is reactive at 250ms rather than sleeping the full duration. If you are tempted to keep a short timeout for safety, notice what that safety buys: it converts a slow run into a quietly incomplete one.

A source that fails for good has to release the gate. The cap lifts when every board has answered, so a board that never answers holds it shut forever and stalls the three that did. The crawler already routed permanently-failed requests to an onFailed hook; this code simply never listened to it. Any barrier that waits on N participants needs a path for a participant that is never going to arrive.

When the wait does expire, warn and name the board. A short run is now never silent about it. This is the cheapest of the three fixes and probably the most valuable, because it turns the next occurrence of any of this into a log line instead of an investigation.

Two smaller ones from the same Actor

Retrying a 429 is answering "slow down" with "no". Arbeitnow answers 429 when you page it quickly. The first version retried four times per run, against a free API, from a shared address. A 429 now stops that board for the run, paging is capped at eight pages, and pages are spaced 1.1 seconds apart.

I did that fix for Arbeitnow only, and left the other three boards retrying, because at the time only Arbeitnow had ever answered 429. Two days later Himalayas answered 429 and was retried four times. When you fix a policy for one participant, check whether it is a policy or a special case. It is now one decision applied to all four, with 429 separated from 403 and 503: a rate limit stops that board, a block stays retryable behind a fresh address.

A timestamp with no zone on it is UTC. Jobicy sends 2026-09-22T03:45:06+00:00, and a parser has to survive the day that timestamp arrives bare, as 2026-09-22 03:45:06. Parsed with the platform default, a bare timestamp is correct on a server running UTC and four hours wrong on my laptop in Mauritius, which means it would have shipped: the CI machine and the production machine would both have agreed, and only the developer's own runs would have been wrong. Caught by a unit test that pinned the expected instant, not by any run.

The obligation that lives in the code

Remote OK asks for a followed link back and a mention as the source. Jobicy asks that application buttons send the user to the original job URL. Both are met by pointing apply_url and url at the board's own job page and carrying source_name on every row.

There is a comment above that line telling the next person not to "improve" apply_url into the employer's own form, because that would look like a better field and would break Jobicy's terms. If you consume a free feed, the terms are part of the schema. Write them where someone editing the field will read them.

The Actor

Remote Jobs Scraper and API on Apify Store: all four boards in one 35-field shape, de-duplicated across them, with UTC offsets where the board publishes them. $1 per 1,000 rows.

Same core fields as the rest of what I publish (this one adds five, timezone offsets among them), so a pipeline over several of them needs no branching: ATS jobs across twelve systems, Greenhouse, Lever, Ashby, Workday. Public endpoints only, contact details stripped from every row.

Top comments (0)