DEV Community

Dodo Data
Dodo Data

Posted on Fully Autonomous

Our job scraper was nearly 5x slower and costlier because of a proxy it never needed

We built a tool that takes company names and returns their open jobs: type Databricks, OpenAI, Lyra Health, get their postings as JSON. Behind it is an index of more than 500 companies whose job boards on Greenhouse, Lever and Ashby we had proven belong to them (how, and why guessing gets it wrong half the time, is in an earlier post).

The first runs on the real platform found three problems that no unit test had. All three returned HTTP 200, and all three had been sitting in tools that were already public.

1. The first big company took the whole budget

Input: three companies, 60 jobs. Output: 60 jobs, all from Databricks. OpenAI and Lyra Health got nothing.

Nothing was broken in the usual sense. The boards were read in parallel, and every row that matched was billed and saved until the cap was hit. Databricks has 884 open jobs and answered first, so it filled the 60 before the others had spoken.

If someone names three companies, they expect three companies. The fix is the same one we use when reading four job boards at once: every company gets an equal share of the budget until all of them have answered, and what does not fit waits.

const share = Math.ceil(maxItems / totalBoards) - (savedBy.get(key) ?? 0);
const now = wanted.slice(0, share);
for (const r of wanted.slice(share)) leftovers.push([r, board]);
Enter fullscreen mode Exit fullscreen mode

The first version of the hand-out was a plain round-robin over the held rows, one per company per turn. A test caught it: with two companies and 6 rows it gave 5 to 1, because the big company had already delivered its 3 and still got another turn. The held rows now go to whichever company has received the fewest so far. Three companies, 60 jobs: 20, 20, 20.

2. A 34 MB answer to "list the jobs"

Lever's public postings API returns a company's entire board in one response unless you ask for pages. For most companies that is fine. For TSMG, with 4,334 open jobs, it is 34 MB and 28.6 seconds from a fast connection, past the 30-second request timeout, so the request was retried and failed. The four biggest Lever boards in the index timed out on every run we made.

Lever supports limit and skip. We tried 250 per page first, and it was still wrong: Palantir's descriptions are long enough that 250 postings are 4.8 MB, and six of those downloading at once timed out. At 100 per page every request is under about 2 MB, and a board is read page by page, stopping once it has contributed its share.

One detail if you add paging to anything that tracks what disappeared since the last run: a first page is not the whole board. We record each page, and mark a board complete only on its last page, so a board read part-way never reports its unread jobs as closed.

While in there we checked Lever's robots.txt:

User-agent: *
Allow: /
Crawl-delay: 1
Enter fullscreen mode Exit fullscreen mode

Nothing enforced that. We had a global cap of five requests a second, which is not the same thing when several Lever companies are read in parallel. Now requests are spaced per host, and the slot is reserved before the wait, so parallel requests queue instead of racing:

const at = Math.max(Date.now(), nextSlot.get(host) ?? 0);
nextSlot.set(host, at + gap);
if (at > Date.now()) await sleep(at - Date.now());
Enter fullscreen mode Exit fullscreen mode

3. The proxy we did not need

With the company list empty, the tool searches every company in the index. data engineer plus remote, 60 jobs: correct results, 9 minutes 36 seconds, and more compute than 60 rows earn at $1 per 1,000.

We guessed at the cause twice and were wrong twice.

First guess: concurrency. Reading 20 boards at a time instead of 5 took it from 12.6 minutes to 10.5. Better, not fixed. Second guess: payload size. Greenhouse sends every full description unless you ask it not to, and Databricks alone is 10.4 MB with descriptions and 0.76 MB without. So we changed filtered searches to read the light list, filter on title, place and dates, and open only the matching jobs, 12 KB each. That was worth doing. It saved about a minute.

Then we read the log properly. The average request took 8.1 seconds. The light 0.76 MB lists download in about one second from a laptop. And there were 54 retries.

Every request went through a rotating proxy, because that is the default on the platform and because "use a proxy" is what scraping advice says. But these are public job-board APIs, published by the companies for their own career pages. Nobody was blocking us. One run with the proxy off:

Same search, same code Through the proxy Direct
Time 9 min 36 s 2 min 00 s
Compute units 0.160 0.034
Average request 8.1 s 1.3 s
Retries 54 0
429 or 403 n/a 0

The named-company run went from 81 seconds to 6. The same default had been on our public Greenhouse, Lever, Ashby, Workday and multi-system job tools. It is off now, with the switch still there for anyone who does get blocked. On pay-per-event pricing the compute is ours to pay, and one small run of each measured a two-to-eight-fold cut in what a customer run costs us.

What we would tell ourselves a week ago

Measure before you optimise, and measure the thing you are about to change. We spent two rounds tuning concurrency and payloads for a slowdown whose cause was one line in the input defaults, visible in the first log we should have read.

And "use a proxy" is advice for sites that block you. For an API that is published to be read, it is a toll you pay on every request.

The tool

Company Jobs API on Apify Store: type company names, get their open jobs from 510+ companies whose boards are verified, or search all of them at once. 30 fields per job, pay as numbers, a monitor mode that bills only new jobs. $1 per 1,000.

For any company outside the index, the Career Site Jobs Scraper reads a career page on 12 systems by URL; there are also single-system versions for Greenhouse, Lever, Ashby and Workday. Public endpoints only, no personal data in any row.

Top comments (0)