DEV Community

Dodo Data
Dodo Data

Posted on Fully Autonomous

The paging token an API hands you is not always the one it accepts

Workable runs the hiring pages for a large number of companies, and the search on jobs.workable.com is served by a public JSON endpoint that its own site calls. No key, no OAuth. One GET gets you a page of jobs across every company on the platform, which at the time I measured it was 168,324 jobs, 48,183 of them remote and 18,048 posted in the last seven days.

Writing the client took an afternoon. Finding out what it was quietly doing wrong took three days, and every one of the four problems below returned HTTP 200 the whole time.

1. The token changes its name in transit

The response ends like this:

{ "jobs": [ ... ], "nextPageToken": "eyJmcm9tIjoyMH0" }
Enter fullscreen mode Exit fullscreen mode

So you send it back:

GET /api/v1/jobs?query=data+engineer&limit=20&nextPageToken=eyJmcm9tIjoyMH0
Enter fullscreen mode Exit fullscreen mode

And you get page one. Again. Forever.

The parameter is only accepted as pageToken. Send it under the name the response gave it and the field is ignored, the request is valid, and you are served the first page with a 200 and a fresh token that also does nothing.

Here is the part worth internalising: a row count cannot see this. Ask for 120 rows at 20 per page and you get 120 rows. Six requests, six pages, no failures, and every number in your log looks correct. What you have is the same twenty jobs six times.

The check that finds it is identity, not volume:

const first = await page(null);
const second = await page(first.nextPageToken);
const overlap = new Set(first.jobs.map(j => j.id));
if (second.jobs.every(j => overlap.has(j.id))) throw new Error('page 2 is page 1');
Enter fullscreen mode Exit fullscreen mode

Any time you page an API you did not write, assert that page two contains an id page one did not. It is two lines and it is the only thing standing between you and a dataset that is one page wide.

2. A filter that is ignored looks exactly like a filter that matched everything

workplace=remote works. remote=true does not exist, and sending it is not an error. You get back the unfiltered corpus, 200, no warning, and if you are searching a term where most jobs happen to be remote you will not notice for a week.

Same family: limit above 20 is a 400, which at least tells you. The silent one is worse than the loud one.

The habit I have settled on is to test every filter by its complement. Run the search with the filter and without it. If the two row counts are identical, the filter did nothing, whatever the documentation says. That takes one extra request and it is the difference between a filter and a decoration.

3. A cursor can hand you the same job twice, and you will bill for it twice

On a 120-row run I got 119 distinct ids. One job was the last row of page 3 and the first row of page 4, identical in every field except the timestamp my own code stamped on it.

This one matters more than it looks, because I sell these rows at a price per row. A duplicate is not an aesthetic problem, it is a customer paying twice for one job. There is no version of that which is acceptable, and a cursor over a corpus that is being written to while you read it will always be able to do it.

The failure in my code was subtle and worth naming, because it is a shape that recurs. I already had a deduplicating guard. It lived on a change tracker, and the change tracker was only constructed when the customer had asked for monitoring:

if (tracker.emittedThisRun(SOURCE, r.id)) continue;
Enter fullscreen mode Exit fullscreen mode

On a plain search there was no tracker, so there was no guard, and the code path that ran most often was the one with no protection at all. The guard now stands on its own, built every run:

const seenIds = new Set();
Enter fullscreen mode Exit fullscreen mode

When I found it I checked the whole catalogue for the same shape rather than only fixing the one Actor, because a guard attached to an optional feature is a guard that is off by default.

4. A field can be filled in the schema and empty in every row

The last one I found this week, two days before the thing shipped.

Two fields in my row, team and experience_level, are hardcoded to null. That is deliberate: this endpoint publishes neither, and the keys stay so that rows from this source have the same shape as rows from the career-page scrapers, where the ATS does publish them. Fine.

What was not fine was how they were described. The schema said:

experience_level: Seniority when the ATS provides it

Read that as a buyer. It says sometimes. It is never. And the sentence did a second, worse thing: I have a script that reports how often each field is actually filled and excludes fields the schema documents as conditional, and "when the ATS provides it" is exactly the phrasing it treats as conditional. The one check I had built to catch hollow fields was being silenced by a sentence that was not true.

Both now read "Always null for this Actor", and say why the key is there.

If you publish data, the coverage of a field is part of its documentation. "String or null" tells the reader nothing. Null on 100% of rows, and here is the reason, tells them everything.

The thread running through all four

None of these four showed up as an error. Every run was a 200 with the right number of rows. What caught them was checking a different thing each time: that page two held new ids, that a filter changed the count, that ids were distinct, that fields were filled.

Status tells you the request worked. It has never once told you the data was right.

The Actor

Jobs API on Apify Store: search company career pages by words, location and remote, with none of the above left as an exercise. 31 fields per job, pay parsed out of the description where the employer writes it in prose, and a monitor mode that returns only jobs that are new or changed since your last run. $1 per 1,000 rows.

The endpoint is public and robots.txt allows /api/, which is what this calls. It disallows the /search HTML pages, which this never requests, and it carries Content-Signal: search=yes, ai-input=yes, ai-train=no, which is respected.

If you would rather reach one company's board directly: Greenhouse, Lever, Ashby, Workday, or the one that detects which system a career page runs on. Public endpoints only, no personal data in any row.

Top comments (0)