DEV Community

Nikita Iakovlev
Nikita Iakovlev

Posted on Edited on

LinkedIn's Own JSON-LD Lists One of Satya Nadella's Three Schools

I run a visa agency in Bali, and the boring half of that business is people data: who works where, who just changed jobs, which of two namesakes is the one who wrote to us. So I ended up building a LinkedIn profile scraper on Apify, and then spending a month finding out that the easy-looking part of the job is the part that quietly loses data.

The easy-looking part is this: LinkedIn serves logged-out visitors a profile page with a <script type="application/ld+json"> block in it, containing a schema.org Person. Structured data, handed to you, no parsing. Ship it.

Here is what five public profiles actually returned on 8 October 2026, read logged out:

followers connections experience rows education rows skills
Satya Nadella 12,210,986 500+ 5 3 0
Bill Gates 40,712,037 8 3 2 0
Jeff Weiner 10,358,463 500+ 24 0 26
Reid Hoffman 2,796,044 500+ 13 6 47
Justin Welsh 884,055 500+ 18 1 0

Three things in that table are worth a paragraph each, and none of them come from the JSON-LD.

The JSON-LD is a subset of the page it sits in

The Person object is real and useful. interactionStatistic.userInteractionCount gives you the exact follower count — 12,210,986, not the "12M" the visible page prints. worksFor gives current employers with company URLs. alumniOf gives schools.

It gives you some schools. Satya Nadella's public page lists three; the JSON-LD alumniOf array has one. Bill Gates: two on the page, one in the JSON-LD. The pattern, once you look at enough of them, is that alumniOf only carries schools that have dates attached. Nadella's Manipal Institute of Technology entry has no years on it, so it exists in the HTML section and nowhere in the structured data.

And alumniOf has no degree field at all. Everything a recruiter would actually filter on — "degree": "Bachelor's Degree", "fieldOfStudy": "Electrical Engineering" — lives only in the HTML. Reid Hoffman's row from Università degli Studi di Perugia is a good example of how much is down there:

{
  "school": "Università degli Studi di Perugia",
  "schoolLinkedinUrl": "https://it.linkedin.com/school/universit-degli-studi-di-perugia/",
  "startDate": "2024-05",
  "endDate": "2024-05",
  "degree": "Honorary Doctorate",
  "fieldOfStudy": "Human Sciences"
}
Enter fullscreen mode Exit fullscreen mode

So the useful shape is not "parse the JSON-LD" but "parse both, and merge with the JSON-LD winning on the fields it is better at". In my implementation the HTML section drives the list (it is complete), and a school that also appears in the JSON-LD takes the JSON-LD's dates and canonical URL.

The roles hidden inside a company group

Second data loss, same flavour. The Experience section has two kinds of list item. A single role at a company is an experience-item. Several roles at the same company are collapsed into an experience-group whose rows are experience-group-position items — and the company name is not on the row, it is up in the group header.

If your selector only matches experience-item, every promotion-within-one-company disappears. That is not an edge case: it is most senior careers. Mine dropped them for weeks before I noticed, because the output looked perfectly healthy — just shorter.

// both kinds, and walk back to the group header for the company name
for (const m of section.matchAll(
  /<li class="[^"]*(?:experience-item|experience-group-position)[^"]*"[^>]*>([\s\S]*?)<\/li>/g
)) { /* … */ }
Enter fullscreen mode Exit fullscreen mode

HTTP 999 is a refusal, and it is numerically ≥ 500

LinkedIn's "no" to a logged-out request is status 999. It is not a documented code, it is an old Yahoo-era artefact LinkedIn kept, and it has one property that will cost you real money: status >= 500 is true for it.

Every generic HTTP client wrapper I have seen treats 5xx as transient and retries it. So does the obvious hand-rolled version:

if (res.status === 429 || res.status >= 500) return retry();   // wrong
Enter fullscreen mode Exit fullscreen mode

A 999 is per-profile, not per-address. On a 25-profile check, retrying each 999 twice on a fresh residential IP opened zero of eight. On a wider sample, 26 retried refusals produced one extra profile. You are paying for three proxied requests to learn what the first one told you.

The policy that survived contact with reality:

if (res.status === 404) return notFound(res);                    // confirmed, free
if (res.status === 429 || (res.status >= 500 && res.status < 600)) return retry();
if (res.status === 999 || res.status === 403) throw refused();   // never retry
Enter fullscreen mode Exit fullscreen mode

The 200 that is a login wall

Worse than a refusal is a refusal that looks like success. LinkedIn answers a blocked request in two other ways: a 302 into /authwall, /checkpoint/ or /uas/login, and — on some routes — a 200 carrying a page that simply has no profile in it.

Two cheap tests catch both, and I check them on every single response:

const blocked = /\/authwall|\/checkpoint\/|\/uas\/login/.test(res.url || '')
             || !/application\/ld\+json/.test(body);
Enter fullscreen mode Exit fullscreen mode

The absence of application/ld+json is the reliable signal. A real public profile page always has it. A wall never does. If you skip this check, your dataset fills with rows that parsed cleanly out of nothing.

What the address you come from decides

Measured on 1 October 2026, same 150-profile list, logged out:

Route Behaviour
Direct connection HTTP 429 after ~25 profiles
Datacenter proxy HTTP 999 after ~25–50 profiles
Residential, new address per profile 71–83% delivered across the whole list, no decay

The important word is decay. Direct and datacenter do not degrade gracefully, they fall off a cliff at a few dozen requests, which is exactly far enough for your test run to look fine. Residential with address rotation holds a flat success rate from the first profile to the last — but it never reaches 100%, because roughly one profile in five is simply not shown to logged-out visitors on that route at all, and more than one in five on lists of ordinary employees rather than famous ones.

Pages come back gzipped at about 46 KB per request, so bandwidth is cheap: around $0.60 per 1,000 delivered profiles at $8/GB residential pricing.

Three smaller traps, from the scar tissue

Scope your regex to the section. My "website" field was lazily matching the next href after the websites block, which on 71 of 120 profiles was a LinkedIn sign-in or help link. Find the block boundaries first, then search inside them, then reject linkedin.com hosts.

Metro areas have no city. "Redmond, Washington, United States" splits into city / region / country cleanly. "New York City Metropolitan Area" and "San Francisco Bay Area" do not split into anything — two of the five profiles above have city: null. Leave them null rather than guessing; a wrong city is worse than a missing one.

connections is a string, not a number. Logged out, LinkedIn shows the bucket, not the count: four of the five profiles say 500+. Bill Gates's page printed 8. Do not compute on this field.

One nice thing: memberId and profileUrn are in the page, and they are stable in a way the vanity URL is not — people rename their slug, their member id does not change. Reid Hoffman's is 1213, which is a pleasing thing to find in a response body.

What logged-out data simply cannot be

No parsing trick recovers these, so I return nothing rather than inventing something: the exact connection count, the Featured section, interests, "open to work" and "hiring" badges, and on long careers the oldest roles. Reid Hoffman's public page gave me 13 positions; his own headline mentions the PayPal founding team, which is not among them. If your use case needs the full career history or a degree-of-connection, logged-out scraping is the wrong tool and no amount of proxy money fixes it.

Running it

All of the above is packaged as LinkedIn Profile Scraper — No Cookies. URLs or bare slugs in, one row per person out, $2 per 1,000 profiles ($3 from 15 October 2026); profiles that are private, missing or refused come back as error rows that say why and cost nothing.

curl -s -X POST "https://api.apify.com/v2/acts/lergassy~linkedin-profile-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H 'content-type: application/json' \
  -d '{"profileUrls":["https://www.linkedin.com/in/satyanadella","justinwelsh"],"includePosts":false}'
Enter fullscreen mode Exit fullscreen mode
{
  "type": "profile",
  "fullName": "Satya Nadella",
  "headline": "Chairman and CEO at Microsoft",
  "title": "Chairman and CEO",
  "company": "Microsoft",
  "companyLinkedinUrl": "https://www.linkedin.com/company/microsoft",
  "location": "Redmond, Washington, United States",
  "city": "Redmond", "region": "Washington", "countryCode": "US",
  "followers": 12210986,
  "connections": "500+",
  "badges": "Creator, Top Voice",
  "website": "https://snscratchpad.com/",
  "memberId": "19186432",
  "experience": [ /* 5 roles, each with dates, duration, company URL and logo */ ],
  "education": [ /* 3 schools, degree and field of study where the page shows them */ ],
  "similarProfiles": [ /* 20 of LinkedIn's "people also viewed" cards */ ]
}
Enter fullscreen mode Exit fullscreen mode

Keep concurrency low — 3 is the default — because LinkedIn's rate limiting reacts to bursts more than to volume. If you also need the person's work email, that is a different Actor with a different cost structure; this one never touches email and never logs in.

The summary I would have wanted a month ago: treat the JSON-LD as a hint, not as the answer; never retry a 999; and check every 200 for the login wall hiding inside it.

Top comments (2)

Collapse
 
0xgollum profile image
0xGollum •

The 999 being numerically >= 500 is the kind of trap that only shows up on the invoice. One question on the ~1-in-5 profiles hidden to logged-out visitors: did that share change with the exit country of the residential IP, or is it per-profile whatever the route?

Collapse
 
stae_him_ae35092536b300a4 profile image
STAE HIM •

The undated-school gap is a useful test case. For a BrowserAct Bot, I would compare visible education rows with JSON-LD before trusting a repeat run.