DEV Community

Nikita Iakovlev
Nikita Iakovlev

Posted on

What you can (and can't) scrape from LinkedIn without logging in — a head-to-head test of 45 runs

Most "no cookies" LinkedIn scrapers on the market return data that a logged-out visitor never sees: exact connection counts, every past job, people search with "15,666 results". That data has to come from somewhere — usually logged-in accounts on the seller's side. Sometimes that's fine for you. Sometimes it's a compliance problem you only discover later.

I wanted a clear answer to a narrower question: how far can you get using only what LinkedIn shows to a logged-out visitor? So I ran the same inputs through the most popular LinkedIn scrapers and through ours, then upgraded ours and ran everything again.

The setup

  • Same inputs for everyone: "python developer" jobs in New York, 3 well-known public profiles, 3 company pages, posts from 1 profile and 3 companies, ads of 1 advertiser.
  • 3–10 rows per run (we cared about structure and completeness, not volume).
  • 45 runs on the first pass, 26 after the upgrade. Total spend: about $1.80.

What a logged-out visitor really gets

Data type Available without login Not available without login
Jobs everything that matters: title, company, location, seniority, applicants, salary (when posted), full description, poster (≈20%) the external apply link (hidden for guests now), "job function"
Company pages website, size, HQ and offices, industry, specialties, followers, company id, affiliated pages, similar companies, latest posts with engagement open-jobs count, verified badge
Profiles headline, about, current and past roles (incl. grouped roles), education with degree and field of study, languages, courses, projects, location, similar profiles exact connections ("500+"), roles LinkedIn hides from guests, open-to-work flags
Posts text, date, reactions, comments count, media, mentions, reposts, top 8–10 comments on a post page the full comment thread, per-reaction lists
People/employee search, post search — all of it (login wall)

Three bugs we only found by comparing side by side

  1. "Remote only" filter. The filter was applied in the search, but every row still said isRemote: false: the flag was derived from words in the title. Fixed — the flag now follows the filter.
  2. E-mails labelled "deliverable" that were never checked. A default first.last@domain guess with zero evidence was marked deliverable and billed. Now the status says exactly what was checked (found-on-website, pattern-confirmed, unverified-guess), and only confirmed addresses are billed.
  3. Missing posts. We read posts only from the page's JSON-LD block, and LinkedIn puts only some posts there (1 of 7 on one big company page). Reading the visible cards too fixed freshness: the newest post went from 25.09 to 30.09 for the same profile.

Before → after (filled fields on identical inputs)

Actor Before After
Jobs 29 48
Profiles 64 84
Companies 38 67
Posts 17 47

Accuracy where both sides overlap was identical: the same applicant counts, seniority and salary on the same jobs; the same follower counts on the same profiles.

Takeaways

  • For jobs, companies and posts, logged-out data is enough for almost every use case, and it's the safest kind of data to build on.
  • For profiles, you lose some depth without login; decide whether you need it before you pay for it.
  • If a "no login" tool returns things a logged-out visitor can't see, ask where they come from.

The actors used in the test: LinkedIn Jobs Scraper, LinkedIn Profile Scraper, LinkedIn Company Scraper, LinkedIn Posts Scraper (Apify, no login, pay per result).

Top comments (0)