If you open a Threads search without being logged in, you get one page of results, roughly 20 posts, and that's it.
No "load more", no cursor, nothing. Every no-login Threads scraper I looked at returns about that many per keyword, and
most of them say so in their README.
I build scrapers, and I wanted more than 20, so I spent some time poking at it. A few notes, in case you're doing the
same.
Search: the 20 posts aren't always the same 20
What I noticed first: if you load the same search twice from two different IPs, you don't get identical results.
There's overlap, but the ranking shifts. And there are actually three different pages that answer a keyword query:
the normal search, the tag search (serp_type=tags), and the /tag/<name> page.
So instead of trying to paginate (you can't, logged out), you sample. Load each of the three surfaces from fresh IPs,
merge by post ID, and stop when three loads in a row bring nothing new. In my tests that lands on 60 to 85 unique
posts per keyword. It's not unlimited, and I'd be lying if I said it was, but it's 3 to 4 times what one page gives
you.
If you need it sorted newest first, sort what you collected. That's not the same as Threads' real "Recent" tab, which
is only for logged-in users, and it's worth being honest about that difference if you sell data to someone.
The Replies and Reposts tabs work logged out
My notes said these needed a login. They don't. The profile page loads them through GraphQL queries with a flag that
the logged-out page just doesn't send by default. Send it and you get the account's replies (with the post they were
replying to) and their reposts (with who reposted and when). That's handy for influencer research: what someone says in
replies is often more revealing than their posts.
The "user not found" false alarm
This one cost me a while. Under load, some real accounts came back as "not found". Threads answers an unknown username
by redirecting to the login page. It also redirects a perfectly fine username to the login page when your IP is being
throttled. Same response, different meaning.
The fix I ended up with: when a profile redirects to login, load a profile that definitely exists from the same IP. If
that one is walled too, the IP is the problem, so rotate and retry. Only if the control profile loads fine is the
username really unknown.
Funny side note: the accounts that kept "failing" in my test list (bbcnews, vogue, samsung, theverge) turned out to
genuinely not exist on Threads under those handles. The scraper was right and my test data was wrong.
E-mails in bios
A lot of business accounts put a contact e-mail in their bio. If you search accounts by keyword and open each profile,
you can pull those out, plus phone numbers and bio links. It only works for people who chose to publish one, which is
also why it's fine to use, but it's surprisingly effective for finding contacts in a niche.
The tool
Disclosure: I packaged all of this as a Threads scraper on Apify. Profiles,
posts, replies, the Replies and Reposts tabs, keyword and hashtag search, account search with bio contacts, and a
monitoring mode that only returns new posts on a schedule. No login. It's $2.50 per 1,000 results on Apify's free plan
and $1.50 on the bigger plans.
Example, hashtag plus keyword search:
{
"searchQueries": ["#buildinpublic", "indie hacker"],
"maxPosts": 80,
"requireKeywordMatch": true
}
If you already use another Threads scraper, the main input fields (usernames, searchQueries, maxPosts,
postedAfter) are named the same, so trying it is just swapping the Actor ID.
There's also an optional mode where you paste the cookie of your own logged-in account. Then search uses the real
Recent tab and pages through it like the app does (300 recent posts for "news" in my test, versus 155 without). It's at
your own risk, use a spare account if you try it.
Top comments (0)