DEV Community

Nikita Iakovlev
Nikita Iakovlev

Posted on

A Threads Search Result Is a Thread, Not a Post

I run a visa agency in Bali. The part of my week that has nothing to do with visas is building scrapers on Apify, and one of them searches Meta Threads: you give it a keyword or a hashtag, it gives you rows of public posts.

The thing that took me longest to get right was not blocking, or pagination, or parsing. It was the unit. Threads does not answer a keyword with posts. It answers with conversations, and the difference decides how many of your rows are about the thing you searched for.

Every number below comes from runs I made on 10 October 2026 from an ordinary paying account, build 0.1.26.

What comes back when you ask for ten posts

Run the keyword coffee shop and visa, ten posts each, top/recent/tag pages all enabled. Twenty rows, and the run summary the Actor writes at the end says this:

{"posts":20,"errors":0,"pages":2,"requests":4,"retries":2,"bytes":5752143}
Enter fullscreen mode Exit fullscreen mode

Two pages. Not twenty requests, not a crawl — one search page per keyword, 2.9 MB of embedded JSON each, and the limit was full. All 20 rows contained the keyword.

But look at where the rows sat:

rows inside a multi-item thread the matching row was a reply
coffee shop 10 2 1
visa 10 5 4

Seven of twenty rows were not standalone posts. They were items inside a conversation, and five of them were replies — someone answering a post that itself never mentioned the keyword:

visa | thr=1/2 | isReply=true | "Free visa kan yah?"
visa | thr=1/2 | isReply=true | "Untuk yg kenegara lainnya perlu visa jg atau bebas visa"
coffee shop | thr=1/2 | isReply=true | "I'm wondering why the owner of the building would lease…"
Enter fullscreen mode Exit fullscreen mode

That is Threads working as designed. A search hit surfaces the thread the hit lives in, and the thread arrives as a bundle: the root post, the reply that matched, sometimes more.

Why the unit matters

Flatten that bundle naively and you get two opposite failure modes, and both of them look fine in a spreadsheet.

Emit every item in every thread and roughly half your rows are context, not matches. On 26 September I measured the same 20-post coffee query against the six most-used Threads search Actors on the Store: they returned between 4 and 17 rows that actually mentioned the word. The rows are not wrong — they are the posts around the match — but if you are counting brand mentions, a dashboard built on them is inflated by whatever the average thread length happens to be that week.

Emit only root posts and you lose the mention entirely whenever the person said your brand name in a reply. In the visa run that would have dropped 4 of 10 rows.

The answer I settled on is to keep the thread structure on the row and let the caller decide. Every row carries:

{
  "threadId": "4004788010406204417",
  "positionInThread": 1,
  "threadLength": 2,
  "isReply": true,
  "rootPostUsername": "…",
  "containsKeyword": true,
  "source": "search-top",
  "searchVariant": "visa",
  "query": "visa"
}
Enter fullscreen mode Exit fullscreen mode

containsKeyword is computed per row against your query — literal substring, or every word present for a multi-word query, so "agents for AI" still matches "AI agents":

const containsKeyword = (row, query) => {
  const hay = `${row.text || ''} ${(row.hashtags || []).join(' ')}`.toLowerCase();
  const q = String(query).toLowerCase().replace(/^#/, '').trim();
  if (hay.includes(q)) return true;
  const words = q.split(/\s+/).filter((w) => w.length > 1);
  return words.length > 1 && words.every((w) => hay.includes(w));
};
Enter fullscreen mode Exit fullscreen mode

Rows that match go into the dataset first. A thread with no literal match anywhere contributes exactly one row — the item Threads itself ranked — and only if the limit is not yet full. That is why both my runs came out 20 of 20 on keyword presence while still being able to hand you the conversation around a mention when you want it. Set strictKeywordMatch and the context rows never appear at all.

The provenance fields do the unglamorous work. source is search-top, search-recent or search-tag; searchVariant is the exact string that was queried; query is what you asked for. Search three keywords at once and these are the only way to tell which term produced which row, and whether your "recent" numbers are actually from the recent feed or from the ranked one.

The part that surprised me: a query can legitimately return nothing

Logged-out Threads serves about one page per query and result type, with no cursor and no date operators. So the only way to widen a search is to ask in more ways. For a two-word term the Actor expands bali visa into four queries — the phrase, a plural of the last word, the words glued together, and that glued form as a hashtag — and reads the pages that exist for each. A hashtag page only exists for a single token, so:

bali visa  → /search?q=…&serp_type=default  +  …&serp_type=recent
bali visas → the same two
balivisa   → those two  +  /tag/balivisa
#balivisa  → those two  +  /tag/balivisa
                                      = 10 pages
Enter fullscreen mode Exit fullscreen mode

Ten pages requested, ten fetched without an error, 9.2 MB downloaded, 17 seconds. Zero rows.

My first instinct was that something was broken. It was not: coffee shop, a two-word query on the same build four minutes later, filled its limit of 10 from the very first page. bali visa simply is not a phrase people post on Threads, and the search index said so ten different ways.

This is worth internalising before you build on any search endpoint: a short answer is data. The expensive mistake is a retry loop that treats an empty result as a transient failure and re-reads ten pages every few minutes forever. Mine charges per row, so that zero-row run billed $0.00 — but the time and the bandwidth were real, and on a per-run pricing model the same mistake costs money.

Running it

Input is the keyword list plus which result pages to read:

{
  "mode": "search",
  "keywords": ["bali", "#ai"],
  "searchTypes": ["top", "recent", "tags"],
  "maxPosts": 50,
  "postedAfter": "2026-10-01"
}
Enter fullscreen mode Exit fullscreen mode

From Python, with the Apify client:

from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("lergassy/threads-search-scraper").call(run_input={
    "mode": "search",
    "keywords": ["coffee shop"],
    "searchTypes": ["top"],
    "maxPosts": 10,
})

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row.get("containsKeyword"):
        print(row["username"], row["likeCount"], row["text"][:60])
Enter fullscreen mode Exit fullscreen mode

A real row from the bali run, trimmed:

{
  "type": "post",
  "code": "DeT3XopAVQB",
  "url": "https://www.threads.net/@wedangputihanget/post/DeT3XopAVQB",
  "text": "Yang di bali yok keluar, bosen bgt kemana mana sendiri 🥲",
  "createdAt": "2026-10-10T10:25:55.000Z",
  "username": "wedangputihanget",
  "likeCount": 15, "replyCount": 14, "shareCount": 1,
  "viewCount": 1014,
  "mediaType": "text",
  "threadLength": 1, "positionInThread": 0,
  "source": "search-top", "query": "bali", "containsKeyword": true
}
Enter fullscreen mode Exit fullscreen mode

Two notes on that row. viewCount comes from the post's own page, not the search page — the bali run did 10 extra lookups for 1.8 MB and came back with a number for 8 of 10 posts and null for the other two, because Threads has not published views for very new posts. Null, not zero: a zero would quietly empty every "over 1,000 views" filter you write. And turning those lookups off (includeViews: false) is what made the two-keyword run finish in 20 seconds against 26 for a single keyword with views on.

One pitfall from today's logs, since it cost me a confused ten minutes: in search mode, startUrls accepts search URLs and /tag/ URLs. Hand it profile URLs and there is nothing to search for — profile URLs belong to the posts mode, which is a different Actor for a different job.

When this is the wrong tool

  • You want one account's full history. That is profile pagination, not search. Search gives you one page per query, and no amount of variants turns that into an archive.
  • You want private accounts, follower lists, or DMs. No login means none of that, ever. The flip side is that nothing you run here can get an account of yours restricted.
  • You need a guaranteed result count per keyword. Threads decides how much it shows a visitor. Anyone promising thousands of posts per keyword from logged-out pages is either logged in or counting duplicates.

What it is good at is the narrow thing: turn a term into the public conversations that mention it, with engagement numbers and enough structure on each row to know whether you are looking at a post, a reply, or the context around one. It is up on the Store as Threads Search Scraper — $2 per 1,000 posts, no start fee, error rows free.

Top comments (0)