DEV Community

Cover image for Why `maxResultsPerProfile` Can Deliver Fewer Results Than Expected
Crawler Bros
Crawler Bros

Posted on

Why `maxResultsPerProfile` Can Deliver Fewer Results Than Expected

The Apify Store is full of pre-built Actors for web scraping, and while they streamline data extraction, relying solely on their public documentation can sometimes lead to unexpected data quality issues or silent failures. This article focuses on the common failure modes that documentation often implies but doesn't explicitly spell out. Using the Tiktok Profile Mention Scraper as our example, we'll examine how its input schema constraints, session limits, and the presence of partial or sentinel values in the output can break your data pipelines, and how to detect these issues defensively in code.

Why maxResultsPerProfile and confirmedMentionsOnly Can Deliver Partial Results?

The maxResultsPerProfile input parameter, which specifies the maximum number of posts to collect per username (1–500), interacts non-trivially with confirmedMentionsOnly. If confirmedMentionsOnly is set to true, the Actor first retrieves results up to maxResultsPerProfile and then filters those results to include only formally confirmed mentions.

This means if many of the initially scraped posts do not contain confirmed mentions, you could receive significantly fewer results than maxResultsPerProfile, or even an empty dataset, without any explicit error. The scraper does not iterate further to meet the maxResultsPerProfile count after filtering.

To guard against this, you should always check the actual number of results returned against your requested maxResultsPerProfile. If isMentionConfirmed is crucial for your use case, enabling confirmedMentionsOnly is appropriate, but be aware of the potential for a lower-than-expected output count. A robust data pipeline should compare the len(output) to the maxResultsPerProfile value you provided. If you have chosen confirmedMentionsOnly: false, then isMentionConfirmed will still be set accurately in the output for each row, allowing you to filter downstream without incurring the implicit "pre-filter" truncation.

Here's how you might configure your input for the scraper, considering this interaction:

{
  "usernames": ["natgeo"],
  "maxResultsPerProfile": 500,
  "confirmedMentionsOnly": true
}
Enter fullscreen mode Exit fullscreen mode

This configuration requests up to 500 results for "natgeo," but only those with confirmed mentions will be returned. If TikTok search returns 500 posts, but only 100 of them have confirmed mentions, your dataset will contain 100 items.

How Do I Detect Missing Data When confirmedMentionsOnly Is Active?

To detect missing or incomplete data, you must compare the number of results received against the maxResultsPerProfile setting in your input. If confirmedMentionsOnly is true, and the number of results is significantly lower than maxResultsPerProfile, it indicates that many potential mentions were filtered out. When confirmedMentionsOnly is false, a mismatch between requested maxResultsPerProfile and actual results might indicate other issues like rate limits or profile availability.

A common pattern is to fetch the run's dataset items and then count them. This Python snippet demonstrates checking the result count:

import apify_client

client = apify_client.ApifyClient("YOUR_APIFY_TOKEN")

# Example input for the run
run_input = {
  "usernames": ["khaby.lame"],
  "maxResultsPerProfile": 100,
  "confirmedMentionsOnly": True
}

# Start the Actor and wait for it to finish
run = client.actor("crawlerbros/tiktok-profile-mention-scraper").call(run_input=run_input)

# Fetch all items from the default dataset
items = client.dataset(run["defaultDatasetId"]).list_items().items

requested_results = run_input["maxResultsPerProfile"]
actual_results = len(items)

print(f"Requested {requested_results} results per profile.")
print(f"Received {actual_results} actual results.")

if run_input["confirmedMentionsOnly"] and actual_results < requested_results:
    print("Warning: Actual results are fewer than requested. Many non-confirmed mentions were likely filtered out.")
elif not run_input["confirmedMentionsOnly"] and actual_results < requested_results:
    print("Warning: Actual results are fewer than requested. Consider checking logs for other issues or increasing maxResultsPerProfile.")

Enter fullscreen mode Exit fullscreen mode

This code explicitly checks if actual_results falls short of requested_results and provides context based on the confirmedMentionsOnly setting. This allows you to differentiate between intentional filtering and potential scraping failures.

What Are the Proxy Session Limits and How Do They Affect Runs?

The Apify platform uses proxies to ensure reliable scraping and avoid IP blocking. Understanding their session limits is crucial. Datacenter proxies typically maintain sessions for about 26 hours, while residential proxies are much shorter, lasting around 30 minutes.

The Tiktok Profile Mention Scraper uses proxies internally, and while you don't directly configure them via the Actor's input, long-running tasks, especially those targeting a large number of usernames or maxResultsPerProfile at its upper limit, might silently hit these proxy session expiry limits.

When a proxy session expires, the scraper might experience a brief dip in performance or even encounter temporary blocks before a new session is established. For most users, this is handled gracefully by the Actor itself. However, if you are running highly parallelized or extremely long jobs, monitoring the run's logs for proxy-related warnings or errors can be beneficial.

The maximum maxResultsPerProfile is 500. If you're targeting many usernames with this setting, the overall run duration can extend significantly, increasing the chance of hitting residential proxy limits if the Actor happens to be routed through them.

When Does The Apify Platform Hard-Cap Synchronous Runs?

Apify's synchronous run endpoint, used for quick API calls, has a hard cap of 300 seconds (5 minutes). If a run exceeds this duration, the API call will return an HTTP 408 timeout error. For the Tiktok Profile Mention Scraper, this is a critical limitation to consider if you're processing a large volume of data or many usernames.

Since maxResultsPerProfile can be up to 500, and you can supply multiple usernames or profile URLs, a single synchronous call can easily exceed 5 minutes. If your run exceeds 5 minutes, you must switch to an asynchronous approach. This involves initiating the run with a POST request to /v2/acts/<actor>/runs and then polling for its completion or using a webhook.

Here's an example of initiating an asynchronous run using curl:

curl -X POST \
  https://api.apify.com/v2/acts/crawlerbros~tiktok-profile-mention-scraper/runs?token=YOUR_APIFY_TOKEN \
  -H 'Content-Type: application/json' \
  -d '{
    "usernames": ["khaby.lame", "charlidamelio"],
    "maxResultsPerProfile": 500,
    "confirmedMentionsOnly": false
  }'
Enter fullscreen mode Exit fullscreen mode

After executing this, you'll receive a runId that you can use to poll for the run status or set up a webhook to be notified upon completion. This prevents the 300-second synchronous timeout from silently failing your data extraction.

What Does an Empty Field in the Output Shape Indicate?

The Tiktok Profile Mention Scraper's README states that "Empty fields are omitted" from the output. This is an important detail for data engineers, as it means you won't always find every field listed in the output schema present in every record. For example, if a post has no hashtags, the hashtags array might simply be absent from that specific JSON object, rather than being an empty array []. Similarly, if no textExtra type=0 tag is found, the mentionEntry object will be missing.

This implies that your downstream processing code should always defensively check for the presence of fields before attempting to access their values. Instead of assuming a field exists and might be empty, assume it might not exist at all.

Consider the output structure for a single post:

{
  "postId": "730335041234567890",
  "postUrl": "https://www.tiktok.com/@creator/video/...",
  "caption": "Check out this awesome mention of @khaby.lame!",
  "targetUsername": "khaby.lame",
  "isMentionConfirmed": true,
  "mentionEntry": {
    "username": "khaby.lame",
    "secUid": "MS4wLjABAAAA...",
    "start": 30,
    "end": 41
  },
  "author": {
    "id": "700000000000000001",
    "username": "creator",
    "displayName": "A Creator",
    "verified": false,
    "avatarUrl": "https://p16-sign-va.tiktokcdn.com/..."
  },
  "likeCount": 1500,
  "commentCount": 50,
  "shareCount": 10,
  "playCount": 150000,
  "createTime": 1678886400,
  "hashtags": [
    {"id": "123", "name": "tiktok"},
    {"id": "456", "name": "viral"}
  ],
  "mentions": [
    {"username": "khaby.lame", "start": 30, "end": 41}
  ],
  "scrapedAt": "2026-09-24T12:00:00.000Z"
}
Enter fullscreen mode Exit fullscreen mode

Now, consider a post where isMentionConfirmed is false and there are no hashtags:

{
  "postId": "730335041234567891",
  "postUrl": "https://www.tiktok.com/@anothercreator/video/...",
  "caption": "This video features khaby.lame in the description, but no formal tag.",
  "targetUsername": "khaby.lame",
  "isMentionConfirmed": false,
  "author": {
    "id": "700000000000000002",
    "username": "anothercreator",
    "displayName": "Another Creator",
    "verified": false,
    "avatarUrl": "https://p16-sign-va.tiktokcdn.com/..."
  },
  "likeCount": 200,
  "commentCount": 5,
  "shareCount": 1,
  "playCount": 20000,
  "createTime": 1678886500,
  "scrapedAt": "2026-09-24T12:01:00.000Z"
}
Enter fullscreen mode Exit fullscreen mode

Notice that mentionEntry and hashtags are entirely absent from the second example, and mentions is also omitted if empty. Your code must account for this. In Python, this means using .get() methods or explicit in checks rather than direct attribute access:

for item in items:
    caption = item.get("caption")
    is_confirmed = item.get("isMentionConfirmed", False) # default to False if not present
    mention_entry = item.get("mentionEntry")
    hashtags = item.get("hashtags", []) # default to empty list if not present

    if mention_entry:
        print(f"Confirmed mention found: {mention_entry['username']}")
    else:
        print("No confirmed mention entry.")

    if hashtags:
        print(f"Hashtags: {[h['name'] for h in hashtags]}")
    else:
        print("No hashtags found.")
Enter fullscreen mode Exit fullscreen mode

This defensive coding pattern ensures your pipeline won't crash when encountering records with omitted fields.

Input Field Constraints and Their Hidden Impacts

The Actor's input schema defines several fields that, while seemingly straightforward, have implicit constraints that can affect your run's outcome.

  • usernames (array) or profileUrls (array): You must provide at least one username or profile URL. If both arrays are empty, the input may not be valid.
  • maxResultsPerProfile (integer): The range is explicitly 1–500. Attempting to set a value outside this range will result in an input validation error. However, as discussed earlier, the actual number of results can be less than this if confirmedMentionsOnly is true.
  • confirmedMentionsOnly (boolean): The default is false. This flag changes the effective meaning of maxResultsPerProfile. If you truly need only validated mentions, setting this to true is correct, but accept that your result count may be lower. If maximum coverage is your goal and you plan to filter downstream, leaving it as false is generally better.

The Apify platform's prefill values in the Console UI are not applied to API calls or existing Actor tasks; only default is. When interacting with the API, always provide an explicit input dictionary to ensure your desired parameters are used, even if they match the UI's pre-filled suggestions.

What Does This Actor Cost Me?

Understanding the cost structure of Apify Actors is essential for managing your budget. The Tiktok Profile Mention Scraper uses a PAY_PER_EVENT model, meaning you are charged for specific events that occur during the run, in addition to general platform usage (which is billed separately at your Apify plan's rates). The named charge events and their prices are:

  • "result" (apify-default-dataset-item): This event is charged at $0.005 per event for each single result item pushed into the default dataset. This is the primary driver of cost, as it scales directly with the number of mention posts collected.
    • Discount tiers apply: FREE $0.005, BRONZE $0.00433, SILVER $0.00367, GOLD $0.003, PLATINUM $0.003, DIAMOND $0.003.
  • "Actor Start" (apify-actor-start): This event is charged at $0.02 per GB of memory allocated to the run. This is a flat charge incurred once per run for the resources provisioned when the Actor starts.

To manage costs, you should monitor the number of results generated. The maxResultsPerProfile input directly influences the potential number of "result" events. Setting it to 500 for multiple usernames will incur significantly more cost than setting it to 50 for a single username. Additionally, if you enable confirmedMentionsOnly, you might get fewer "result" events, potentially reducing this component of the cost, but this comes with the trade-off of potentially incomplete data if unconfirmed mentions are valuable to you.

You can set a maxTotalChargeUsd parameter when starting a run (either via API or the Console) to automatically terminate the run if its estimated cost exceeds a certain threshold. This is exposed to Actor code as ACTOR_MAX_TOTAL_CHARGE_USD. While it terminates the run, it's not an instant kill and may consume resources briefly beyond the threshold.

Real Limitations and Caveats

While powerful, the Tiktok Profile Mention Scraper has specific limitations worth noting:

  • TikTok API Reliability: The Actor relies on TikTok's public search API. Changes to TikTok's internal APIs or website structure can temporarily affect the scraper's functionality until it's updated. This is a common challenge with any web scraping tool.
  • Rate Limits and Blocks: Although the Actor uses proxies to mitigate this, aggressive scraping (e.g., very high maxResultsPerProfile across many accounts in rapid succession) can still trigger rate limits or temporary IP blocks from TikTok. The Actor attempts to handle these gracefully, but in extreme cases, they can lead to slower runs or fewer results.
  • Data Freshness vs. Historical Data: The scraper finds current mentions. While it provides createTime for each post, it does not guarantee finding all historical mentions if TikTok's search API itself has limits on how far back it indexes.
  • textExtra Type 0 for Confirmed Mentions: The Actor specifically validates textExtra type=0 tags for confirmed mentions. This is the most reliable method. If TikTok changes how it represents formal mentions, or if you need to detect mentions that don't create this specific textExtra entry, this Actor's isMentionConfirmed might not cover those cases. However, if confirmedMentionsOnly is false, you still receive mentions[] for other text matches.
  • No Native S3 or Slack Integration: If your data pipeline requires sending results directly to AWS S3 or Slack, you'll need to use webhooks to trigger external services like n8n, Make, or Zapier, or integrate directly with the Apify API client in your custom code. Apify's webhooks support POSTing to a URL on run completion, which is the foundational primitive for such integrations.

Checked against the Actor's input schema and Apify docs on 2026-09-24.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

Top comments (0)