DEV Community

Cover image for Build a YouTube influencer shortlist from subscriber counts and recent views
Joshua Smith
Joshua Smith

Posted on

Build a YouTube influencer shortlist from subscriber counts and recent views

Most sponsorship shortlists start the same way: open a channel, glance at the subscriber count, move on. That works for two creators. For forty, you want a sortable table that answers a better question than "who is biggest?"

Why subscriber count alone misleads

A subscriber total is a running tally of everyone who ever clicked subscribe. It says nothing about how many of them see the next video. A big, quiet channel can reach fewer people per upload than a smaller one that publishes steadily.

These are the numbers that tend to matter for a sponsorship:

  • Typical views per recent video. This is the reach you are buying. Use the median of the last ten or so uploads, not the average: one viral video drags an average far above a normal upload.
  • Views per video as a share of subscribers. If recent videos get views close to the subscriber count, the audience still shows up. A low ratio is a flag, not a verdict: some channels grow through search rather than subscribers, so compare creators in the same niche.
  • Upload consistency. The gap between uploads tells you whether there will be a fresh video to carry your sponsor next month.
  • Engagement per view. Likes and comments divided by views is noisy, so use it as a tie-breaker, not a ranking.

The vanity numbers are the subscriber total on its own, lifetime views (old hits dominate), and a creator's single best video. I am not quoting benchmark percentages: they differ by niche, and nothing in the data below supports one. Compare the creators on your list against each other. Running everyone through the same formulas makes the odd ones out obvious, and gives you something to show whoever approves the budget.

Step 1: channel stats for the whole list

The YouTube Data API publishes these numbers, and the Best Damn YouTube Scraper Actor calls it for a whole list at once. Set Output to Channels only and you get one record per channel: title, handle, url, subscriberCount, viewCount, videoCount, country, publishedAt, description, keywords, thumbnailUrl and uploadsPlaylistId.

This input pulls one row per channel. @handles, channel URLs and channel IDs all work.

{
    "startUrls": [
        "https://www.youtube.com/@YouTube",
        "https://www.youtube.com/@jawed",
        "https://www.youtube.com/channel/UCBR8-60-B28hp2BmDPdntcQ"
    ],
    "outputMode": "channels"
}
Enter fullscreen mode Exit fullscreen mode

The same run over HTTP:

curl -X POST "https://api.apify.com/v2/acts/josh99smith~youtube-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "startUrls": ["https://www.youtube.com/@YouTube", "https://www.youtube.com/@jawed"], "outputMode": "channels" }'
Enter fullscreen mode Exit fullscreen mode

Download the dataset as CSV or Excel, or send it to Google Sheets.

Two things to expect. YouTube rounds public subscriber counts to three significant figures (for example 46,300,000), so rank with them, do not measure with them. And channels can hide the count entirely, which comes back as null. Do not read that as zero: leave the cell blank and rank that creator on views instead. A handle that cannot be found comes back as a free record, such as {"success": false, "input": "@thisHandleDoesNotExist", "errorType": "not-found"}, so a typo never silently vanishes.

Step 2: add recent videos

The channel record only has lifetime totals. For recent reach and upload frequency, run the same list in the default Videos mode with maxResults set to 10. Channels are read newest first, so you get each creator's latest ten uploads, with viewCount, likeCount, commentCount, publishedAt, durationSeconds, channelUrl and channelSubscribers.

Clean up before comparing:

  • likeCount and commentCount are null when the owner hides likes or turns comments off. Skip those rows rather than counting zero.
  • Scheduled streams appear with liveStatus: "upcoming". Drop them.
  • Shorts are included, and durationSeconds lets you separate them from long-form. Mixing them blurs the median.
  • Very recent uploads are still collecting views, so ignore the last few days.

Step 3: compute the metrics

On the channels sheet, with subscriberCount in column B, viewCount in C and videoCount in D:

E2: =IFERROR(C2/D2, "")        lifetime views per video
F2: =IFERROR(E2/B2, "n/a")     that, as a share of subscribers ("n/a" when hidden)
Enter fullscreen mode Exit fullscreen mode

That is a rough first pass, since lifetime numbers lean on old hits. For recent-video metrics, a short script beats spreadsheet formulas:

from collections import defaultdict
from datetime import datetime
from statistics import median
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("josh99smith/youtube-scraper").call(run_input={
    "startUrls": ["https://www.youtube.com/@YouTube", "https://www.youtube.com/@jawed"],
    "maxResults": 10,
})

by_channel = defaultdict(list)
for v in client.dataset(run["defaultDatasetId"]).iterate_items():
    if v["success"] and v["liveStatus"] != "upcoming":
        by_channel[v["channelUrl"]].append(v)

rows = []
for channel, vids in by_channel.items():
    views = [v["viewCount"] for v in vids]
    dates = sorted(datetime.fromisoformat(v["publishedAt"].replace("Z", "+00:00")) for v in vids)
    gap_days = (dates[-1] - dates[0]).days / max(len(dates) - 1, 1)
    subs = vids[0]["channelSubscribers"]
    like_rates = [v["likeCount"] / v["viewCount"] for v in vids if v["likeCount"] is not None and v["viewCount"]]
    rows.append((channel, median(views), median(views) / subs if subs else None,
                 round(gap_days, 1), median(like_rates) if like_rates else None))

# channel, median views, views / subscribers, days between uploads, like rate
for row in sorted(rows, key=lambda r: r[1], reverse=True):
    print(row)
Enter fullscreen mode Exit fullscreen mode

Sort by median views, then read the ratio and the upload gap beside it. The top of that list is your shortlist.

Cost and limits

  • Channels only: $0.001 per channel delivered, so 100 channels cost $0.10.
  • Videos mode: $0.001 per video, 1,000 for $1.00. A hundred channels at ten videos each is 1,000 videos, or $1.00.
  • Private, deleted and unavailable items, invalid inputs and quota errors cost nothing. There is no start fee, and a run stops at the maximum cost you set.
  • It uses the official YouTube Data API v3, not page scraping, on a shared API key with a daily quota of 10,000 units. Reading channels and videos costs about 1 unit per 50 videos, so lists like this barely register. Searches cost 100 units per page, so the shared key allows up to 5 search terms and 100 results per term per run. For heavy use, create a free key in the Google Cloud Console and paste it into YouTube API key. When a quota runs out, the run stops cleanly with free rate-limited records; quotas reset at midnight Pacific Time.
  • It returns public channel and video data only. It does not look for emails or other contact details and collects nothing about viewers. The description text comes back as the creator published it.
  • Audience location, brand fit and whether views are genuine are not in this data. The metrics narrow a list; they do not replace looking at the content.

If this saves you time, a review on the Store page helps other people find it, and a star on GitHub helps too.

Disclosure: I built this Actor. Source: github.com/josh99smith/youtube-scraper. YouTube is a trademark of Google, and the Actor is not affiliated with Google.

Top comments (0)