GitHub search is great for browsing and terrible for comparing. If you want the top 100 Python repositories for "vector database" with their stars, license, last push date and latest release in one table, you end up opening tabs and copying numbers by hand. The same goes for keeping an eye on the releases of a list of dependencies.
This guide shows how to get that table in one run with a small Apify Actor published by Hay Equipos called GitHub Repository Search, Stats and Releases. It uses GitHub's official REST API, works without a key, and returns one clean row per repository.
What the tool returns
You can give it search queries, a list of repositories, or both. Search queries accept any GitHub qualifier you would type on github.com, such as topic:rag, org:vercel, license:mit or in:readme.
Each repository row includes the full name, owner, link, description, homepage, main language, topics, stars, forks, open issues, license, fork and archive flags, default branch, and the creation and last push dates. Optionally, it adds:
- the latest stable release (tag, name, date and link)
- recent releases as separate rows, with release notes, asset count and total download count
- the README in Markdown, cut to a length you choose, ready for an LLM or a search index
A repository row looks like this (values are illustrative):
{
"type": "repository",
"source": "search",
"query": "vector database",
"rank": 1,
"fullName": "example-org/example-vectordb",
"url": "https://github.com/example-org/example-vectordb",
"language": "Python",
"topics": ["vector-database", "embeddings", "rag"],
"stars": 18450,
"forks": 1320,
"openIssues": 214,
"license": "Apache-2.0",
"isArchived": false,
"createdAt": "2022-05-03T09:12:44Z",
"pushedAt": "2026-09-29T17:40:02Z",
"latestRelease": { "tag": "v2.4.0", "publishedAt": "2026-09-20T12:00:00Z" },
"readme": "# Example VectorDB ..."
}
Searches with no matches and repositories that do not exist are listed in a RUN_SUMMARY record in the run's key value store and cost nothing.
Step by step in the Apify Console
- Open the Actor from its Apify Store page and sign in to Apify Console.
- In the Input tab, add one or more Search queries, for example
vector database, or list Repositories to look up asowner/nameor github.com links. - Pick Sort search results by (best match, most stars, most forks or recently updated) and add filters if you want them: Language, Minimum stars, Created on or after and Last code push on or after (dates like
2026-01-01). - Set Maximum repositories per query (up to 1,000, GitHub's own cap for one search).
- Switch on Add latest release and Add README text if you need them, and set Release rows per repository above 0 for release notes.
- Click Start, then export from the Output tab as CSV, JSON or Excel.
To find new and rising projects, set Created on or after to a recent date and sort by stars.
How to call it from code
With curl:
curl -X POST "https://api.apify.com/v2/acts/pistachio_implementation~github-repo-search-stats/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchQueries": ["vector database"], "language": "Python", "minStars": 1000, "sort": "stars", "maxReposPerQuery": 50}'
In Python, with the apify-client package:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/github-repo-search-stats").call(
run_input={
"repositories": ["apify/crawlee"],
"includeLatestRelease": True,
"includeReadme": True,
}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["type"], row["fullName"], row.get("stars"), row.get("latestRelease"))
Pricing
Pay per event, with no start fee and no platform usage charge on top:
- Repository saved: $0.001, which is $1 per 1,000 repositories. The latest release and the README are included in that price.
- Release saved: $0.0003, which is $0.30 per 1,000 release rows.
Searches with no matches, repositories not found, and releases not fetched because of GitHub's rate limit are free. As an example, the top 100 Python repositories for a topic with READMEs cost $0.10. You can cap the spend of any run with the maximum charge setting in Apify.
Limits and what it does not do
-
GitHub's own limits apply. Without a token, GitHub allows 10 searches a minute and 60 other calls an hour from one address. Plain searches need no extra calls, so they work fine without a token. The latest release and release rows use one extra call per repository, so without a token they cover about 50 repositories an hour. Once that allowance runs out, the remaining repositories are still saved without release data, and
RUN_SUMMARYshowscoreLimitHit. - Add a token for big runs. A free personal access token with no scopes raises the limits to 30 searches a minute and 5,000 calls an hour. It is a secret input and is never written to the output.
- One search returns at most 1,000 repositories. Split large searches by language, date or star range.
- Public repositories only. No user profiles, emails or contributor lists are collected.
-
openIssuesis GitHub's own count and includes open pull requests.watchersis filled only for repositories looked up by name.
Please use the data in line with GitHub's terms of service and API terms.
Try it on the Apify Store: https://apify.com/pistachio_implementation/github-repo-search-stats
Top comments (0)