Engineering teams evaluating developer tools frequently encounter a visibility problem when auditing the GitHub ecosystem. While GitHub Marketplace lists thousands of apps, integrations, and GitHub Actions, GitHub provides no native REST endpoint to bulk-query marketplace listings, compare pricing tiers across categories, or track installation counts over time. Scraping these listings directly requires dealing with pagination, distinct HTML structures between apps and actions, and omitted fields when developers skip optional profile data.
The GitHub Marketplace Scraper resolves this by pulling structured records directly from public marketplace listings without requiring GitHub personal access tokens or OAuth credentials.
Scraping Modes and Input Configuration
The scraper operates in four distinct extraction modes controlled by the mode parameter:
-
searchApps: Scrapes standard GitHub Apps and integrations matching a keyword. -
searchActions: Queries reusable CI/CD pipeline actions. -
byCategory: Pulls all listings tagged with a specific marketplace category. -
bySlug: Retrieves exact metadata for a provided list of listing slugs.
When building a targeted dataset, combining mode with precise filtering properties prevents unnecessary record generation.
{
"mode": "byCategory",
"category": "security",
"planType": "paid",
"sortBy": "most-installed",
"maxItems": 100
}
The category field accepts 17 predefined values matching GitHub's taxonomy, including continuous-integration, code-quality, security, dependency-management, and deployment. The planType property filters the query upstream to return only listings marked as all, free, paid, or free_trials. Sorting options via sortBy include best-match, most-installed, and newest.
If you already maintain a list of specific tools to track, set mode to bySlug and pass the identifiers into the slugs array:
{
"mode": "bySlug",
"slugs": ["dependabot", "snyk-security", "codecov"]
}
This bypasses the search index entirely and targets the underlying listing pages directly.
Handling the Extracted Data Structure
The scraper outputs records under the githubMarketplaceListing record type. Each item contains publisher metrics, categorization, and granular pricing breakdowns.
{
"listingId": 1024,
"slug": "sample-security-app",
"name": "Sample Security App",
"shortDescription": "Automated dependency vulnerability scanner",
"fullDescription": "Deep vulnerability analysis integrated into your PR workflow...",
"url": "https://github.com/marketplace/sample-security-app",
"logoUrl": "https://avatars.githubusercontent.com/ml/1024",
"planType": "paid",
"pricing": [
{
"name": "Open Source",
"monthlyPriceInCents": 0,
"yearlyPriceInCents": 0,
"description": "For public repositories",
"bullets": ["Public repos only", "Community support"]
},
{
"name": "Team",
"monthlyPriceInCents": 4900,
"yearlyPriceInCents": 49000,
"description": "For private repositories",
"bullets": ["Unlimited private repos", "Priority support", "SSO"]
}
],
"categories": ["security", "dependency-management"],
"rating": 4.8,
"installedCount": 12500,
"verifiedPublisher": true,
"publisherName": "Sample Security Inc",
"publisherUrl": "https://github.com/sample-security",
"recordType": "githubMarketplaceListing",
"scrapedAt": "2026-05-15T10:00:00+00:00"
}
A key technical detail of the output schema is missing-field handling: the scraper omits keys entirely when a listing lacks data rather than returning null. For example, listings without user reviews will not contain the rating key, and listings without publicly surfaced installation counts will omit installedCount. Downstream parsers should use safe dictionary lookups (such as Python's .get()) to avoid KeyError exceptions when transforming records.
Execution Walkthrough
You can execute a marketplace run using the Apify API or Python client.
Step 1: Define the Input Payload
Create your configuration targeting specific segments. For example, to audit code review tools by popularity:
run_input = {
"mode": "searchApps",
"searchQuery": "code review",
"category": "code-review",
"planType": "all",
"sortBy": "most-installed",
"maxItems": 50,
}
Step 2: Trigger the Run via Client
Pass the parameters to the Actor endpoint using the official Python client:
from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
actor_call = client.actor("crawlerbros/github-marketplace-scraper").call(
run_input=run_input
)
dataset_items = (
client.dataset(actor_call["defaultDatasetId"]).list_items().items
)
Step 3: Process the Pricing Structure
Extract the raw price metrics from cents into standard currency units for analysis:
for item in dataset_items:
name = item.get("name")
installs = item.get("installedCount", 0)
pricing_plans = item.get("pricing", [])
print(f"Tool: {name} | Installs: {installs}")
for plan in pricing_plans:
plan_name = plan.get("name")
monthly_usd = plan.get("monthlyPriceInCents", 0) / 100
print(f" - Plan: {plan_name} (${monthly_usd:.2f}/mo)")
Platform Pricing and Event Billing
This Actor uses a PAY_PER_EVENT pricing model with platform usage paid by the user. Platform usage consumed during the run is billed separately at your Apify plan's standard rates.
The named events charged for this Actor are:
-
Actor Start (
apify-actor-start): Charged at $0.005 per GB of memory allocated to the run when the Actor starts running (minimum one event). -
result (
apify-default-dataset-item): Charged at $0.005 per dataset item written on the FREE tier. Users on discounted tiers pay reduced per-result rates:- BRONZE: $0.00433 per item
- SILVER: $0.00367 per item
- GOLD: $0.003 per item
- PLATINUM: $0.003 per item
- DIAMOND: $0.003 per item
Filtering your extraction query upstream via category and planType ensures you only write relevant records to the dataset, reducing event charges compared to pulling broad lists and filtering downstream.
Operational Boundaries and Tool Limitations
This scraper processes publicly available listings from the web interface of GitHub Marketplace; it does not access internal GitHub telemetry, private enterprise marketplace portals, or private repository installations. If an app publisher does not list pricing on GitHub Marketplace and instead redirects users to an external website for enterprise licensing, the pricing array will only capture the base public plans visible on the marketplace page.
For data pipelines tracking ecosystem market share, combining category-level searches with monthly scheduled runs provides a consistent view of app adoption and pricing shifts across the GitHub ecosystem.
Everything above runs on GitHub Marketplace Scraper. Start with a small input and a low result limit before you widen the run -- the output shape is easier to check that way.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-26. Check the Actor page for the current rates.
Top comments (0)