Prior art searches, legal technology platforms, and competitive intelligence pipelines frequently require programmatic access to global patent datasets. While official government registries like the USPTO or EPO expose APIs, their data models vary significantly, rate limits are strict, and crossing international jurisdictions requires building separate collectors for each patent office.
Google Patents aggregates data from over 80 worldwide patent offices into a unified interface. Scraping this interface directly via headless browsers typically leads to high infrastructure overhead and aggressive IP blocking.
The Google Patents Scraper solves this by making direct HTTP requests via curl_cffi using a Chrome 131 fingerprint. It bypasses browser rendering entirely, providing structured JSON output covering titles, claims, CPC/IPC classifications, assignees, backward/forward citation graphs, and direct PDF download links.
Scraping Modes and Parameter Selection
The actor handles querying through seven distinct operation modes passed via the mode parameter:
-
searchPatents: Standard free-text search supporting native Google Patents syntax. -
byPatent: Direct extraction of single or multiple publication numbers. -
byInventor: Returns patents associated with an inventor name. -
byAssignee: Filters by patent owner or corporate assignee. -
byClassification: Extracts records matching specific Cooperative Patent Classification (CPC) or International Patent Classification (IPC) codes. -
byCitation: Traverses one layer of the citation graph for a publication number. -
byUrl: Direct scraping of rawpatents.google.comURLs.
For standard research pipelines, balancing response size and execution speed depends on the fetchFullDetails boolean flag.
When fetchFullDetails is set to false during search operations (searchPatents, byInventor, byAssignee, or byClassification), the actor yields rapid search result snippets. These include the title, dates, primary inventor, primary assignee, and figure thumbnails. Setting fetchFullDetails to true causes the actor to issue an additional HTTP request per patent page, populating full abstract, claims, backwardCitations, forwardCitations, and description text (truncated at 50,000 characters).
Here is an example JSON configuration for querying AI patents within specific jurisdiction and date constraints:
{
"mode": "searchPatents",
"query": "neural network hardware accelerator",
"country": "US",
"status": "GRANT",
"filingDateFrom": "2020-01-01",
"publicationDateTo": "2024-01-01",
"fetchFullDetails": true,
"maxItems": 100,
"language": "en"
}
Traversing Citation Graphs and Metadata Schemas
Tracking prior art requires traversing citation networks. The byCitation mode allows targeting a known publication number (such as US11704715B2) and defining citationDirection as backward (patents cited by the target), forward (subsequent patents citing the target), or both.
{
"mode": "byCitation",
"publicationNumber": "US11704715B2",
"citationDirection": "both",
"maxItems": 500
}
Each returned dataset item includes a standardized schema regardless of the source patent office:
{
"publicationNumber": "US11704715B2",
"country": "US",
"title": "Method and system for distributed data processing",
"abstract": "A system and method for scaling pipeline execution across clusters...",
"claims": [
"1. A system comprising a processor and memory...",
"2. The system of claim 1, further comprising..."
],
"claimCount": 2,
"inventors": ["Jane Doe", "John Smith"],
"primaryInventor": "Jane Doe",
"assigneesOriginal": ["Tech Corp LLC"],
"assigneesCurrent": ["Holding Co Inc"],
"primaryAssignee": "Holding Co Inc",
"priorityDate": "2018-05-12",
"filingDate": "2019-05-10",
"publicationDate": "2023-07-18",
"grantDate": "2023-07-18",
"legalStatus": "Active",
"classifications": ["G06F9/50", "H04L67/10"],
"backwardCitations": [
{
"publicationNumber": "US9876543B1",
"title": "Distributed load balancing"
}
],
"forwardCitations": [],
"pdfUrl": "https://patentimages.storage.googleapis.com/pdfs/us11704715.pdf",
"scrapedAt": "2024-03-29T10:00:00.000Z"
}
The emitted records automatically drop null, empty string, or empty array keys via an internal recursive omit-empty filter before storage, keeping dataset sizes minimal.
Executing a Targeted Extraction Pipeline
To use the scraper in a automated workflow using Node.js or Python, follow these steps:
-
Select the Lookup Axis: Identify whether your pipeline relies on free text (
searchPatents), specific company portfolios (byAssignee), or classification codes (byClassification). -
Define Boundaries: Pass
filingDateFrom,filingDateTo,country, andstatusto limit the candidate set before requesting full details. -
Configure Anti-Blocking Options: The scraper runs directly over HTTP without proxies by default. Set
autoEscalateOnBlocktotrue(enabled by default) so that if Google returns HTTP 429 or 403 status codes, the scraper automatically routes subsequent requests through the Apify proxy pool without terminating the run. - Execute and Stream Results: Poll or stream items directly from the default dataset using the Apify SDK or REST API.
Here is a minimal Python integration using the apify-client library:
from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run_input = {
"mode": "byAssignee",
"assignee": "Google LLC",
"classification": "G06N3/00",
"fetchFullDetails": False,
"maxItems": 250,
"autoEscalateOnBlock": True
}
run = client.actor("crawlerbros/google-patents-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(f"{item.get('publicationNumber')}: {item.get('title')}")
Actor Event Pricing Structure
This actor uses a pay-per-event pricing model alongside standard platform usage billed separately at your Apify plan's rates. Costs are explicitly assigned based on operational run events:
- Actor Start (
apify-actor-start): $0.005 per GB of memory allocated to the run upon initialization. - Dataset Result (
apify-default-dataset-item): $0.005 per result added to the default dataset under the FREE tier.
For users on higher account tiers, the per-result event price scales as follows:
- BRONZE: $0.00433 per result item.
- SILVER: $0.00367 per result item.
- GOLD: $0.003 per result item.
- PLATINUM: $0.003 per result item.
- DIAMOND: $0.003 per result item.
For example, extracting 1,000 full patent records on a FREE tier account generates 1,000 result events costing $5.00, plus the flat startup event cost per GB of memory allocated, in addition to the underlying platform usage.
Technical Boundaries and Limits
This actor does not perform multi-hop citation graph traversals automatically in a single execution. If your analysis requires building a multi-generational network graph (citing patents of citing patents), you must capture the first-hop publication numbers from the initial run's output and pass those array values into a subsequent run using mode="byPatent" or mode="byCitation". Additionally, single search queries are capped at a maximum of 1,000 items per run by Google Patents, requiring query pagination via date slices (filingDateFrom and filingDateTo) for large-scale portfolio extractions.
Everything above runs on Google Patents Scraper. Start with a small input and a low result limit before you widen the run -- the output shape is easier to check that way.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-29. Check the Actor page for the current rates.
Top comments (0)