The problem
Patent monitoring usually starts with an export: a PatentsView search export under CC BY 4.0, a rights-holder-authorized dump, or your own portfolio system's records. That export is rarely clean — duplicates, malformed dates, missing attribution — and every downstream consumer (analysts, warehouses, review queues) needs the same things: stable identifiers, preserved provenance, freshness flags, and a defensible answer to "what exactly was delivered and billed?"
Doing that normalization by hand in scripts means re-solving deduplication, licence statements, transformation disclosures and run-level reconciliation for every batch. And the obvious alternative — pointing a scraper at Google Patents or a patent-office site — raises both fragility and rights questions the README deliberately avoids.
What the actor does
The Patent Evidence Normalizer is explicit about what it is: it does not scrape Google Patents, does not call PatentsView, does not log in to a patent office, and makes no legal conclusions. It operates in a zero-network mode. You supply a buyer-owned, rights-holder-authorized, PatentsView CC BY 4.0, or otherwise licensed export (1–100 records), and it:
- validates the closed input contract — rejecting unknown fields, control characters, impossible dates, credentialed or non-HTTPS URLs, malformed identifiers and records missing attribution;
- suppresses duplicate
(sourceName, patentId)pairs before billing, with a free diagnostic reporting the suppression; - preserves source attribution, licence statements, retrieval timestamps and transformation disclosures on every row;
- computes freshness from your supplied retrieval time and threshold, plus a mechanical evidence-confidence score with an explicit gap list and a recommended human-review action;
- assigns deterministic stable identifiers and a canonical row digest for downstream joins and integrity checks;
- writes a current-run KVS
OUTPUTreceipt reconciling requested, unique, duplicate, invalid, delivered, paid, free, withheld and settlement state.
Every row carries safeToAutomate: false because a normalized record is not a legal conclusion. Legacy query-style inputs still parse, but produce a free migration diagnostic instead of any external fetch.
Example: input and output
The documented authorized-export input (a contract fixture, per the README — replace fields with your authorized export):
{
"schemaVersion": "2.0",
"authorization": "I confirm I may process and commercially use these patent records.",
"sourceContext": "patentsview_cc_by_4_export",
"watchName": "battery-thermal-management-review",
"freshnessDays": 30,
"patents": [
{
"sourceQuery": "battery thermal management",
"patentId": "US12345678B2",
"title": "Thermal management system for an energy storage assembly",
"patentDate": "2026-07-14",
"assignee": "Example Energy Systems Inc.",
"abstractSnippet": "A supplied excerpt describing a thermal management assembly.",
"url": "https://patents.google.com/patent/US12345678B2/en",
"jurisdiction": "US",
"recordStatus": "granted",
"sourceName": "PatentsView PatentSearch export",
"sourceUrl": "https://search.patentsview.org/docs/",
"sourceLicense": "CC BY 4.0; attribution and indication of changes required.",
"changesMade": "Selected fields, normalized whitespace, and truncated the supplied abstract excerpt.",
"sourceRetrievedAt": "2026-08-12T10:00:00.000Z"
}
]
}
A delivered evidence row adds stable identity, provenance, freshness, confidence with gaps, a review action, billing intent metadata and a row digest; the full checked example ships with the actor as examples/dataset-item.json. A free diagnostic row, quoted from the README, looks like this:
{
"recordType": "patent_evidence_advisory",
"found": false,
"matched": 0,
"summary": "1 duplicate patent record(s) were suppressed before billing.",
"failureDiagnostics": { "failureType": "duplicate_patent_records", "retryRecommended": true },
"billing": { "billingEligible": false, "billingIntent": "free_diagnostic_dataset_write" }
}
Pricing and the free limit
Pay-per-event: $0.005 per run start plus $0.005 per delivered normalized patent evidence record. Duplicate, invalid and legacy-migration diagnostics emit no result event and are free. A run delivering 100 records costs $0.005 + 100 × $0.005 = $0.505. Apify's free plan gives $5 of usage credits per month, so $5 covers about 9 such runs — roughly 900 normalized records.
Try it
Confirm the rights statement, map an authorized export to the input contract, and start with one record: Patent Evidence Normalizer
For AI agents and MCP
The actor takes a closed JSON contract and returns schema-enforced rows plus a machine-readable OUTPUT receipt, callable from the Apify API, the SDKs, or the hosted Apify MCP server documented in the README. An agent can reconcile delivery and spend from the receipt counters, route rows by decision.outcome and failureDiagnostics.failureType, and rely on safeToAutomate: false to keep normalized evidence headed for human review rather than automated decisions.
Top comments (0)