DEV Community

Cover image for How to Detect Newly Registered UK Companies Without a Paid API
Tim Zinin
Tim Zinin

Posted on Originally published at apify.com

How to Detect Newly Registered UK Companies Without a Paid API

The problem

Newly incorporated companies are one of the oldest lead seeds in B2B: a business that registered last week does not yet have an accountant, a bank, an insurer, or a software stack. The raw material is public — Companies House publishes every registration — but turning it into a usable list is tedious. You have to search the register keyword by keyword, check incorporation dates against your window, copy company numbers and status by hand, and somehow avoid counting the same entity twice when two keywords match it.

Commercial "new company data" feeds exist, but they bundle the same public register behind a subscription and often dress a registration up as qualified demand. For a small team that just wants "active companies matching these keywords, incorporated in the last N days," a bounded, pay-per-result tool is a better fit.

What the actor does

The New UK Company Detector queries the public Companies House search-results pages and returns the first results page of active companies matching your keywords, filtered to a configurable incorporation lookback window. There is no register API key and no bulk-data contract involved.

From the README, the concrete behavior:

  • Accepts up to 25 case-insensitive company-name keywords per run; case variants are deduplicated, and each keyword returns the first search page — typically up to about 20 active matches incorporated within the window.
  • windowDays bounds how recent "new" means: 1 to 365 days, default 60. maxConcurrency (1–20) controls parallel keyword searches.
  • Emits one row per unique company number with the official register facts: company name, company number, direct Companies House URL, incorporation date, legal type, status, and registered-office locality (town/region only).
  • Deduplicates across keywords inside the run: if two keywords return the same company, it is delivered and charged once, and the repeat increments duplicateCompanyCount instead of selling a duplicate row.
  • Wraps each row in decision metadata: which keyword and window produced it, observation timestamp, incorporation age and freshness band, an evidence confidence score with explicit dataGaps (website, industry, contacts, decision maker, commercial fit), a recommended enrichment-first action, action priority, and a plain-language interpretationBoundary.
  • Confirmed no-match and operational outcomes are explicit, free rows — never a silent empty dataset — and a run-level KVS OUTPUT receipt reconciles delivered, paid, free, withheld, duplicate, and failed counts.

The README is equally clear about what it never claims: a found row is a recent official registered-name observation — a lead seed requiring enrichment — not proof of trading activity, buyer intent, revenue, or consent to outreach. It also does not page beyond the first search result.

Example: input and output

The input is three fields:

{
  "items": ["capital", "consulting"],
  "windowDays": 60,
  "maxConcurrency": 5
}
Enter fullscreen mode Exit fullscreen mode

The README does not ship a fixed output fixture; the dataset schema is a row per unique delivered company number. Per the field dictionary, a delivered company row carries entityId, companyNumber, officialCompanyUrl, companyName, incorporatedOn, status, companyType, addressLocality, plus keyword and searchWindowDays explaining why the row matched, observedAt / incorporationAgeDays / freshnessBand for timing context, confidenceScore / confidenceBand / confidenceConflict for evidence support, dataGaps listing what enrichment is still missing, sourceEvidence with the official register link, and routing fields recommendedAction (for example ENRICH_COMPANY_BEFORE_OUTREACH), actionPriority, and safeToAutomate. A live sample dataset is linked on the actor page for inspection before you integrate.

Pricing and the free limit

Pricing is pay-per-event: $0.005 per actor start plus $0.005 per delivered company evidence row (result-found), with lower per-event prices on paid Apify plans — the README headline works out to $4.25 per 1,000 rows at the entry discount tier. At list prices, a run delivering 100 companies costs $0.005 + 100 × $0.005 = $0.505. Apify's free plan includes $5 of usage credits per month, which covers about 9 such 100-row runs — roughly 1,000 individual company rows — before anything is charged. Explicit no-match and operational advisory rows are free.

Try it

Start with one narrow keyword and a 30–60 day window, inspect the official URLs and confidence fields, then schedule the watch: New UK Company Detector & Lead Seeds.

For AI agents and MCP

The actor takes JSON in and returns structured JSON rows plus a KVS OUTPUT receipt via the Apify API, so an agent can call it directly. The README documents both an apify-client recipe and an Apify MCP call (call-actor against zinin/new-company-detector, with an MCP server setup block), and its webhook consumer policy tells agents to branch on failureType and safeToAutomate and to send safeToAutomate=false rows to a human-review queue rather than treating a registration as outreach permission.

Top comments (0)