After making one business-registry Actor usable through the Apify MCP server, I tried the next obvious step: give an agent several company-research tools at once.
My test question sounded harmless: “What evidence can you find about Goldman Sachs as a New York entity, public filer, and website?” The first version of my workflow tried to flatten every result into one company object. That broke the moment production data arrived.
The New York registry returned GOLDMAN SACHS & CO. LLC. SEC EDGAR returned GOLDMAN SACHS GROUP INC for ticker GS. RDAP returned a registered goldmansachs.com domain, but public registration data did not prove which legal entity controlled it.
A polished merged answer would have hidden the most important fact: the sources described different records. I changed the Apify MCP agent workflow to keep registry, domain, and filing evidence separate. This article shows the three production Actors, the exact Streamable HTTP calls, and the interpretation contract I added after the failed merge.
Screenshot note: The production Apify MCP tools/list response contains three narrowly scoped Actor tools and one dataset reader.
What I built
I used three Actors I already run in production:
- US Business Entity Search for state registry evidence.
- Domain Availability Checker and WHOIS Scraper for authoritative RDAP, DNS, and domain-age evidence.
- SEC EDGAR Company Filings for the public filer, Central Index Key (CIK), ticker, and filing links.
Each Actor has one source job. The workflow does not ask the registry Actor to infer a public ticker. It does not ask WHOIS to identify a legal entity. It does not treat an SEC filing as a state good-standing record.
The final output is a source-separated evidence bundle:
{
"subject": {
"companyQuery": "Goldman Sachs",
"ticker": "GS",
"domain": "goldmansachs.com"
},
"warning": "The three sources describe different records. Name, ticker, and domain similarity do not establish that they are the same legal entity.",
"evidence": [
{"tool": "business registry", "items": ["..."]},
{"tool": "RDAP and DNS", "items": ["..."]},
{"tool": "SEC EDGAR", "items": ["..."]}
]
}
The warning is part of the data contract. I do not rely on a prompt alone to remember it.
Prerequisites
To reproduce this Apify MCP agent workflow, you need:
- An Apify account and an API token or OAuth-capable MCP client.
- Python 3.10 or newer.
- The
requestspackage. - A small usage budget for three bounded Actor calls.
All three Actors use pay-per-event pricing. Each charges $0.0001 per start event and $0.002 per default dataset item. This test wrote one item per Actor, so the Actor event charges totaled $0.0063. Apify platform usage is separate.
Set your token in the environment instead of putting it in the URL:
export APIFY_TOKEN="your-token"
pip install requests
python company_evidence.py
Scoping one MCP endpoint to three Actors
The Apify MCP server accepts a comma-separated tools parameter. I scoped the endpoint to the three Actors instead of exposing the full Store catalog:
ACTORS = [
"pink_comic/us-business-entity-search",
"pink_comic/whois-domain-lookup",
"pink_comic/sec-edgar-company-filings",
]
URL = "https://mcp.apify.com?tools=" + ",".join(ACTORS)
HEADERS = {
"Authorization": f"Bearer {os.environ['APIFY_TOKEN']}",
"Accept": "application/json, text/event-stream",
"Content-Type": "application/json",
}
That choice made tool selection easier to inspect. The live tools/list response contained the three Actor tools plus get-dataset-items. I failed the script immediately if any expected tool was missing:
_, tools = rpc("tools/list", {}, 2)
available = {tool["name"] for tool in tools["result"]["tools"]}
expected = {
"pink_comic--us-business-entity-search",
"pink_comic--whois-domain-lookup",
"pink_comic--sec-edgar-company-filings",
"get-dataset-items",
}
missing = expected - available
if missing:
raise RuntimeError(f"missing MCP tools: {sorted(missing)}")
This check caught a class of error I had seen before: trusting local Actor names without confirming the hosted MCP names. An agent cannot call a tool that the live server exposes under a different identifier.
Running an Actor and retrieving its dataset
A successful Actor call does not place every result in the MCP response. It returns run metadata and storage references. My helper therefore performs two tool calls: run the Actor, then read its default dataset.
def run_actor(tool_name, arguments, request_id):
_, result = rpc(
"tools/call",
{"name": tool_name, "arguments": arguments},
request_id,
)
run = result["result"]["structuredContent"]
if run["status"] != "SUCCEEDED":
raise RuntimeError(f"{tool_name} ended with {run['status']}")
dataset_id = run["storages"]["datasets"]["default"]["id"]
_, dataset = rpc(
"tools/call",
{
"name": "get-dataset-items",
"arguments": {
"datasetId": dataset_id,
"limit": 10,
"clean": True,
},
},
request_id + 1,
)
return {
"tool": tool_name,
"runId": run.get("id") or run.get("runId"),
"datasetId": dataset_id,
"items": dataset["result"]["structuredContent"]["items"],
}
I kept both IDs. The dataset ID lets the workflow retrieve output. The Actor run ID gives me a production record to inspect when a result looks wrong.
The helper also rejects any non-SUCCEEDED run. A transport-level HTTP 200 only means the MCP request completed. It does not mean the Actor succeeded.
Giving each Actor a bounded request
I used one result from each source. The inputs were intentionally specific:
registry = run_actor(
"pink_comic--us-business-entity-search",
{
"searchQuery": "Goldman Sachs",
"states": ["NY"],
"maxResults": 1,
"fetchDetails": True,
"waitSecs": 45,
},
3,
)
whois = run_actor(
"pink_comic--whois-domain-lookup",
{
"domains": ["goldmansachs.com"],
"maxResults": 1,
"concurrency": 1,
"includeRawWhois": False,
"waitSecs": 45,
},
5,
)
sec = run_actor(
"pink_comic--sec-edgar-company-filings",
{
"tickers": ["GS"],
"formTypes": ["10-K", "8-K"],
"maxFilings": 4,
"waitSecs": 45,
},
7,
)
maxResults caps paid registry and domain items. maxFilings caps the nested filing array inside one paid company item. Those limits are different, and the workflow preserves that distinction.
I disabled raw port 43 WHOIS because RDAP and DNS already supplied the evidence needed for this test. The smaller output is easier for an agent to reason about and avoids carrying a large unstructured response into context.
What the production Apify MCP agent workflow returned
I ran the complete workflow against production on July 30, 2026. All three Actors succeeded and wrote one dataset item:
- Registry run
ZtiagTQxjUQwUaYiFcompleted in 1.27 seconds. - WHOIS/RDAP run
ghqLAKD0Znw5Ai1yZcompleted in 1.54 seconds. - SEC EDGAR run
Vp3Fl5xBaZeFf5CvVcompleted in 1.73 seconds.
Screenshot note: Three production Apify Actor runs succeeded in under two seconds and exposed their run and dataset IDs.
The New York registry returned entity ID 1560743, type DOMESTIC LIMITED LIABILITY COMPANY, and source status Active for GOLDMAN SACHS & CO. LLC.
The domain Actor returned goldmansachs.com as registered. Verisign's authoritative .com RDAP endpoint reported a creation date of July 25, 1995. DNS records were present.
SEC EDGAR resolved ticker GS to CIK 0000886982 and GOLDMAN SACHS GROUP INC. The bounded response contained four recent 8-K filings.
Screenshot note: The state registry, domain registry, and SEC source returned different identifiers and legal names.
The results were individually useful. They were not safe to collapse into one entity row.
The failed merge was the important result
My first model used fields such as legalName, status, domainCreated, ticker, and latestFilings. It implied that one resolved company owned every value.
Production output disproved that assumption. The registry record and SEC filer had different names and entity types. RDAP supplied domain evidence but no reliable bridge to either legal record.
I replaced the merged model with three evidence sections. Each section keeps its own:
- Source identifier.
- Source URL.
- Retrieval time.
- Status vocabulary.
- Coverage and interpretation notes.
The agent can now say, “The New York source returned this LLC record, Verisign RDAP returned this domain record, and SEC EDGAR returned this filer.” It cannot silently turn similarity into identity.
Adding an interpretation contract for the agent
I wrote down what each source can and cannot support.
Screenshot note: An interpretation matrix separates the evidence supplied by each Actor from conclusions it cannot establish.
The state registry supports a named record, entity ID, and point-in-time source status. It does not prove certified good standing, ownership, or onboarding approval.
RDAP and DNS support domain registration and infrastructure observations. They do not prove who controls the domain, whether the website is safe, or whether the domain belongs to a particular registry entity.
SEC EDGAR supports a named public filer, CIK, ticker mapping, and company-filed disclosures. A filing is not SEC verification of every claim, and the ticker does not link itself to a state registry result.
This separation also improves failure handling. If RDAP is unavailable, the workflow can keep registry and SEC evidence while labeling the missing domain source. It should not downgrade the whole subject to “unverified” or treat the missing result as proof that no domain exists.
What I would change next
I would add an explicit per-source outcome object: found, no_match, source_unavailable, or invalid_input. An empty array is still too easy for an agent to misread.
I would also add an optional record-linkage stage after evidence collection. That stage would compare names, addresses, identifiers, and disclosed relationships. It would output a confidence explanation, never overwrite the raw source sections.
Finally, I would execute independent Actor calls concurrently. The production runs each took under two seconds, but my reference script runs them sequentially for easier debugging. Parallel calls would reduce end-to-end latency without changing evidence semantics.
The main lesson was not about adding more tools. It was about refusing to let more tools create a stronger conclusion than their sources support. The Apify dataset model made it easy to preserve each Actor's output. The hard part was keeping the agent from erasing those boundaries during synthesis.
FAQ
Why use one MCP endpoint instead of three separate servers?
One scoped endpoint gives the client a small, coherent tool set and one authentication path. The Actor inputs and datasets remain separate.
Does this workflow perform KYB verification?
No. It collects source-linked evidence for review. It does not verify identity, beneficial ownership, certified good standing, sanctions status, or onboarding eligibility.
Why not let the language model decide whether the records match?
A model can suggest possible relationships, but the raw records should remain visible. Legal-entity linkage needs explicit evidence and review, not name similarity alone.
Where is the complete code?
All workflow-specific code—the scoped endpoint, live tool check, Actor runner, dataset reader, and exact inputs—is included above. The three linked Actor pages document the live tools, while Apify's MCP documentation covers the hosted server and client setup.
Top comments (0)