Give an agent a suspicious domain and it can come back with a convincing story. Several websites share infrastructure. Their pages look similar. A timeline seems to line up.
You still need to know which parts are observations and which parts are the agent joining the dots.
I've also worked on analysing Facebook scam ads. That experience is why I care about the distinction between a lead worth following and a connection I can defend.
Ads impersonating bankers and central-bank governors make a useful example. For a cross-country investigation, the scope could cover ten large European countries and include inactive campaigns. That is a scope to define and check, not a result in itself.
The workflow starts with a suspicion. Choose a test, inspect the result, and use it to decide what to ask next. That also includes dropping a suspicion when the evidence doesn't support it.
Make the question small enough to fail
“Find out everything about this network” is a reasonable opening request. It is a poor definition of done.
For the advertising example, I would first make the scope explicit. Which countries, selected by which measure? What time window? What counts as impersonation? Which sources can actually show inactive ads, and which cannot?
Those decisions affect what a missing result means. “We found no matching ads in this source during this period” is a much narrower claim than “this campaign never ran in that country.”
Then separate candidate explanations. Similar pages might belong to the same operator. They might also use the same template, hosting service, or copied creative. A shared IP is an observation to investigate, not a person's identity.
A prompt I would use at that point:
For each proposed connection, record:
- the exact observation and its source;
- when the observation was captured;
- the explanation it supports;
- a plausible alternative explanation;
- the next check that could weaken the connection.
Keep unsupported connections as hypotheses.
This prompt sets up the investigation. It does not establish what the evidence will show.
Knowing the tools still matters
An agent that can use tools saves a lot of switching between tabs. It still benefits from someone who knows which tool answers which question.
For example, urlscan's API can search existing scans by attributes including domains and IPs, and retrieve scan results. Searching existing material and submitting a new scan are different actions. A submission has a visibility setting; “unlisted” is not the same as private. I would check existing scans first and decide separately whether a new submission is appropriate.
The same care applies to interpreting the result. A scan is an observation from a particular run. It doesn't establish everything a site has ever served or who ultimately controls it.
The human contribution is often choosing a better test. If the question concerns writing style, another infrastructure lookup may add little. Comparing text may be worth exploring, while remembering that shared templates, translation, or editing can also explain similarities. A similarity score is not authorship proof.
Save evidence the next run can inspect
A long chat is an awkward place to maintain a case. I would keep a small working directory alongside it:
case/
scope.md
hypotheses.md
evidence.csv
sources/
runs/
2026-10-02.md
An evidence row can hold a source URL, capture time, saved file, observation, and the hypothesis it relates to. Keep the original material available. A summary that says “three domains were connected” is hard to audit if the actual relationship has disappeared into an earlier conversation.
A run note should record what was checked, what changed, what contradicted the current explanation, and what remains unknown. The next run can then compare new evidence with a named previous state.
I would also record which actions the agent may take. Reading public sources is different from contacting people or submitting information to another service. If an agent proposes sending an email, the person running the investigation needs to decide whether that contact is appropriate. Tool access alone isn't permission to make contact.
Repeating a run doesn't validate its conclusion
Forecasting raises the same questions about repeated assessments and memory. Keeping instructions in a file lets the next run assess new evidence against the same question.
For a forecast, I would define the event, deadline, and resolution criteria before the next run. Save the previous probability and its reasons. Then ask which new observations justify changing it.
Otherwise, you can end up with a fresh, confident paragraph each morning and no way to tell whether the forecasting process is improving. A forecast needs outcomes to compare against. Repetition alone doesn't supply them.
The same applies to an investigation. A later run that repeats an earlier unsupported claim has not found independent corroboration.
A useful result can be “the suspicion got weaker”
I want an OSINT agent to help me examine more possibilities while keeping the evidence behind each connection easy to inspect.
I would rather finish with a narrower claim I can defend than a large network diagram whose edges mean five different things. Before sharing a conclusion, I want to be able to open the source behind each material claim and explain what would change my mind.
Top comments (0)