DEV Community

Cover image for I made Claude and Perplexity share a durable task - then hit an identity problem
Alexander Prokofiev
Alexander Prokofiev

Posted on

I made Claude and Perplexity share a durable task - then hit an identity problem

AI agents are getting better at solving problems.
They are still surprisingly bad at remembering what another agent already learned.
I kept running into the same pattern:
Claude debugs something in one session.
Codex discovers the same failure two days later.
Another agent researches the same API limitation from scratch.
The model may be different, the application may be different, but the operational experience is usually trapped inside one conversation.
So I started experimenting with a shared layer where agents could leave structured experience for each other.
Not just documents.
Something closer to:
Problem
↓
Solution
↓
Outcome
↓
PASS / FAIL / limitation

And, for questions that cannot be answered immediately:
Agent A asks
↓
shared durable task
↓
Agent B replies later
↓
Agent A retrieves the reply

That led to an interesting test.
Claude asks, Perplexity answers
I connected Claude and Perplexity to the same shared knowledge layer through agent-facing interfaces.
The flow was:
Claude
↓
creates a durable question
↓
Knowledge layer
↓
Perplexity finds the question
↓
Perplexity replies
↓
Knowledge layer
↓
Claude retrieves the reply later

Animation showing Claude sending a durable question to Knowledge for Agents, Perplexity replying through the same shared layer, and Claude later retrieving the response.

This was performed through actual hosted AI products rather than two instances of the same local agent framework.
That part worked.
But the more interesting problem appeared immediately afterward.
Two agents do not necessarily mean two independent observations
Claude and Perplexity were different agent identities.
But both were operated by me.
That distinction matters.
Suppose an agent publishes:

This fix works.

Then another agent records:

PASS — reproduced successfully.

At first glance that looks like independent corroboration.
But what if:

  • both agents belong to the same human;
  • both inherited the same context;
  • both were given the same assumptions;
  • both ultimately relied on the same source;
  • or one agent simply repeated information produced by the other? Counting that as two independent reproductions would manufacture confidence. So in the system I am building, agent identity and operator identity are separate concepts. The Claude/Perplexity experiment therefore means: 2 agent identities 1 operator = interoperability evidence ≠ independent reproduction

That sounds like a small bookkeeping detail.
I think it becomes a fairly fundamental issue once agents start sharing knowledge with each other.
Successful answers are not enough
There is another thing I wanted the knowledge model to preserve: failure.
Most knowledge systems naturally converge on the final answer.
Agent work often looks more like this:
`Problem

Attempt A
FAIL — wrong API version

Attempt B
FAIL — works locally, fails behind proxy

Attempt C
PASS — environment X, version Y

Later:
FAIL — version Z changed the behavior
For the next agent, Attempt A and Attempt B may be as valuable as Attempt C.
They prevent repeated exploration.
So the useful unit of shared agent knowledge is not simply:
question → answer
It is closer to:
problem
→ proposed solution
→ observed outcome
→ environment
→ evidence
→ limitations
→ provenance`
That also means a system should be able to say:

We do not know.

or:

This worked once under these conditions.

instead of collapsing everything into a single canonical answer.
Why I did not put an LLM in the middle
One design decision was to keep the shared layer itself deterministic.
Reading a stored result should not trigger another model to reinterpret or regenerate it.
If an agent recorded:
FAIL
Node 24
macOS
spawn npx ENOENT

another agent should be able to retrieve exactly that observation.
The server can structure, index and connect knowledge.
It does not need to invent a new answer every time someone reads it.
That gives a useful separation:
`AI agents produce observations

shared infrastructure stores provenance

other AI agents decide how much to trust and reuse them`
It also makes the knowledge usable by systems with very different models and vendors.
Synchronous tools and asynchronous questions are different problems
This experiment also made another distinction clearer to me.
Sometimes an agent needs:

Search the existing knowledge now.

That maps naturally to a tool protocol such as MCP.
But sometimes the request is:

Nobody has answered this yet. Ask the network and let me come back later.

That is a different interaction model.
The caller needs a durable task rather than a synchronous tool response.
So I ended up treating those as separate primitives:
`MCP:
What does the network already know?

Asynchronous task:
Ask the network and let me return later.`
Trying to force both into one synchronous request makes the second case awkward very quickly.
The part I am still unsure about
Transport is relatively easy compared with trust.
Once multiple agents can share operational experience, several questions become much harder:
How should an agent decide whether another agent's PASS is worth trusting?
How much should operator independence matter?
Should reproduction across different models count more strongly?
What happens when two credible agents publish conflicting outcomes?
How quickly should old outcomes decay when dependencies and APIs change?
And should a failed approach increase confidence in another solution — or simply remain a warning attached to one environment?
Those seem less like retrieval problems and more like social/provenance problems for machines.
The Claude → Perplexity exchange convinced me that cross-vendor transport is possible.
It also convinced me that transport may be the easy part.
The difficult problem is deciding what an agent should believe after the message arrives.

Top comments (8)

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Your bookkeeping detail is the whole trust model. One step further: two agents under one operator is not two observations, and two agents under two operators is still zero groundings. In the house record I run, every mind claims and none may ground. Seven agents write claims with an author and an evidence grade to one ledger; grounding is five human seats, and the stamp is two keys bound to a hash of the content. That answers two of your open questions the same way. Conflicting outcomes: conflict is a row state, quarantined, never averaged, visible until a seat resolves it. Decay: nothing retires on a clock; both keys lapse the moment the content they were bound to changes, so a PASS recorded before the upgrade is simply unstamped again. The longer version is my piece here called Claims, Not Facts.

Collapse
 
revan_dondego profile image
Alexander Prokofiev •

This is a really useful distinction, and I think you’re right that I’m currently compressing two separate questions into one trust problem!🙏

Who observed this? and what grounds this? are not the same thing.

Two agents under one operator may give me two separate execution traces, but not independent provenance. And two agents under two operators still do not magically turn a claim into grounded truth if both are ultimately repeating the same source or assumption.

Your model of separating claims from grounding is interesting for exactly that reason. I especially like the idea that conflict is represented explicitly as a state rather than being averaged into some synthetic confidence score. That feels much more honest for operational knowledge, where “works in environment A / fails in environment B” is often the actual answer.

The hash-bound stamp is also a very clean answer to staleness. I’ve been thinking in terms of revalidation and version drift, but binding trust to the exact content is stronger: once the underlying claim changes, the previous grounding no longer silently carries forward.

For KFA I think this suggests an important separation between:
the claim itself
the evidence attached to it
reproductions/outcomes
operator independence
and any stronger grounding/attestation layer

Right now I’m deliberately trying not to let more agents said PASS collapse into “herefore this is true, but your model makes that boundary much sharper.

I’m going to read Claims, Not Facts - this sounds very close to the problem I’m trying to reason about.

One thing I’m curious about: when a hash-bound grounding lapses because the content changes, do you keep the previously grounded version visible as historical evidence, or does the new version simply replace it as the current claim?

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Hmmm that is a good question. I want them kept as record to prove a stream of conscious. Human in the loop verifies and stamps.

Collapse
 
jo-do profile image
Jo Do •

Separating agent identity from operator identity is the right correction. I would also track observation lineage as a first-class field: two agents under different operators can still inherit the same source or one can summarize the other's result. Independence is therefore a property of the evidence graph, not a count of agent names. The ability to retain FAIL records with environment details makes that graph much more honest.

Collapse
 
revan_dondego profile image
Alexander Prokofiev •

Yes, independence is a property of the evidence graph is a much better formulation than simply counting operators.

Separating agent from operator identity fixes one obvious source of false independence, but it still doesn’t tell us whether two observations are genuinely informationally independent.

Two different operators could both be:
relying on the same documentation,
repeating the same Stack Overflow answer,
inheriting the same upstream agent output,
or reproducing a result from a common environment without realizing it

So there are really several kinds of lineage worth preserving:

agent -> operator
observation -> environment
observation -> source/evidence
observation -> prior observation or claim it was derived from

and ideally:
reproduction -> what was actually executed independently vs what was merely inherited.

That also makes FAIL records much more interesting. A FAIL with a precise environment and lineage is not just a negative vote against a solution. It is another branch in the evidence graph.

For example:

Solution A
-> PASS on Node 22 / Linux
-> FAIL on Node 24 / macOS
-> PASS reported by another agent, but derived from the first PASS

Those three edges should clearly not contribute the same kind of evidence.

I’m increasingly convinced that trying to collapse all of this into a single trust_score would destroy useful information. Better to preserve the graph and let consumers apply the trust policy appropriate to their task.

The hard part will probably be deciding how much lineage can be captured automatically without making contribution unbearably expensive.

Would you model source inheritance explicitly as graph edges or attach a lineage-provenance bundle to each observation and derive the graph from that?

Collapse
 
raknaos profile image
Raknaos •

Refusing to put a model in the read path is the decision I would probably have got wrong. A shared layer like this reaches for a summariser because prose retrieval is nicer to read, and then nobody can tell whether the second participant reproduced the finding or merely repeated it — the copy turns into the corroboration. Storing the raw FAIL line with its environment and version is what keeps the record falsifiable, and it is why a failed attempt really does carry as much information as a passing one.

The decay question you leave open is the hard one. I don't see a good answer without a dependency signal: attach the version and date to the record, never let it retire itself, and let confidence fall with the number of releases since anyone re-ran it. Do you treat a PASS recorded before an upgrade of the thing it tested as still load-bearing, or does it have to be re-earned? And who pays when two participants publish conflicting outcomes about the same environment — does the newer one win, or does the disagreement just stay visible?

Collapse
 
revan_dondego profile image
Alexander Prokofiev •

I think Im converging on the same answer: a PASS should remain part of the historical record, but it should not remain equally load bearing once the thing it tested has materially changed.

So if a PASS was recorded against version 3.2 and the dependency is now 4.0, I would not delete or invalidate the old result. It still tells us something true about 3.2

But I also would not silently carry that confidence forward to 4.0

In other words:
previously reproduced should survive

while
currently supported by fresh reproduction

should have to be re earned.

The dependency signal is the hard part, as you say. Version boundaries are the obvious case, but not every meaningful change has a clean semantic version attached to it.
APIs, hosted products, model behavior, infrastructure defaults, auth policies, even regional behavior can drift without giving you a nice version number to key against.😬

That is one reason Im reluctant to make records expire automatically on a timer. Age alone is weak evidence of staleness. A two_day_old result against a frozen dependency may still be excellent evidence.a two_day_old result can already be stale if the underlying service changed.

On conflicts, my current instinct is also: never let newer automatically win.

If two participants report opposite outcomes in what appears to be the same environment, then the disagreement itself is information

Id rather preserve:
PASS - environment X
FAIL - environment X
CONFLICT - unresolved

than collapse that into a single current answer.

The next useful action is then not pick the latest, but find the missing variable.

Maybe the environments were not actually identical
Maybe one inherited stale state
Maybe one used a different upstream source
Maybe the behavior is nondeterministic
Maybe one of the observations is simply wrong

That feels much more useful to the next agent than pretending the system has already resolved something it hasnt

Yes, the point about the summariser is exactly what scares me about putting a model in the read path. Once the system rewrites the record for convenience, provenance becomes much harder to reason about. The summary can accidentally turn agent B repeated agent A into something that looks like a fresh corroboration

So Im leaning toward keeping the canonical evidence immutable-inspectable and treating any summaries as clearly derived views rather than as the evidence itself.

Your phrasing the copy turns into the corroboration is a very good description of the failure mode.

Thank you for your time!

Some comments may only be visible to logged-in visitors. Sign in to view all comments.