This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
A few weeks ago, I was ch...
For further actions, you may consider blocking this person and/or reporting abuse
OriginTrace is officially live. π‘οΈ
It does 3 things:
πScan the web for copies of your DEV.to posts
βοΈ Check whether those copies actually credited you
π¬ Ask the AI Agent to analyze the evidence and generate a DMCA notice when needed.
So here's the fun part: if you found your own post copied tomorrow, which of these 3 would you use first? π
Sir @francistrdev, I'd love an opinion in here. Cause after Hacktoberfest gets over, I'm planning to continue this OSS.π
But no rush, give me review when you're free! Have A Great Day!π
Solid framing. One gap I keep hitting in practice is proving who accessed training or customer data after the fact. Without signed access logs and a clear purpose tag on each read, takedown and breach response both stall. Curious how you would bind that to an MCP or agent tool boundary. iin1004h2028
Thanks for the thoughtful question! That is a massive gap in enterprise AI deployments right now.
Currently, OriginTrace is acting as a public ledger, but if we were to lock this down for private customer or training data, binding identity and purpose to the MCP boundary would be critical. Here is exactly how I would architect that integration:
1. Identity Propagation via MCP Context
When the Next.js frontend initializes the AI SDK and connects to the MCP server, we'd pass a signed JWT in the connection headers. The MCP server validates this token before exposing any tools or resources.
2. Tool-Level Purpose Tagging
We would modify the agent's system prompt and the MCP tool definitions to make a
purposeargument strictly required. For example, the agent couldn't just callfetch_provenance(). It would have to explicitly callfetch_provenance(articleId, purpose="dmca_investigation").3. Immutable Audit Logging in Sanity
Since we're already using Sanity Content Lake for structured data, we would use it as the audit trail. Before the MCP server returns the requested data to the agent, it writes an
accessLogdocument to Sanity containing:By intercepting the request at the MCP server level, we guarantee that the LLM cannot bypass the logging mechanism. The Sanity Content Lake becomes a fully queryable, timestamped ledger of every agent action for compliance teams.
Itβs definitely the next logical step for secure agentic systems! Thanks again for reading.
The public-ledger angle is the interesting one. We syndicate on purpose β mirrors carry canonical_url back to the original β and the only thing separating that from theft is the declared link, which is invisible to any scanner that checks content alone. Does OriginTrace treat a live canonical back-link as declared provenance and de-scope the copy, or does everything unlisted land in the same bucket? That one distinction decides whether syndication keeps working without every mirror having to register first.
Spot on observation! That exact distinction is why I had to build the "Verdict Engine" on top of the NLP scanner.
If OriginTrace only checked for content overlap, every legitimate cross-post and syndication would get flagged as theft.
To answer your question directly: No, they don't land in the same bucket.
When OriginTrace finds a high-overlap copy, it runs a secondary attribution pass over the DOM. If it finds a live back-link to the original DEV URL (or the original author's name), it flips the
hasOriginalLinkboolean totruein Sanity. The Verdict Engine then classifies that specific copy as acredited_syndicationrather than anunattributed_repost. This automatically de-scopes it from the DMCA generation flow.Regarding the invisible
<link rel="canonical">tag specifically: currently, the scraper is looking for visible<a>tags in the body to confirm attribution. However, because the extraction pipeline uses Cheerio to parse the full HTML, adding a check for the<head>canonical tag is a brilliant idea and incredibly easy to add to theverdictlogic.The goal is exactly as you said: let syndication flow freely without friction, while catching the bad actors who strip out the links!
BTW thanks for the read! Have A Great Day!
This one is really good concept, and your future plan with it is also good.
Thanks!π
Good concept and great execution bro, ALL THE BEST!
Thanks bro!
The verdict engine is the part I'd want to stress-test hardest, since labeling something unattributed_repost when it was legitimate syndication has real DMCA fallout. Word-overlap plus 5-gram shingling plus LCS is a reasonable ensemble, but I'd want a precision number on the "unattributed" class specifically before letting it act. Did you hold out a labeled set of known-syndicated versus stolen pairs to measure false positives?
The DMCA draft generation alone makes this worth it. Drafting those from scratch is such a painβhaving the agent pull the evidence and structure it automatically is a huge time saver.