Anthropic posted on September 23 that Claude agents found a new enzyme system hiding in phage DNA. The numbers in the post: roughly 950 agents, 21 hours, 210 million tokens, one giant database of DNA sequences. One agent noticed a repeating pattern next to a gene that looked odd. Humans did the lab work, named the system ART, and released a pre-print. When I checked around 1 PM Singapore time the Hacker News thread sat at 567 points and 588 comments, and most of the fight is about one word: discovery.
I want to argue about a different part. I build small agent pipelines myself, a security scanner harness and the automation that helps publish these articles, and the published numbers describe an architecture more interesting than the headline. This is a teardown of that pipeline by someone who did not run it. Five things I would steal from it, and one question I cannot shake.
# The funnel Anthropic published, sketched the way I sketch any agent campaign
funnel = {
"wall_clock": "21 hours",
"agents": "roughly 950 parallel Claude sessions",
"tokens": 210_000_000,
"gathered": 200_000, # RTs pulled from the database ("over 200,000")
"candidates": 3_500, # new candidate systems
"reported": 20, # human-readable reports
"lab_named": 1, # ART, after human lab work
}
print(f"survival rate: {1 / funnel['gathered']:.6f}")
# survival rate: 0.000005
What were roughly 950 agents doing for 21 hours?
Per the post, human involvement was "limited to the initial prompt and the lab work, while Claude agents combed through the database, investigated the distinct RT families, and used their own judgement to identify interesting candidates." The tools were Claude Science and Claude Code, which Anthropic describes as "the same tools available to any scientist," plus a harness of their own "that coordinates many Claude sessions running in parallel." No specific model is named anywhere on the page.
The funnel is the finding
The agents gathered over 200,000 reverse transcriptases, picked 3,500 new candidate systems, wrote human-readable reports for the 20 most compelling, and the lab work named exactly one previously uncharacterized system. The post adds that for an expert scientist "this type of analysis can take weeks to months of work."
That shape is not exotic. Fan out cheap, filter hard, escalate survivors. What is new is running it over a scientific database where the final ground truth check lives in a wet lab instead of a test suite. The engineering lesson sits in the middle of the funnel: almost everything died in software, not in the lab.
Where is the verification loop in this design?
This is the part of the post I re-read twice, because it is the most transferable. Before hunting anything new, Claude "reads the relevant literature and reproduces the established results from public data to check its methods." Then it searches for genomic neighbors that fit no described system. Then a human-readable report for each candidate. Then adversarial self-review, where "typically most candidates are eliminated at this stage. A survey may end with a single candidate worth testing, or with none."
Here is that gate in the shape I use locally:
from dataclasses import dataclass, field
@dataclass
class Candidate:
claim: str # what the agent thinks it found
evidence: list[str] = field(default_factory=list) # pointers, not vibes
status: str = "open" # open, eliminated, reported, lab
def gate(c: Candidate, calibration_ok: bool) -> Candidate:
if not calibration_ok:
c.status = "eliminated" # known results failed to reproduce
elif not c.evidence:
c.status = "eliminated" # a claim without pointers is noise
return c # humans own the "lab" transition
Reports are artifacts, not vibes
The agent that found ART "counted the repeats and measured their spacing, compared the layout with the known RT systems, and searched the literature for any previous report of the pattern" before filing its report for human review. That trace is what makes the claim checkable. I learned to care about this the hard way when my own agent reported a prompt injection that had never happened, and the logs said otherwise. A security alert from your own tooling is still just an output that can be wrong. What separates a finding from a story is the artifact trail underneath it.
Why would reruns miss the find the pipeline already made?
Here is the detail that got less attention than the exclamation mark. The Next Web's read of the pre-print reports that reruns of the same search missed the repeat array 10 times out of 10, and that detection was around 90% when the DNA region was handed to the model directly versus as low as 32% when it had to be found through the tool loop. I could not extract the pre-print text myself, so treat that as secondhand, but it was consistent across the coverage I checked.
Read it slowly: the campaign found the thing, and the campaign could not reliably re-find it. That does not bother me the way it seems to bother some people. Stochastic search over a giant space is supposed to look like that. But it changes the discovery headline into something more precise: one lucky pass surfaced a candidate, and the verification chain made it real. If you are buying or selling agent discovery, ask for the rerun distribution, not the highlight.
Is this discovery or expensive search?
The Hacker News thread split exactly where you would expect. One side: "LLM's are giant cross-domain search engines. Not thinking machines. They can discover patterns extremely well. This discovery is well within that space." The other: "isn't a 'conserved, highly transcribed array sitting next to reverse transcriptases' in itself the description of an unknown mechanism?"
Both are right, at different layers. The surface layer was search. The part worth arguing about is the chain that let a search result survive contact with reality. Anthropic's own framing stays modest: "Our work to understand the primary function of ARTs is ongoing." AlphaSignal's summary is the narrowest defensible version I found: "Claude helped Anthropic prioritize an unusual phage locus for laboratory study." Feng Zhang, after reviewing the pre-print, called it "genuinely intriguing" work that "merits further investigation." Meanwhile Startup Fortune reports CRISPR Therapeutics shares dropped more than 4% on the news, which tells you the market read a stronger headline than the science supports.
My position, held loosely: this is tool-assisted hypothesis generation with unusually good verification. I will argue about the word discovery when the function of ART is known. If that work shows the system does something programmable, the discovery framing earns itself back.
What happens when this architecture points at your own stack?
Flip the picture toward the thing I usually write about. Roughly 950 agents reading a giant text database means roughly 950 context windows ingesting text nobody vetted. That is the same shape as your support agent reading tickets, or your coding agent reading issues. In the stateless MCP world a planted prompt is a credential, and the inbox your agents read is the delivery mechanism.
Anthropic's design contains that blast radius with three moves worth copying: calibration against known results before trusting new ones, a report artifact per survivor, and the irreversible step behind a human. Where the architecture looks thinner than the headline: the fan-out itself. A harness coordinating hundreds of parallel tool-using sessions is a lot of access under one operator, and the post does not describe the auth boundaries inside that harness. I would not assume they exist. I would not assume they do not, either.
Which parts would I steal for a pipeline a thousand times smaller?
- Calibrate before you hunt. Reproducing known results first is the cheapest possible trust signal an agent pipeline can emit. If it cannot re-derive what is already known, nothing it says about the unknown deserves weight.
- Fan out cheap, filter brutally. Of 3,500 candidates, 20 got reports and 1 got lab time. Kill in the cheapest tier. My scanner harness philosophy is the same: fixtures that never fire cost less than a wrong alarm.
- Write a report per survivor. Human readable, pointer rich, dated. Reports turn a moment of agent confidence into something another human can audit later.
- Keep the irreversible step human. The lab was the write API. Every agent system I trust has one step the model cannot take alone.
- Publish the miss rate. The rerun detail is the most valuable number in the whole story because it is the least flattering. I tried to hold my own work to this standard by planting 10 vulnerabilities and reporting the 3 a free scanner caught. A pipeline that only shows hits is marketing.
What is still unknown?
The function of ART. Whether it does anything like what the repeat layout hints at. Whether other labs confirm the locus. The pre-print is not peer reviewed. No specific model is named on the page for this work. And the market already moved on a headline the authors themselves hedged.
So here is the question I keep circling: is "the agent discovered" a claim about the model, the harness, or the verification chain that did the actual vetting? My money is on the third. I think that answer decides what we all build next. If you think it is the first or the second, the comments are open and I genuinely want the argument.
Top comments (0)