AI coding agents are getting very good at navigating repositories.
Ask an agent:
"Where is the API timeout configured?"
It will usually find something useful.
Ask:
"Is the API timeout 60 seconds?"
It will probably search the codebase, find a 60, find a timeout configuration, and give you an answer.
The harder question is:
How do we know the answer is actually supported by the repository?
That distinction is what led me to build Repo X-Ray.
Repo X-Ray is a Claude Code skill for evidence-grounded repository analysis. It checks claims against concrete repository evidence instead of treating the agent's interpretation as the evidence itself.
The underlying engine and evidence protocol are not tied to Claude. The same approach can be adapted to other coding agents and IDE environments.
GitHub: https://github.com/Yudeeswaran/repo-xray
The problem isn't finding code
One of the things coding agents are already good at is finding relevant code.
The problem starts after the search.
Suppose a repository contains:
docs/api.md
timeout: 60 seconds
src/config.py
API_TIMEOUT = 30
tests/test_api.py
assert timeout == 30
An agent can find all three.
But what should the answer be?
The repository doesn't have one single source of truth.
There are at least three different statements:
- the documentation says 60 seconds
- the implementation says 30 seconds
- the tests expect 30 seconds
A normal repository search gives you the files.
It doesn't necessarily give you the contradiction.
And this becomes more important when the repository also contains Git history, deployment configuration, pull requests, issues, and old implementations.
The question changes from:
"Can the agent find something relevant?"
to:
"Can the agent distinguish evidence from interpretation?"
That's the problem I wanted to work on.
Repo X-Ray
Repo X-Ray turns repository analysis into an evidence workflow.
The basic pipeline is:
Requirement / Claim
│
▼
Scope + Plan
│
▼
Deterministic discovery
│
├── source code
├── configuration
├── tests
├── documentation
├── Git history
└── GitHub
│
▼
Evidence ledger
│
▼
Targeted reasoning
│
▼
Deterministic validation
│
▼
VERIFIED / CONTRADICTED /
UNVERIFIED / UNCHECKABLE
The important part is the ordering.
I don't want the model to invent an explanation first and then search for something that supports it.
The repository gets searched first.
Reasoning happens over the collected evidence.
Then the result is validated against that evidence.
Four outcomes instead of "probably"
Repo X-Ray uses four classifications:
| Classification | Meaning |
|---|---|
VERIFIED |
Evidence supports the claim within the recorded scope |
CONTRADICTED |
Evidence contains incompatible behavior, values, scope, or semantics |
UNVERIFIED |
The claim wasn't established within the searched scope |
UNCHECKABLE |
Verification requires unavailable runtime, external, private, or operational evidence |
There is intentionally no ASSUMED.
That was an important design decision.
An agent saying:
"I couldn't find evidence of this."
is not the same thing as:
"This does not exist."
If I search only:
src/orders/
and don't find something, I can say:
No matching evidence was found within
src/orders/.
I cannot honestly conclude:
The repository doesn't implement it.
That distinction sounds small, but it matters a lot when an agent is being used for engineering decisions.
The evidence ledger
The central artifact is an evidence ledger.
Instead of returning only prose, Repo X-Ray records structured information about each claim.
A simplified example:
{
"claim": "API timeout is 60 seconds",
"classification": "CONTRADICTED",
"search_scope": {
"paths": [
"src/",
"tests/",
"docs/"
],
"patterns": [
"timeout",
"60",
"30"
]
},
"risk": "HIGH",
"confidence": 0.95
}
The actual ledger also preserves evidence and provenance.
For repository evidence, that means things like:
src/config.py:18
tests/test_api.py:91
docs/api.md:42
For Git:
commit abc123
For GitHub:
issue #118
PR #201
review/comment reference
This matters because an agent's answer should be reproducible.
If someone asks:
"Why did you say this was contradicted?"
there should be an answer other than:
"The model thought so."
Documentation isn't automatically truth
This is one of the more interesting parts of the system.
Suppose the documentation says:
Maximum upload size: 25 MB
but the application contains:
MAX_BODY_BYTES = 10 * 1024 * 1024
and the tests enforce 10 MB.
The correct result isn't to blindly trust the documentation.
It's also not to blindly trust the implementation and throw the documentation away.
The useful result is:
CONTRADICTED
Documentation:
25 MB
Runtime:
10 MB
Tests:
10 MB
The sources are preserved separately.
That lets the engineer decide what the intended behavior actually should be.
Git history can add another layer.
For example:
60 seconds introduced
↓
implementation changed
↓
change reverted
↓
documentation still says 60 seconds
Now the contradiction has a historical explanation.
GitHub discussions are evidence, not truth
Repo X-Ray can also use GitHub issues and pull requests when available.
But an issue saying:
"We support 25 MB uploads"
doesn't prove that the running code supports 25 MB.
It proves that somebody wrote that statement in an issue.
That's still useful evidence.
But it has different provenance.
I therefore treat sources differently rather than flattening everything into one giant context window.
For example:
Current code/config/tests
↓
implementation evidence
Merged commits
↓
historical implementation evidence
Documentation
↓
declared behavior
Issues / PRs / reviews
↓
engineering discussion
External runtime behavior
↓
requires reproduction
This is particularly useful when a repository has accumulated years of documentation and implementation changes.
The uncomfortable case: UNVERIFIED
One of the things I deliberately wanted to avoid was fake certainty.
Imagine an issue says:
"Requests larger than 10 MB cause the service to run out of memory."
Static repository analysis may find the 10 MB limit.
It may find tests.
It may find configuration.
But it cannot necessarily prove the runtime memory behavior.
That should become:
UNVERIFIED
with something like:
No matching runtime evidence found within the searched scope.
Searched:
- source files
- configuration
- tests
- Git history
- GitHub issue references
Runtime memory behavior was not reproducibly established.
If the evidence genuinely requires a production runtime, external service, or private system, the result can instead be:
UNCHECKABLE
That is preferable to inventing a conclusion.
Requirement changes are another problem
I didn't want Repo X-Ray to stop at:
"The repository is inconsistent."
A more practical question is:
"What happens if the requirement changes?"
Suppose the requirement changes:
10 MB → 50 MB
A naive search might find every occurrence of:
10
MB
50
That produces a lot of noise.
Instead, Repo X-Ray tries to follow the requirement through concrete repository concepts:
Requirement
│
├── identifiers
├── numeric encodings
├── configuration
├── environment variables
├── validation
├── application code
├── tests
├── deployment
├── documentation
├── Git history
└── GitHub discussion
The output distinguishes between two important things.
Direct value
The requested value or a concrete encoding of it is actually present.
For example:
MAX_BODY_BYTES = 10 * 1024 * 1024
That's strong evidence.
Semantic anchor
A location is strongly related to the requirement, but the exact control point isn't proven.
For example:
def validate_request_body(request):
...
This function is clearly relevant to request-size validation, but static evidence alone doesn't prove that this is where the 10 MB limit originates.
That distinction is important.
I don't want a tool that says:
"I found 17 places that must be changed."
when what it really found is:
"I found 17 places that might be related."
Mutation is read-only
Another deliberate decision: Repo X-Ray doesn't automatically rewrite the repository when analyzing a requirement change.
Mutation analysis answers:
"What evidence suggests would be affected?"
It does not mean:
"Go modify all of these files."
For a change such as:
30 seconds → 60 seconds
the system can trace:
runtime configuration
tests
documentation
deployment configuration
Git history
related GitHub discussions
and classify the resulting impact.
This makes mutation analysis useful before an engineer or coding agent actually makes the change.
Risk is separate from truth
Another thing I didn't want to collapse into one number is truth and risk.
A claim can be:
VERIFIED
and still be high risk.
For example:
"Production authentication is enabled."
That might be completely supported by the repository.
But if there is a separate internal route that bypasses authentication, the security implications may still be significant.
Likewise, a small numerical change isn't automatically low risk.
Risk considers things such as:
- severity
- blast radius
- production relevance
- security implications
- configuration impact
- test coverage
- confidence
- dependency surface
- semantic scope
So classification answers:
What does the evidence say?
Risk answers:
How much should we care about it?
Why deterministic first?
This is probably the most important architectural decision in the project.
It's tempting to solve repository analysis by launching many agents:
repository
│
├── agent 1
├── agent 2
├── agent 3
├── agent 4
├── ...
└── agent N
That can work.
It can also become expensive very quickly.
More importantly, adding more agents doesn't automatically give you better evidence.
Repo X-Ray instead starts with deterministic work:
files
patterns
identifiers
values
Git
GitHub
↓
concrete evidence
↓
reasoning
The model is useful for interpreting relationships and ambiguous cases.
It shouldn't have to rediscover the same repository facts independently.
Cost is part of the design
This came from another problem I ran into while looking at agent-based workflows.
If every request turns into a large number of sub-agent calls, the system can become expensive before the user realizes what is happening.
Repo X-Ray therefore has a planner.
Before expensive work, it can estimate:
Files in scope
Claims
Evidence items
Git lookups
GitHub lookups
Expected operations
Reasoning calls
Token range
The intended execution model is:
cheap discovery
↓
find ambiguity
↓
escalate selectively
↓
reason
↓
validate
Not:
launch 50 agents
↓
hope one of them finds the answer
I tested it against deliberately inconsistent repositories
I didn't want the benchmark to be:
"I ran it on my own repository and it looked good."
So I built adversarial fixtures containing contradictions such as:
- documentation says 25 MB, runtime says 10 MB
- documentation says 60-second timeout, runtime says 30 seconds
- retry policy differs between configuration and API behavior
- documentation says idempotency is optional, implementation requires it
- authentication is described as universal while a specific internal route bypasses it
- queue visibility timeout differs from the worker configuration
- PII is documented as being redacted before persistence while code persists the raw object first
- rate limits differ by scope
- documentation describes inventory reservation as atomic while implementation performs separate operations
- currency precision differs between documentation and storage behavior
The benchmark also included decoy information.
For example, an architectural document saying that some internal endpoints may bypass authentication should not cause the system to claim that every internal endpoint is insecure.
That distinction matters.
The final adversarial benchmark detected all 10 expected contradiction domains.
I also tested mutation analysis across several public-source stacks, including Flask, FastAPI, Django, Kubernetes, and Go.
Those tests are useful as generalization checks, but they obviously aren't evidence that the tool works perfectly on every repository.
Repo X-Ray is a Claude Code skill
The project is packaged as a Claude Code plugin/skill.
The Claude-facing instructions live in:
skills/xray/SKILL.md
That skill tells the agent:
- when X-Ray should be used
- how to establish scope
- how to gather evidence
- how to handle uncertainty
- how to interpret Git/GitHub evidence
- how to reason about contradictions
- how to perform requirement mutation
- how to report the evidence ledger
The deterministic implementation sits alongside it.
So there are really two layers:
Claude Code
│
▼
Repo X-Ray Skill
│
▼
Deterministic engine
│
├── repository
├── Git
└── GitHub
The skill is the agent interface.
The Python engine is the deterministic execution layer.
It isn't limited to Claude
Although I built Repo X-Ray first as a Claude Code skill, the underlying idea isn't specific to Claude.
The important pieces are:
evidence model
scope model
classification model
provenance
risk model
mutation protocol
JSON schemas
Those can be consumed by other coding agents and IDE environments.
That's why the repository also contains an agent-oriented protocol and adapter documentation.
The goal is not to create another Claude-only prompt.
The goal is to define a reasonably strict way for coding agents to reason about repository evidence.
What Repo X-Ray does not try to do
There are some boundaries I think are important.
It isn't a compiler.
It isn't a complete dependency graph.
It doesn't prove runtime behavior that can't be observed from the available evidence.
It doesn't magically know which architectural component will be affected by every requirement change.
And a semantic mutation anchor is not the same thing as a confirmed dependency.
Those limitations are intentional.
A tool that says "I don't know" when the evidence isn't sufficient is more useful to me than a tool that produces a confident architecture diagram from weak signals.
The bigger idea
I think we're entering an interesting phase of coding agents.
The first problem was:
Can the agent write code?
Then:
Can the agent understand a codebase?
Now there's another question:
Can the agent distinguish what it knows from what it inferred?
That's the problem Repo X-Ray is trying to address.
The interesting part isn't the phrase "AI repository analysis."
It's the evidence boundary.
Repository
│
▼
What can we actually observe?
│
▼
What does that evidence imply?
│
▼
What remains uncertain?
│
▼
What should the agent do next?
That separation becomes increasingly important as coding agents move from answering questions to making changes.
If an agent is going to modify a production repository, "I found something that looks relevant" isn't quite enough.
I'd rather have:
I found this.
Here's where I found it.
Here's what conflicts with it.
Here's what I couldn't verify.
Here's what Git history says.
Here's what would potentially change.
And here's how confident I am.
That's what I wanted Repo X-Ray to provide.
Try it
Repo X-Ray is open source and available here:
https://github.com/Yudeeswaran/repo-xray
The core workflow is:
repo-xray plan .
repo-xray scan .
repo-xray claim .
repo-xray mutate .
repo-xray validate .
If you're building Claude Code skills, coding-agent workflows, or repository-analysis systems, I'd be particularly interested in feedback on the evidence model and the boundary between deterministic evidence and agent reasoning.
Top comments (0)