
Sentinel mr-report 0.2.0 URL adress: https://gitlab.com/jakubrojicek11/sentinel-mr-report/-/tree/v0.2.0?ref_type=tags
Sentinel MR Report 0.2.0: Before AI Reviews Your Code, Ask What Changed
Let's start with a deliberately uncomfortable question:
If an AI agent changes your code, who checks what the code is now capable of doing?
Not whether the code looks reasonable.
Not whether the AI says the tests passed.
Not whether the pull request description says:
«"Implemented requested security improvements."»
I mean something much simpler:
What can this code do now that it could not do before?
That question is the reason "sentinel-mr-report" exists.
Version 0.2.0 is now released.
It is small. It is JavaScript-only. It does not use an LLM. It does not need an account. It does not send your source code to a remote service.
It simply looks at the old and new versions of changed JavaScript files, extracts facts from the syntax tree, and tells you how the capability or risk surface changed.
And if you want it to, it can fail your GitLab pipeline when a dangerous capability appears.
That's the boring version.
The interesting version starts when you ask:
Why should we make an AI rediscover facts that a deterministic program can establish directly?
First, forget AI for five minutes
Imagine you have a small web application.
Yesterday, one of your JavaScript files could:
- read a configuration value,
- read a file,
- make an HTTP request.
Today someone opens a merge request.
The diff adds:
const { exec } = require("child_process");
exec(command);
A human reviewer can obviously notice that.
But now imagine the change is 1,500 lines long.
Or the developer is an AI coding agent.
Or there are 30 merge requests a day.
Or the dangerous capability isn't sitting next to a big red comment saying:
// WARNING: I CAN NOW RUN SHELL COMMANDS
The important change can be buried inside otherwise boring code.
And this is where I think we often ask AI to do something slightly backwards.
We give the model thousands of lines of source code and say:
«"Please understand all of this and tell me whether anything important changed."»
The model then has to discover the facts before it can reason about them.
But some facts don't need reasoning.
If the syntax tree establishes that a file gained a call to "child_process", we don't need a language model to have an existential crisis about whether "child_process" exists.
We can just check.
Think of it as a witness, not another reviewer
This is the mental model I find easiest.
Imagine a code review with three people.
The first person says:
«"I think this change adds process execution."»
The second says:
«"I agree."»
The third says:
«"I disagree. I think it's just a function named "exec"."»
Now everybody starts reading the code.
That's a terrible way to spend everyone's afternoon if the question can be answered mechanically.
Instead, imagine there is a small machine sitting next to the review:
«Before: no process execution
After: process execution
Evidence: "child_process.exec"
File: "runner.js"»
Now the humans can argue about what the change means.
The machine already answered what it observed.
That distinction is the whole idea.
What does Sentinel MR Report actually do?
"sentinel-mr-report" is a small GitLab CI job for JavaScript repositories.
For each changed ".js" file, it can compare the old and new versions directly from the Git objects available inside the CI runner.
It uses Tree-sitter to parse the code and extracts structured facts.
For example:
Before
filesystem read
network access
becomes:
After
filesystem read
network access
child_process
process spawning
The report can then say:
Risk raised: low → critical
Dangerous operations introduced:
- child_process
- process_spawn
That's much more useful to a reviewer than:
«"Something security-related may have changed somewhere in this enormous diff."»
The report is also machine-readable, so CI can make a decision from it.
And no, the AI is not doing the checking
This is important.
The MR Report itself uses:
- Git
- Tree-sitter
- deterministic fact extraction
- ordinary program logic
It does not need:
- an LLM
- an OpenAI API key
- an external analysis service
- telemetry
- a Sentinel account
The analysis happens inside your CI environment.
That matters for two reasons.
First, privacy.
Your source code does not need to be uploaded somewhere just to answer a basic question about what changed.
Second, predictability.
If the same input is processed twice, the fact extraction does not suddenly wake up in a different mood.
There is no:
«"Yesterday the model thought "exec()" looked suspicious, today it feels pretty chill about it."»
The same input produces the same result.
That's a useful property when the output can actually block a merge.
So where does AI fit?
This is where the project gets more interesting.
I am not trying to remove AI from software development.
Quite the opposite.
I want AI to reason about the things that actually require reasoning.
Suppose the deterministic layer says:
{
"child_process": true,
"process_spawn": true,
"filesystem_read": true,
"network_access": true
}
An AI can now reason about that.
For example:
«"The new process execution is used by the deployment helper and receives only a fixed argument."»
Or:
«"The new network request sends an environment-derived value to an external host."»
Those are reasoning questions.
But the basic observation:
«"This file gained process spawning."»
doesn't need an AI.
That's the boundary I'm interested in.
Machine establishes facts.
AI reasons about facts.
Humans decide what is acceptable.
At least, that's the direction.
The first version had a hole
And this is where 0.2.0 becomes more interesting than just "we added a feature."
Version 0.1.0 had a simple gate.
If the risk level of a file increased to the configured threshold, the gate could fail the pipeline.
For example:
low → critical
↓
BLOCK
That sounds sensible.
Until you think about what happens when a file is already critical.
Suppose the file already contains process execution.
Its state is:
critical
Now someone adds another dangerous operation.
The risk level is still:
critical
So numerically:
critical → critical
No increase.
The old gate could therefore say:
«"Nothing happened here."»
Meanwhile the report itself could see:
Dangerous operations introduced:
process_spawn
The observation was correct.
The gate simply wasn't looking at the right thing.
We found that by trying to break it
This part matters to me more than the fix.
We didn't discover the problem because a production incident happened.
We asked an annoying question:
«"What happens if a capability is introduced gradually across multiple merge requests?"»
Then we wrote synthetic before/after cases and pinned the answers in tests.
That produced a much more interesting result.
The obvious scenario was actually already caught.
If the first merge request adds:
require("child_process")
the file becomes critical immediately.
So the first merge request is blocked.
The interesting hole was the next step.
The file is already critical.
Then another merge request adds:
exec(command)
The report sees the new process-spawning capability.
But the risk level remains critical.
So the old gate sees:
critical → critical
and does not block.
That was the real gap.
There was another problem hiding next to it
The old acknowledgement model was also too broad.
Imagine this:
{
"acknowledged": {
"scripts/deploy.js": "This file intentionally runs Terraform."
}
}
That makes sense.
Terraform needs process execution.
The repository owner has reviewed it.
Fine.
But what happens if somebody later adds:
eval(userInput);
to the same file?
The acknowledgement says:
«"This file runs Terraform."»
It does not say:
«"This file may acquire every dangerous capability imaginable for the rest of eternity."»
Yet a path-level acknowledgement could effectively behave that way.
That is a granularity problem.
The gate knows the file.
But it also needs to know which capability was intentionally accepted.
This is what 0.2.0 changes
Version 0.2.0 changes the gate in two important ways.
- A new dangerous capability can trigger the gate
The gate no longer asks only:
«"Did the numerical risk level increase?"»
It asks:
«"Did the risk increase, or did this file gain a new dangerous capability?"»
Conceptually:
gated =
risk_is_raised
OR
dangerous_fact_was_added
But only when the resulting risk is at or above the configured threshold.
So:
critical → critical
+
new process_spawn
BLOCK
while:
critical → critical
+
rename a function
PASS
That's an important distinction.
A file that legitimately uses process execution should not become permanently frozen.
The gate should care about new capability, not simply the existence of an old capability.
- Acknowledgements can name the capability
The new form looks like this:
{
"acknowledged": {
"scripts/deploy.js": {
"reason": "shells out to terraform; reviewed 2026-08",
"facts": [
"child_process",
"process_spawn"
]
}
}
}
Now the repository is saying something much more precise:
«"We understand that this file uses these capabilities and we intentionally accept them."»
If somebody later adds:
eval
that is a different capability.
It is not silently covered by the Terraform acknowledgement.
The gate can block it.
What if you already use the old format?
We didn't make everybody rewrite their repository configuration overnight.
The old form still works:
{
"acknowledged": {
"scripts/deploy.js": "shells out to terraform"
}
}
That's important for backwards compatibility.
But there is a trade-off.
A legacy acknowledgement is broad.
So 0.2.0 makes the waiver visible in the report, including the dangerous facts it absorbed.
In other words:
legacy configuration still works, but silent acceptance becomes harder to hide.
I like that compromise.
Why "capability" instead of just "risk"?
This is probably the most important conceptual part.
Imagine two files.
File A:
risk: critical
capabilities:
- process_spawn
File B:
risk: critical
capabilities:
- process_spawn
- eval
If you only look at the risk label, they are identical:
critical
But they are not identical systems.
File B can do something File A could not.
The number saturated.
The capability surface did not.
That's why a security gate that only watches the number can miss meaningful changes once the number hits its ceiling.
The interesting thing isn't always:
«"Did the score go up?"»
Sometimes it is:
«"What new thing can this code do?"»
What counts as a dangerous operation?
The current project is deliberately small.
Among the dangerous fact families are:
eval
child_process
process_spawn
process.exit
destructive_fs
The exact report can also expose less severe facts such as filesystem, network, environment and route changes.
The idea is not that every one of these things is automatically malicious.
"child_process" is not malware.
"fs.unlinkSync()" is not malware.
"process.env" is not malware.
Sometimes your application genuinely needs them.
The point is visibility.
If a web handler suddenly gains the ability to execute operating-system commands, that deserves a look.
If a deployment script already does exactly that because it has to call Terraform, perhaps it doesn't.
Context determines whether a capability is acceptable.
The fact layer's job is to tell you that the capability changed.
This is not a replacement for security scanners
It is worth being very clear about this.
There are already mature security tools.
Semgrep, for example, can run blocking rules in CI and use exit code 1 to prevent a merge when configured findings are detected. That is a well-established pattern: finding → policy → pipeline result.
"sentinel-mr-report" is much narrower.
It isn't trying to become Semgrep.
It asks a different question:
«What capabilities were introduced by this particular change?»
That makes it closer to a change-aware capability diff than a general-purpose vulnerability scanner.
You might use both.
You might use neither.
The useful experiment is to put the tiny thing in your CI and see what happens.
Why this matters even more with AI coding agents
Now we get to the original reason I built this.
AI coding agents are increasingly capable of:
- reading repositories,
- editing files,
- running tests,
- calling tools,
- opening merge requests,
- changing configuration,
- interacting with infrastructure.
That changes the question.
When a human developer writes:
exec(command);
there is at least a human decision somewhere in the loop.
When an autonomous coding agent writes it, the agent may have reached that solution because:
- it interpreted a ticket,
- followed a prompt,
- copied a pattern,
- reacted to an error,
- optimized for passing tests,
- or simply found a shortcut.
The model does not need to be malicious for the result to be dangerous.
It just needs to produce a capability that nobody explicitly approved.
That is where deterministic observation becomes useful.
The model can propose.
The fact layer can observe.
The gate can enforce.
And this is where supply-chain attacks become interesting
Consider what malicious JavaScript actually wants to do.
A recent Socket Threat Research report on the SANDWORM_MODE campaign describes malware that harvested credentials and CI environment secrets, used network exfiltration, interacted with GitHub credentials, modified repositories and workflows, and targeted AI development tooling.
The exact attack is obviously much more sophisticated than this tiny project.
But look at the underlying capabilities:
read environment
+
read files
+
spawn processes
+
network communication
+
write files
+
modify repository/workflow state
Those are capabilities.
A capability diff is not a complete defense against a supply-chain attack.
It won't magically recognize every malicious package.
It won't prove intent.
It won't replace dependency security tooling.
But it gives you another question at a very useful boundary:
«Did this change suddenly give this repository a capability it did not have before?»
That's cheap to ask.
And sometimes cheap questions are the ones worth asking everywhere.
There is academic work pointing in the same direction
This isn't an idea invented in a vacuum.
A 2026 paper from Aarhus University, Defensive Capability Analysis for JavaScript Libraries, studies capability analysis for JavaScript packages and evaluates which security-sensitive capabilities packages exercise.
One result is particularly interesting for this project: after deduplication, the authors report that at least 72.9% of the npm packages in their benign dataset used no security-sensitive capabilities at all.
That doesn't mean "72.9% of npm is safe."
It doesn't mean capability analysis solves supply-chain security.
It means something much simpler:
A lot of software may not need complicated security reasoning before you can establish that its capability surface is relatively small.
That is exactly the space where a deterministic fact layer can be useful.
The part I really want people to test
Here's where I don't want this article to become another:
«"Look at my cool open-source project!"»
Please don't just read this and tell me it sounds interesting.
That gives me almost no useful information.
I want you to install it.
Preferably on something boring.
Not your production crown jewels.
Pick a small JavaScript repository.
Add the GitLab CI job.
Run it without blocking for a while.
Look at what it reports.
Then ask:
- Did it identify the changes correctly?
- Did it miss something obvious?
- Did it flag something that wasn't actually a capability?
- Does the report make sense to a human reviewer?
- Is the risk threshold useful?
- Would you actually want this to block a merge?
- Does the acknowledgement model make sense?
- What happens when your repository has an unusual pattern?
And then do something even more useful.
Try to break it.
Add a weird binding.
Rename a function.
Wrap a dangerous operation.
Move code between files.
Split a change across multiple merge requests.
Put something legitimate behind an acknowledgement.
Try to create a false positive.
Try to create a false negative.
If it fails, tell me.
That's much more valuable than a star.
You don't even have to trust the project
This is one of the reasons I like the current shape of the tool.
You can inspect it.
The current implementation is only a small set of source files.
You can run the tests locally.
You can see the GitLab CI configuration.
You can see the fact extraction.
You can see the gate.
There isn't a mysterious cloud service sitting behind it saying:
«"Trust me, bro. The AI security oracle has spoken."»
The interesting property is that the thing making the basic gate decision is visible code.
If the gate is wrong, we should be able to reproduce why.
And if we can't reproduce it, that's a problem.
One experiment I would especially like to see
Here's a small challenge for anyone running coding agents.
Take a repository where an AI agent is allowed to modify JavaScript.
Run Sentinel MR Report before the agent starts.
Then let the agent implement a normal task.
Don't tell the agent that the capability gate exists.
When the agent opens the MR, look at the report.
Ask:
«Did the agent introduce a capability that wasn't part of the task?»
For example:
The task is:
«"Add a health-check endpoint."»
The resulting code suddenly gains:
network access
filesystem access
process spawning
Maybe there is a perfectly legitimate explanation.
Maybe there isn't.
Either way, the fact layer has done something useful.
It has changed the review question from:
«"Can we understand everything the AI changed?"»
to:
«"Why did this change gain these capabilities?"»
That's a much smaller question.
The gate is deliberately not omniscient
This is where 0.2.0 draws a line.
There are still known boundaries.
Threshold creep
Suppose you configure:
--gate critical
and the application evolves like this:
MR 1:
low → medium
MR 2:
medium → high
Neither reaches critical.
The gate doesn't block.
That is not an invisible bug.
It is what the configured threshold means.
If you want high-risk changes blocked, use:
--gate high
The reports still show the intermediate changes.
The policy is yours.
Cross-file reasoning is another boundary
Suppose:
a.js
gains:
spawnSync("make");
and:
b.js
later gains:
a.build();
The current analysis is per file.
It can see the dangerous capability in "a.js".
It does not build a complete call graph and infer that "b.js" can now indirectly cause a process spawn.
That is deliberately out of scope.
And I'm actually happy that this is written down.
A security tool saying:
«"We don't currently perform cross-file call-graph analysis."»
is much more useful than quietly pretending that it does.
JavaScript only, for now
The current parser focuses on JavaScript.
TypeScript and Flow-annotated files that cannot be handled by the current parser are reported rather than guessed at.
Again, this is intentional.
A deterministic system should be willing to say:
«I don't know.»
That sentence is surprisingly valuable in an industry where software sometimes prefers producing an answer to producing an accurate answer.
What about false positives?
They will happen.
And I actually want them.
Not because false positives are good.
They're not.
But because an early open-source project should expose them rather than hide them behind a marketing page.
If the tool says:
process_spawn
because you wrote:
exec();
but that "exec" is actually your own local function, that's a bug in binding resolution.
That's useful.
Now we have:
input
→ wrong fact
→ reproducible case
→ test
→ fix
Compare that with an LLM saying:
«"This code appears potentially suspicious."»
Good luck writing a regression test for that sentence.
The deeper idea: facts before reasoning
This project started from a broader experiment around autonomous coding agents.
The more I worked with agents, the more I kept running into the same question:
What should an AI be allowed to believe?
If the model says:
«"This code does X."»
why should that statement automatically become the truth?
Maybe it is right.
Maybe it misunderstood the code.
Maybe the context was incomplete.
Maybe an earlier agent wrote something misleading into its history.
Maybe the model simply made a mistake.
Instead, we can sometimes establish a smaller set of facts independently:
This import exists.
This process operation exists.
This network call exists.
This environment value is read.
This route was added.
This capability did not exist in the previous revision.
Then the AI can reason about those facts.
This is the same general idea behind the larger Sentinel-IR experiment:
«Don't make the AI rediscover facts that a deterministic program can establish reliably.»
In our live Sentinel-IR experiments, that approach produced a measured reduction of more than 70% in input tokens in one tested workflow.
That's a benchmark observation, not a universal promise.
But the architectural idea doesn't depend on token savings.
Even if the model cost were zero, deterministic facts would still be useful because they give the reasoning system a separate source of evidence.
So is Sentinel MR Report an AI security tool?
Not really.
And that's deliberate.
It is a deterministic code-change observer with an optional CI gate.
The larger Sentinel project uses these facts as part of an AI governance system.
The standalone MR Report doesn't need any AI at all.
That distinction matters.
You can install it even if your organization doesn't want AI touching source code.
Or you can put it in front of an AI coding agent.
The second use case is where I think things get particularly interesting.
Because then you get something like:
AI agent
↓
writes code
↓
opens MR
↓
Sentinel MR Report
↓
"What changed?"
↓
deterministic facts
↓
policy / human review
↓
merge
The AI doesn't get the final vote simply because it produced a convincing explanation.
Start in observation mode
If you're thinking:
«"There is absolutely no way I'm putting some random open-source security gate in front of production merges."»
Good.
Don't.
Start with reporting.
Install the CI job.
Run it for a week.
Don't block anything.
Collect reports.
See what it thinks changed.
Compare those reports with actual code review.
If it produces garbage, remove it.
If it catches things you care about, turn on the gate for a narrow threshold.
This is the safest way to test an early tool.
You don't need to believe the author.
You need to measure the tool against your own repository.
The install is intentionally boring
Add the GitLab CI template:
include:
- remote: 'https://gitlab.com/jakubrojicek11/sentinel-mr-report/-/raw/main/ci/mr-report.gitlab-ci.yml'
Pin the release:
MR_REPORT_REF: v0.2.0
Start without:
--gate
Read the reports.
Then, if the results are useful, enable something like:
node scripts/mr-report.js \
--base origin/main \
--gate critical
Now the process becomes:
No new critical capability
↓
PASS
New critical capability
↓
FAIL
And if the capability is intentional, document it in the repository acknowledgement.
The policy becomes visible.
Why I want people to try this before it gets bigger
There's a temptation with projects like this.
You get an idea.
Then you add:
- TypeScript
- Python
- Java
- call graphs
- dependency graphs
- cloud dashboards
- AI explanations
- automatic remediation
- enterprise policy management
- 47 configuration files
- a Kubernetes operator
- a blockchain for some reason
And six months later you have built a spaceship to answer:
«"Did this file start spawning processes?"»
I would rather avoid that.
For now, I want the small thing to become trustworthy at the thing it actually claims to do.
That means more repositories.
More weird code.
More false positives.
More false negatives.
More uncomfortable GitLab issues.
Less PowerPoint.
Version 0.2.0 in one picture
The simplest way I can describe the change is this:
0.1.0
Did risk increase?
YES
↓
BLOCK
NO
↓
PASS
0.2.0
Did risk increase?
│
├── YES ───────────────┐
│ │
NO │
│ │
Did a new dangerous │
capability appear? │
│ │
├── YES ───────────────┤
│ │
NO │
│ │
PASS ↓
Is risk at
or above threshold?
│
┌────┴────┐
YES NO
↓ ↓
BLOCK PASS
And the acknowledgement model changed from:
"I trust this file."
to:
"I intentionally accept these capabilities in this file."
That's a much more useful contract.
One last uncomfortable question
If an AI coding agent opens a merge request and tells you:
«"I've completed the task. All changes are safe."»
What evidence do you actually have?
Maybe the answer is:
«"The tests passed."»
Good.
But tests don't necessarily tell you whether a file gained a capability you didn't intend.
Maybe you also have an AI code review.
Also good.
But now you're asking another probabilistic system to validate the first system.
Maybe you have static analysis.
Better.
But many static analyzers answer a different question, such as whether a known rule or vulnerability pattern exists.
What if the question is simply:
«"What can this changed code do now that it could not do before?"»
That's the question Sentinel MR Report is trying to make cheap.
Not perfect.
Not comprehensive.
Not magical.
Just cheap, deterministic, visible and testable.
Try it and break it
The repository is still an early open-source project.
That is exactly why I want people to test it now.
If you are a developer, security engineer, DevOps engineer, or someone experimenting with coding agents, try it on a repository you understand.
Run it without blocking first.
Look at the output.
Then deliberately try to make it wrong.
If it misses something, I want to know.
If it flags something incorrectly, I want to know.
If the gate policy doesn't make sense for your repository, I want to know.
If you think the entire idea is unnecessary, tell me why.
The most useful outcome isn't:
«"Nice project."»
It's:
«"I tried this and here's where it broke."»
That's how version 0.3.0 should happen.
Not because we invented another feature.
Because somebody tried to break 0.2.0.
Final thought
AI is getting very good at producing software.
That makes one particular question more important, not less:
Who independently checks what the software became?
Maybe the answer in your organization is a human reviewer.
Maybe it's a mature security platform.
Maybe it's several layers.
Maybe it's nothing yet.
"sentinel-mr-report" is an experiment with one small additional layer:
before asking an AI to reason about a change, establish the mechanical facts about that change first.
Let the machine observe.
Let the policy decide.
Let the AI reason where reasoning is actually needed.
And if you think that sounds like a terrible idea, there's a very convenient way to prove it.
Install it.
Try to break it.
I would genuinely rather receive a nasty bug report than another polite star.
Project
"sentinel-mr-report" is open source under the MIT license.
Current release: v0.2.0
The project is intentionally small and experimental.
JavaScript only.
Per-file analysis.
No call graph.
No LLM.
No telemetry.
No remote source-code analysis.
If you try it, please report the false positives and false negatives.
Especially the embarrassing ones.
Those are usually the interesting ones.
URL Adress :
Further reading — the same problem from three angles
Semgrep — "Handling blocking findings and errors in CI"
https://docs.semgrep.dev/semgrep-ci/configuring-blocking-and-errors-in-ci
The reference design for rule → finding → exit code 1 → MR blocked, and the Monitor / Comment / Block modes that let you phase a gate in. sentinel-mr-report is deliberately a tiny subset of this idea: one language, capability deltas only, no platform, no account.Socket Threat Research — "SANDWORM_MODE: Shai-Hulud-style npm worm hijacks CI workflows and poisons AI toolchains"
https://socket.dev/blog/sandworm-mode-npm-worm-ai-toolchain-poisoning
A current, concrete picture of what malicious JavaScript actually does: harvest env secrets, spawn processes, exfiltrate over HTTPS/DNS, inject workflows, persist via git hooks. Every step is one of the capability families a per-MR fact diff reports.Xu & Møller (Aarhus University) — "Defensive Capability Analysis for JavaScript Libraries", ASE 2026
https://www.cs.au.dk/~amoeller/papers/capa/paper.pdf
The academic case for capability analysis as the first stage of every supply-chain tool, and for soundness over heuristics: tracking which code can reach fs, child_process, process.env and friends. Their finding that at least 72.9 % of npm packages use no security-sensitive capability at all is the same intuition behind gating on new capabilities: most changes should produce an empty report.
Honourable mention: Danger JS (https://danger.systems/js/) — the general-purpose "codify review chores in CI" tool. It is where you would put your rules; sentinel-mr-report is one specific rule, done carefully.
Top comments (0)