My coding agent can write a feature, add tests, and hand me a green build. It can also put the right behavior in the wrong place.
The application works. The imports look reasonable. But an application service has quietly learned how the database stores and replaces rows. That is the kind of review comment I want to turn into useful automated feedback.
I built Jev CI to explore this. It compares Git branches, evaluates changes against policies written in YAML, and lets Jev request more evidence when the diff does not answer the review question.
Here is the example from my short video, followed by the mechanism and commands to run it yourself.
Two implementations of the same feature
The feature is small: replace an item's ordered list of labels. I built a demo repository with my coding agent and opened two pull requests against main.
Both branches pass the demo's six checks: five behavior tests and one import check. That import check inspects the Python AST and verifies that the service imports only LabelStore from src.ports. No direct database import slips into the service.
In PR 1, the application service coordinates the replacement:
def set_labels(self, item_id: str, labels: list[str]) -> None:
self.store.delete_labels(item_id)
for position, label in enumerate(labels):
self.store.insert_label(item_id, position, label)
self.store.flush()
The feature works. But the service now knows that replacement means deleting rows, assigning positions, inserting the new values, and flushing the changes. If storage needs a different replacement procedure later, that knowledge is sitting in the caller.
In PR 2, the service uses a single operation:
def set_labels(self, item_id: str, labels: list[str]) -> None:
self.store.replace_labels(item_id, labels)
The storage adapter owns the sequence:
def replace_labels(self, item_id, labels):
self.delete_labels(item_id)
for position, label in enumerate(labels):
self.insert_label(item_id, position, label)
self.flush()
The interface already defines replace_labels as persisting the item's complete ordered label collection. The second implementation uses that contract. The first reaches for the lower-level operations and reconstructs the procedure itself.
This is a limitation of the checks in this example, not a claim that deterministic tools could never detect the pattern. A custom test or static rule could forbid these calls. The more general review question is harder to encode: does this caller depend on another component's internal procedure, or is it using a supported contract? Answering that can require interpreting both sides of the boundary.
Define the review policy
I want teams to express their own conventions, so the tool starts with a policy rather than a universal definition of good architecture.
The demo's responsibility boundary policy includes this statement and interpretation. This is an excerpt; the repository contains the complete policy and configuration.
id: responsibility-boundary-leakage
statement: >-
A responsibility should not depend on another responsibility's
internal procedure or representation without an evidenced need.
interpretation:
definitions:
- >-
A business service that must sequence delete, positional insert,
and flush calls to replace a collection may depend on storage
procedure; a single provider-owned replace operation may preserve
the boundary.
The policy also defines scope, exceptions, and permitted evidence requests. Its exceptions matter: a supported stable contract can be a perfectly valid dependency. Merely calling another component should not produce a violation.
The wording explicitly covers this kind of storage sequence. The evaluator still has to connect the changed code to the interface and implementation evidence. These two examples demonstrate that workflow; they do not establish a general detection rate.
In the recorded runs, Jev CI reported violation for storage-leak and compliant for storage-owned. The policy is configured to report findings, so a violation here means a review finding, not a blocked merge.
The HTML report connects findings to the changed code, selected evidence, and evaluation rounds. I can inspect what was available to the evaluator instead of receiving only a verdict.
Give the evaluator a way to ask for evidence
A diff can show a method call while omitting its contract or implementation. Sending just that diff leaves the reviewer guessing. Sending the entire repository up front is often unnecessary and quickly consumes the available context.
Jev CI starts with the committed branch comparison. It resolves the source and target refs, compares the source with their merge base, and splits the complete patch into bounded chunks. Each relevant chunk is evaluated against the applicable policies, followed by reconciliation where the policy requires it.
Within one evaluation, the controller keeps the current evidence and builds a menu of concrete retrieval requests. Jev can select from that menu when it needs more context.
Each call batches three typed questions:
| Question | What Jev selects |
|---|---|
disposition |
How the change relates to the policy, including whether more evidence is needed. |
next_request |
One candidate ID from the offered request menu, when a menu exists. |
support |
An ID for already delivered source evidence, or none. |
If an accepted verdict finishes the evaluation, the controller ignores the speculative next request from that same response. Otherwise, it executes the selected request, records the result, removes the used request from the menu, and asks again with the additional evidence.
The important boundary is that the model selects IDs. The controller maps those IDs back to requests it constructed. It does not execute model-written shell commands or arbitrary file paths.
The demo uses the git-exact provider: committed files, bounded literal searches, diff chunks, and trusted documents. Optional Ripwire caller/callee retrieval exists, but it was not enabled for these runs. The loop diagram illustrates the mechanism, rather than replaying one particular run's request sequence.
The loop has limits. It checks serialized input size, call count, deadlines, configured follow-up rounds, and the remaining request inventory. If it cannot resolve the question, that outcome stays visible. It does not keep asking until it gets the answer I want.
The feedback text shown in a finding comes from the policy's authored templates. Jev chooses the assessment and supporting evidence; it does not write a free-form explanation. The evaluation protocol describes those transitions in detail.
Run the same branch comparison
You need Git, Python 3.12 or newer, and an AI_GATEWAY_API_KEY for the hosted Jev integration. These commands work from a fresh directory on macOS or Linux.
First install the tool and clone the example repository:
mkdir jev-ci-lab
cd jev-ci-lab
git clone https://github.com/krisitown/jev-quality-gate.git
python3 -m venv jev-quality-gate/.venv
jev-quality-gate/.venv/bin/python -m pip install \
-e ./jev-quality-gate -c ./jev-quality-gate/constraints.txt
git clone --branch main https://github.com/krisitown/jev-ci-video-demo.git
git -C jev-ci-video-demo fetch origin
git -C jev-ci-video-demo worktree add --detach ../jev-controls origin/main
JEV="$PWD/jev-quality-gate/.venv/bin/jev-ci"
REPO="$PWD/jev-ci-video-demo"
CONTROLS="$PWD/jev-controls/.jev-ci/config.json"
KEY_FILE="$PWD/jev-quality-gate/.env.local"
cp jev-quality-gate/.env.example "$KEY_FILE"
chmod 600 "$KEY_FILE"
Edit .env.local and set AI_GATEWAY_API_KEY to your own key. Keep it out of version control. The integration sends selected code evidence to hosted Jev through Vercel AI Gateway, so use a project you are comfortable sending to that service.
The controls above come from a separate checkout of main. This keeps the evaluated branch from supplying its own review rules.
Validate the controls without making an inference request, then evaluate both branches into fresh output directories:
"$JEV" validate --config "$CONTROLS"
OUT="$(mktemp -d "${TMPDIR:-/tmp}/jev-ci-demo.XXXXXX")"
"$JEV" evaluate \
--repo "$REPO" \
--source origin/storage-leak \
--target origin/main \
--config "$CONTROLS" \
--env-file "$KEY_FILE" \
--output "$OUT/storage-leak"
"$JEV" evaluate \
--repo "$REPO" \
--source origin/storage-owned \
--target origin/main \
--config "$CONTROLS" \
--env-file "$KEY_FILE" \
--output "$OUT/storage-owned"
Open $OUT/storage-leak/report.html and $OUT/storage-owned/report.html in your browser. On macOS:
open "$OUT/storage-leak/report.html" "$OUT/storage-owned/report.html"
The tool evaluates committed changes; editing a file without committing it does not change this comparison. Live model outcomes can vary, so inspect your report rather than treating the recorded verdicts as guaranteed assertions. For reproducible automation, use reviewed commit IDs and the repository's GitHub Actions example.
Help turn recurring review comments into policies
I am running a separate broader evaluation of this approach. For now, these two PRs show how the tool behaves on a small example. Whether the additional feedback reliably improves engineering outcomes needs more evidence.
The contribution I would love to see is a policy built around something your team keeps discussing in review, accompanied by examples that should pass and examples that should fail. Architecture and maintainability feel like useful starting points for shared bundles.
Useful counterexamples, minimal reproductions of incorrect findings, and fixes to the tool are welcome too.
Try Jev CI, inspect the two demo pull requests, and tell me which rule you would try first.







Top comments (0)