Git's idea of a merge conflict is narrower than it should be. Two branches change the same lines, you get a conflict. Two branches touch the same function from different files and never collide, git merges cleanly and congratulates you. Then production breaks.
That's the merge I've learned to fear most: the clean one. One person refactors a function while another adds a call to it, the merge goes through without a peep, and you find out a week later from a bug report.
I'm building PRISM, an AI-powered git time machine, and this week I shipped a conflict predictor built around a simple idea: lines are the wrong unit. Symbols are better.
The approach
Parse the merge-base and both branch heads into ASTs with tree-sitter, extract the symbols (functions, classes, methods), and flag overlap between the two sides. Both branches touching the same symbol is risk, even across different files. The score saturates so one hot file can't dominate the whole prediction. It's exposed as POST /repos/{id}/predict-conflict, and there's a branch-compare UI with a risk bar and the overlapping symbols listed, so the number is inspectable instead of magic.
The unglamorous parts took most of the time
Tree-sitter grammar packages aren't always installed on the machine. The first version crashed when one was missing, which made it a demo, not a tool. Now the service degrades to line-level scoring when a grammar is missing. Boring engineering, but it's the difference between something you can actually run and something you can't.
Method detection inside class bodies broke twice before my tree-sitter queries finally caught it. Queries for methods are fiddlier than function queries, and I learned that through two rounds of debugging. Skill issue on my part, honestly.
9 new tests, 39 passing overall.
What it's actually for
The predictor doesn't replace tests or review. It answers one narrow question: did these two branches step on the same symbols? That's cheap to answer, and it catches exactly the class of breakage that git itself can't see. Next up I'm working on semantic code ownership with embeddings, which should catch the cases where even the symbol names diverge.
If you merge a lot, you know the feeling of a clean merge that wasn't. What's the worst one you've shipped?
Top comments (0)