DEV Community

Udit Raj
Udit Raj

Posted on

GitOCX: What If Your Best Developer Left Tomorrow?

This is a submission for the MLH x DEV Writing Challenge.

Imagine joining a project with 500,000 lines of code.

The repository is healthy.

The tests pass.

The CI pipeline is green.

GitHub contains years of commits.

Everything looks fine.

Then your team lead tells you:

"You need to take over the payment feature."

You open the repository.

There are hundreds of files.
Thousands of commits.
Dozens of contributors.

And somewhere inside all of that history is the answer to:

"Why was this feature built this way?"

Now imagine the developer who originally built most of it is no longer available.

They didn't take the code with them.

->They took the context.

That is the problem that led us to build GitOCX.

What We Built

GitOCX is an AI-powered GitHub knowledge discovery and feature documentation system.

Instead of treating a GitHub repository as just a collection of files, GitOCX looks at its development history to understand how a project evolved.

The idea is:

GitHub Repository
↓
Commits
↓
Commit Messages
↓
AI Feature Classification
↓
Features
↓
Related Commits
↓
Related Files
↓
Source Code + Context
↓
AI Analysis
↓
Feature Documentation

Top comments (5)

Collapse
 
octyn profile image
OCTYN •

"they took the context" is the whole problem in four words. one honest limit worth building into the output: commit history records what changed, rarely why. the why lived in the PR discussion, the Slack thread, the meeting. so docs inferred from commits describe the system as built, not the reasons as decided, and the tool should say which one it's handing you. that provenance line is what makes a new dev trust the docs instead of just reading them.

Collapse
 
razzudit profile image
Udit Raj •

Absolutely agree. That distinction is important — Gitocx can reconstruct what the system became from the commit history, but it shouldn't pretend it knows why those decisions were made when that context lives elsewhere.

I really like the provenance idea. We should explicitly label the documentation as something like “Inferred from repository history” and make the boundary clear: what is directly observed from the code/commits vs. what is inferred. That transparency would make the generated docs much more trustworthy for a new developer.

Collapse
 
octyn profile image
OCTYN •

worth going one step further than labels: put observed and inferred in different files, not just different sections. the observed file can be regenerated on every merge so it never goes stale, the inferred one can't, and the day they share a file the inferred half quietly borrows the observed half's credibility.

Collapse
 
dev-saurabh-k profile image
Saurabh Kumar •

this is what an actually useful product looks like...