Every repository has a moment nobody talks about.
Not the first commit. Not the launch. Not the sprint where everything clicked. The moment I keep coming back to — the one that made me build this thing — is the gap. The silence. The chapter where nothing happened and everyone pretended it didn't exist.
When I ran Commit Canvas over Flask's history, it found six chapters.
Five were predictable. The sixth stopped me:
The Silence: 31 days of nothing. November 28, 2024 → January 1, 2025.
A changelog treats that gap as an absence. A story treats it as a beat.
That's everything.
The Problem Nobody Names
You've shipped something. Maybe it took two years. Maybe six. There were nights you pushed at 2am, weeks where the repo barely breathed, a sprint that felt like you were running straight into a wall and somehow came out the other side.
And the artifact of all of that is... a commit graph. Some green squares. A number.
git log is honest. It's complete. It is also about as narrative as a shipping manifest.
GitHub Insights tells you how much. Gource makes a beautiful video of your file tree growing — genuinely hypnotic, fifteen years mature, and about as useful for understanding what happened as watching a timelapse of a city and trying to read its history.
The data for a real story is already sitting in your repo: commit dates, authors, gaps, tags, merge topology, churn. What's missing is something that reads it as a sequence of beats instead of a pile of rows.
So I built the reader.
What It Is
Commit Canvas reads a git repository and generates one thing: a single, self-contained HTML file. Chapters. A scrubbable time machine. A project fingerprint. An export studio. All of it rendered from your real history.
git clone https://github.com/ahmadrrrtx/commit-canvas
cd commit-canvas
./run.sh /path/to/your/repo --theme sunset --open
Python 3.8+ and the git CLI. Zero third-party dependencies. The output is one file — roughly 160 KB. Email it. Commit it to the repo. Put it on a USB stick. Open it offline in ten years. There is nothing left to fetch.
No server processes anything. No account. No API key. No analytics. The analysis never leaves your machine.
There's also a web version — paste a public GitHub URL and the same engine runs entirely in your browser tab. And four live demos you can open right now: Flask · Express · Requests · the project's own story.
The Engine Is Honest by Design
The entire analysis runs on two subprocesses. One git log call with a custom delimiter-separated format — fields split by \x1f because commit messages can contain pipes, and \x1f is the one character git promises never to emit inside them. One git for-each-ref for release dates. Everything after that is pure Python.
Out comes a monthly series: commits per month, active days, merges, new contributors, net lines. Almost every story decision flows from that series.
Here is the part I want you to read carefully:
There is no message parsing. No AI. No LLM anywhere in the pipeline. No model reading your commit messages and deciding what your project "means." Just arithmetic over dates — which is exactly why the results are defensible. When someone asks "why does my repo say The Sprint?" you can open the rule and count the thresholds on one screen.
if gaps >= 60 and comeback_strength >= max(3, median * 0.5):
label, arc = "The Return", "return"
elif peak >= 4 * median and peak >= 15:
label, arc = "The Sprint", "sprint"
elif late_velocity >= 1.8 * early_velocity:
label, arc = "The Climb", "climb"
elif early_velocity >= 1.8 * late_velocity:
label, arc = "The Slow Fade", "fade"
elif len(months) >= 6 and max_month < 3 * median:
label, arc = "The Marathon", "marathon"
Everything is relative to your own median. With absolute floors where a relative rule would be silly — a repo that medians 0.4 commits can't fake a Sprint. The >= 15 floor rejects it before the story begins.
Seven Shapes. One Truth.
The pipeline picks one of seven labels for every project:
| Shape | What Triggers It |
|---|---|
| Fresh Start | Less than 5 commits, or under 60 days old |
| The Return | Longest gap ≥ 60 days, then a real comeback |
| The Sprint | Peak month ≥ 4× median and ≥ 15 commits |
| The Climb | Second half ≥ 1.8× faster than the first |
| The Slow Fade | First half ≥ 1.8× faster than the second |
| The Marathon | Six months or more, no spike above 3× median |
| The Journey | Everything else — quiet parts, loud parts, and all of it |
Each shape ships with a one-line summary. The Sprint's reads: "Built in bursts of deep focus — the history remembers every surge."
People always assume those lines are hand-written per project. They're not. Seven strings, selected by rules. That's the tradeoff — sentences you can't configure are sentences you can't accidentally fake.
The Silence
The rule that generates the most reaction is also the simplest:
if gap and gap["days"] >= 21:
"subtitle": f"{human_gap(gap['days'])} of nothing"
Twenty-one days. Not because three weeks is sacred — because anything shorter is a weekend plus a cold. That number took five repos and one embarrassing live demo of my own project to land on.
The honesty constraint is the design law the whole engine sits under: never claim more than the data shows.
The subtitle says "31 days of nothing" and stops there. It doesn't tell you the maintainer was burnt out. It doesn't suggest the project was dying. It doesn't know any of that. Neither do I. The gap happened. That's the chapter. The reader brings the meaning.
Same philosophy in the fingerprint section — commit-time patterns identified as Night Builder, Weekend Hacker, Marathoner, Lone Wolf — closed with a line I keep coming back to: "A fingerprint of Git behavior — not a personality test."
What Was Harder Than Expected
Which clock is real. Streaks were initially computed from committer date. On any repo with rebases or late pushes, a streak can be authored Tuesday and committed Sunday. One wrong demo later: switched to author date. If you build anything on git data — %at is when the work happened, %ct is when the object landed. They are not the same. Most tools quietly pick one and never say which.
Tests that lied. Fixtures committing multiple commits in the same second made streak tests pass for the wrong reason. I had to stagger fixture dates before the tests could fail correctly. A test suite that can't fail is just documentation with extra steps.
Restraint. It would've been so easy — sentiment analysis of commit messages, a productivity score, a burnout warning. Every single one would have been wrong on at least one repo in the demo set. Wrongness is the one thing a tool built on your history cannot afford. Half the design work was deleting features that would have been genuinely cool and genuinely dishonest.
What I'm Not Claiming
No analytics. There is no way for me to know how many people use this — and that's by design. It would be rich to build a tool about not being surveilled and then install beacons on it.
The README mentions pip install commit-canvas. That's aspirational. The package isn't on PyPI yet. ./run.sh or the website are the real paths today.
It's an MIT-licensed utility. Four public demos. Twenty-seven tests. A website with no trackers. That's the whole pitch.
Try It
git clone https://github.com/ahmadrrrtx/commit-canvas
cd commit-canvas
./run.sh ~/code/your-project --theme sunset --open
Add --max-commits 5000 if you're pointing it at something Linux-sized. Private repos work — the analysis never touches a server. Hit the ◐ button inside the story to flip between eight visual editions without re-running anything.
No install? The website takes a public GitHub URL and does the same thing in your tab.
The Ending I Didn't Plan
When I ran Commit Canvas on its own repository, the engine read the timeline its demo file records — a beginning, a launch, then eight weeks of complete silence — and returned the only label that fit:
The Return. "It went quiet. Then it came back. Projects don't do that by accident."
I didn't write that sentence for my own project. The rule found the pattern. The summary was already sitting there waiting.
That's the moment I'd been building toward since June. Not a dashboard. Not a leaderboard. One file, honest arithmetic, and a story your repo already contained.
If this made you curious about what your own repo's chapters look like — the GitHub repo is here. A star helps more people find it, and it takes two seconds. Either way, run it on something you've spent real time on. The silence chapter has a way of hitting different when it's yours.
website try it here
Top comments (9)
Really interesting approach. The restraint around "The Silence" is probably the part that caught my attention most.
You know there were no commits for 31 days. You don't know why. So the system reports the observable absence and stops rather than turning that absence into an assertion about burnout, abandonment, vacation, or whether development itself actually stopped.
I've been thinking about almost exactly that problem from a different direction: when is "nothing" actually something a system can make a claim about?
A Git history with no commits for 31 days proves there are no commits in the history for that interval. It doesn't prove no code was written. Work could have happened on an unmerged branch, in another repository, locally, or somewhere that was later rewritten out of the history.
That makes your author-date versus committer-date finding especially interesting. Even "when did this happen?" turns out to depend on which event and which provenance you're actually describing.
I wonder if there is another useful distinction hiding in Commit Canvas between observed silence and explained silence. The first can come directly from the Git history. The second would require evidence the repository itself may not contain.
In other words, "31 days with no commits" is defensible. "31 days where nothing happened" might already be saying more than the data knows.
That last little boundary is what I find really interesting.
This is the best kind of feedback: it caught us not living up to our own standard. The chapter body was careful ("No commits for 31 days — Nov 28 to Jan 1"), but the subtitle literally read "31 days of nothing." That's precisely the overstatement you describe, and it had shipped. Your comment is now a commit: the subtitle reads "31 days with no commits," and the chapter closes with an explicit boundary line — "The history records the absence — not the reason."
Your observed/explained distinction is now effectively a design rule for us. Explained silence would require evidence the repository doesn't carry by construction — work on an unmerged branch, in another repo, locally, or rewritten out of the history. Commit Canvas reads only what the history holds, so the honest move is to name that limit inside the artifact rather than paper over it. That's what the new line does.
On author date vs committer date: agreed, and it's why the analyzer uses committer timestamps consistently (falling back to author date only when committer date is absent). The tool describes the history that exists after every rebase, squash, and force-push — not the history the author originally intended. Where the two disagree, that divergence is itself data, but declaring either one "when it really happened" would be exactly the overclaim you're pointing at.
This is pretty much the ideal outcome of leaving a technical comment. :)
I really like the change to "The history records the absence, not the reason." It states the observation boundary without weakening what Commit Canvas can legitimately say, and making the observed/explained distinction a design rule is even better.
Your point about committer timestamps also sharpens it nicely. The tool isn't reconstructing some unknowable "true" development history. It's describing the history represented by the repository as it exists now. Rebase, squash, and force-push are therefore part of the provenance of what the analyzer can observe, not inconveniences to reason around.
Thanks for taking the comment seriously enough to turn it into a commit. That's a pretty satisfying feedback loop.
The interesting design question is whether the history algorithm distinguishes narrative similarity from causal continuity. A useful result should let a reader inspect the commits that formed the cluster and reject a plausible but accidental storyline, rather than treating the generated arc as ground truth.
Fair challenge, and worth answering precisely: the arcs are not causal claims. Shape detection is a statistical classification of rhythm — burst months against the median, silent gaps, runs of consecutive active months — and it's presented as a reading, not a truth. The shape card even lists its runners-up, because "The Comeback" and "The Phoenix" can genuinely tie on the same evidence.
On inspectability, here's where the tool stands today: every number in the story derives from commits. Chapters quote actual commit messages (The Silence chapter names the commit that ended it), the time machine scrubs through the real commits month by month, and the roast section pins every jab to an evidence string. The CLI runs against your local git history, so everything is re-derivable by construction.
But your sharper point stands: the chapter↔commit link isn't surfaced at the chapter level yet. You can inspect the era via the time machine, but you can't yet see "these are the commits that formed this cluster" inside the chapter itself — which is exactly where a reader should be able to reject a plausible-but-accidental storyline. Per-chapter evidence anchors (the actual SHAs that define each chapter) are the right next step for the tool, and this framing is the cleanest argument for it we've seen.
The Flask finding is the most compelling part of the post: a 31-day gap that a changelog renders as nothing, and a maintainer reading it back as a beat. That reframe is real, not decoration.
I'd be curious how the algorithm handles the messy side of history, because that's where narrative tools usually fall apart. Authors collapse under several identities unless you apply a mailmap, and a force-push or a long-lived branch rewrite can move the "chapters" you detected after the fact. Does Commit Canvas run off
git logtopology as it exists today, so a rebase can silently rewrite the story?Also: how are chapter boundaries actually decided — commit-density clustering, tag boundaries, or a gap threshold? If it's density-based, the Silence is a legitimate finding; if it's mostly time-based, you're naming absence and calling it structure, which is a different claim. Worth stating in the README, because people will show the output to managers and it will be read as analysis either way.
Would love to run it against a repo with heavy squash-merge on main.
Really fair pushback — worth addressing properly.
On identity collapse: Commit Canvas doesn't apply mailmap automatically right now. It reads authors exactly as git log surfaces them, which means a contributor with three email addresses counts as three people. The fingerprint and contributor stats will be off on any repo that didn't enforce consistent identity. It's a known gap and honestly one of the first things on the roadmap.
On rebase / force-push: Yes — it runs off the topology as it exists today. A rebase rewrites the story silently, because the story was always built from what git log returns right now. That's not a bug I can fix without lying — the tool can only see what the object store shows it. Worth stating clearly in the README and I'll add it.
On chapter boundaries: Mostly gap and density threshold based, not clustering. The Silence fires on a raw day count between active months — so you're right to call it what it is: naming absence. The Sprint and Marathon fire on velocity ratios relative to the repo's own median. It's not ML-derived segmentation, it's deterministic rules with visible thresholds. The claim is "this gap existed" not "this gap meant something." The meaning is yours to bring.
On squash-merge repos: That's genuinely the hardest case. Heavy squash-merge on main flattens branch history into single commits and the velocity signal goes weird — sprints look like single data points instead of sustained bursts. Would love to know what shape it assigns. Run it and drop the arc here.
The README needs a cleaner honesty section. You just wrote half of it.
The gap is the chapter that separates real projects from demos. Every repo has the commit burst, then silence - and the silence is where the project either became load-bearing in someone's life or quietly joined the graveyard. An algorithm that reads commit history as story has to treat the quiet stretches as data, not absence: a three-month gap followed by a tiny fix commit is the narrative of "came back to it after life happened", which is the most human sentence a git log contains. The hard part is that git records effort but not intent - "refactor" can mean courage or procrastination, and the diff knows which but won't tell you. Curious how your algorithm handles merge commits; those are the chapter breaks written by the calendar instead of the author.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.