DEV Community

ke jia
ke jia

Posted on

git shortlog vs gitpulse: Headcount vs Heatmap, Same Repo

git shortlog is the command everyone knows, and it answers a real question: who committed what, in the period you care about. It is a headcount tool, and as a headcount tool it has been accurate for twenty years. The question it does not answer is the one that actually predicts trouble: where is the change concentrated, and is the concentration in a place that can take it? A file that one person touched fifty times and a file that fifty people touched once have the same total commit count in the shortlog, and they are completely different stories about the project's health.
The comparison on the same repository makes the difference concrete. The shortlog gives you the per-person totals, which is the right input for the credit conversation and the wrong input for the refactor conversation. The hotspot report gives you the per-file totals with the period attached, which is the right input for the refactor conversation and the wrong input for the credit one. Neither replaces the other; they are the two axes of the same grid. The walkthrough below runs both on the same repository over the same month and reads the two reports next to each other, because the interesting findings live in the disagreement between them: the person with the modest commit count who owns the hottest file, and the file with the high count that everyone is touching and nobody understands. That disagreement is where the project's actual risk is hiding.

A comparison is a set of questions with two answers each, and the questions below are the ones the choice actually turns on, not the ones the marketing pages argue. The two sides are gitpulse, Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact. installed with npx @wuchunjie/gitpulse, and the other option, which gets the same treatment and the same chance to win a section. The setup is identical where the setup can be identical, and the identical is the part that makes the difference the difference instead of the configuration. The verdict comes at the end of each section, because the section-level verdict is the one the reader can check against their own context, and the context is the part the global verdict does not know.

Branch Strategy as a Number

The branch strategy is a process fact, and the process fact becomes a number when the report counts the pull-request commits against the direct-push commits over the period. The ratio is the strategy: a team that merges through pull requests shows the ratio in one direction, a team that pushes to main shows it in the other, and the mixed team shows the split, and the split is where the policy lives, because the policy is usually by area instead of by rule, and the area is what the ratio reveals when the report breaks the split down by directory. The number is not the verdict. The verdict is the conversation the number starts, and the conversation is the one that used to take the meeting and now takes the slide, because the slide is the last three periods of the ratio, and the three periods show the trend, and the trend is what the policy argument is about. The number ends the fight by making the fight about the trend instead of the memory.

What Measuring Lots of Repos Taught Me

After running the same set of measurements across a large number of repositories — my own projects, client work, and public repositories from a range of stacks and team sizes — a few patterns held up consistently. Hotspots concentrate: a small fraction of files accounts for the majority of churn in almost every codebase. Commit patterns correlate with team size more than with team health: small teams look erratic, large teams look steady, and the middle is where the interesting data is. And the branch ratio is more stable over time than anyone expects — teams do not really change their shipping process, they change the people. The tool was designed to answer one question per repository. Measuring hundreds of repositories let the data answer a question I had not thought to ask: the shape of a codebase is more determined by its size and its people than by its technology, and the technology shows up mostly in the noise. The design decisions that followed — the default period, the impact weighting, the ratio as a headline metric — are all consequences of those patterns.

The Compare Range: Two Branches, One Answer

The compare flag takes a range and answers the question the branch list cannot: what is the relationship between the two branches, and what does the relationship cost? The main-to-develop range is the classic one, and the report is the difference: the commits on develop that main does not have, the distribution of the commits by person, the files that are hot on both sides. The hot-on-both is the collision warning, and the collision warning is the part that saves the bad merge, because the bad merge is the merge where both branches changed the same lines, and the same lines are what the report flags before the merge attempt. The range is also the feature review input, when the question is what did the feature branch change and who will need to know, and the answer is the file list with the contributor column, and the file list is the review scope, and the review scope is what the range computes instead of the human who has to walk the diff. The one command is the answer the meeting used to build by hand.

The Contributor Impact View

The contributor view is the people axis of the report, and the axis is the part the team reads with the most attention, because the axis is the one that names the person. The view measures the contribution per person over the period: the commit count, the file reach, the area concentration. The file reach is the interesting column, because the file reach is the breadth of the knowledge, and the breadth is what the onboarding needs, when the question is who knows the auth module, the answer is the person whose reach covers it. The area concentration is the second column, and the concentration is the depth, and the depth is what the refactor assignment needs, when the decision is who rewrites the hotspot, the answer is the person whose concentration is in the hotspot. The two columns together are the knowledge map, and the knowledge map is the onboarding answer and the refactor answer and the bus-factor answer, and the three answers are the three meetings the map replaces.

What git log Can't Tell You

The log answers one question: who changed what, and when. It is a great question, and it is not the question. The question that actually predicts trouble is: where is the codebase accumulating damage? A file that changes forty-seven times a month is a hotspot no matter how clean each individual commit looks. A team whose commits cluster on late weekend hours is a team that is running hot no matter how green the pipeline is. The log shows you the events; the analytics show you the patterns underneath the events. That difference is the difference between a log and an analysis, and it is the reason a tool that reads your entire history and summarizes the shape of it is a different category of instrument from the version control system itself. The events are facts; the shape is the information. This article is about the shape, and about how to get it in thirty seconds instead of an afternoon of log-reading.

Branch Strategy: The PR-versus-Direct-Push Ratio

How does your team actually ship? The honest answer is in the history, not in the process document. The analytics measure the ratio of changes that arrive via merge — through review — versus changes that land on the branch directly. The number is not moral; direct push is fine for docs, dependencies, and solo work. But the shape of it tells you how much review your code actually gets. A repository where the large majority of commits are direct pushes has a code review process that exists in the wiki, not in the history. The compare flag lets you look at the ratio between two branches, which answers the practical question: is the integration branch cleaner than the trunk, or did the branching experiment produce more direct landings than the mainline? The ratio is a process fact, and process facts are the kind of thing you cannot get from a meeting. The meeting tells you what the process is supposed to be. The history tells you what it is.

When Analytics Are the Wrong Tool

Honesty section: repository analytics are not for every repository. A new project with two weeks of history has no patterns to analyze — the hotspots are noise, the commit shape is just one person working, and the branch ratio is undefined. A solo project where you are the only contributor and you already know the codebase has little to tell you. And a repository where the team will use the output as a performance signal will get a distorted version of the truth, because the data was never collected for that purpose. The right use is a diagnostic: a repository you are inheriting, a repository that feels slower than it should, a repository you are about to present to stakeholders. Use it as a stethoscope, not as a scoreboard. The stethoscope tells you where to listen; the scoreboard tells you who to blame, and the blame is never in the data. Knowing which instrument you are holding is the whole skill, and the wrong instrument, used with confidence, is worse than no instrument at all.

Reading a Report You Did Not Author

The report is a tool for the person who is not the committer, and the non-author is the larger audience: the tech lead reading the team's repository, the new developer reading the project they just joined, the reviewer reading the area before the review. The reading skill for the non-author is the same three passes the author uses, and the passes are the structure. First pass is the shape: the total volume, the active files, the contributor count, and the shape is the one-paragraph summary of the project's state. Second pass is the concentration: the hotspots, the areas where the energy is, and the concentration is the map of where the project is spending itself. Third pass is the change: the period over period delta, and the delta is the direction, and the direction is the part the summary needs, because the summary without the direction is the photograph, and the photograph is what last month's summary was. The three passes take the length of the report, and the report is the part that does not need the author present.

Onboarding a New Repo in 60 Seconds

The worst moment in a developer's life is opening a repository they have never seen: no documentation, a multi-year history, and a codebase that looks the same in every folder. The analytics turn that moment into a sixty-second orientation. The hotspots tell you where the action is — start reading there, not at the README. The contributor impact tells you who to ask when the code does not make sense. The branch strategy tells you how changes actually get in, which is the one process fact that documentation never gets right. Inherited projects are the common case, not the exception: new job, new team, acquired codebase, open-source contribution. A tool that compresses the first hour into the first minute pays for itself on the first use. The sixty seconds buy you something rarer than time: a map, so the first day is spent building context instead of stumbling through directories hoping the important file announces itself.

Commit Patterns and the Burnout Signal

Commits have a rhythm, and the rhythm is a health metric. The analytics look at when commits happen — by hour, by day of week, by streak — and the pattern is more honest than any survey. A team that commits steadily on weekdays is one kind of team. A team whose commits spike on Friday nights and Saturday mornings is another, and the difference is visible in the data without asking a single person how they are doing. I am not saying commit timing equals wellbeing; it is a signal, not a verdict. But it is a signal that a manager who only looks at velocity will never see, and it is exactly the kind of information that is cheap to collect and expensive to guess at. The right use is the trend, not the snapshot: one busy weekend is a fact, four busy weekends in a row is a pattern, and the pattern is the conversation worth having, had with data instead of with hunches, before the hunches become resignations.

Pairing gitpulse With dotguard in One Pipeline

A repository has two kinds of health: structural and security, and they are measured by different tools. One tells you whether the code is organized and maintained — hotspots under control, contributors distributed, branches reviewed. The other tells you whether the repository is leaking — secrets in configs, tokens in history, credentials in compose files. Running both in the same weekly pass is a complete repository health check in under a minute, and the outputs are complementary: a hotspot in a config file that the secret scanner also flags is a refactoring task with a security deadline. Two small CLIs, no shared infrastructure, one habit. That is the whole architecture of the pipeline, and it is the kind of architecture that survives because nothing in it needs to be maintained. The weekly pass becomes the meeting the team does not have to schedule: the file is the agenda, the findings are the action items, and the rotation of the week is the follow-up. Infrastructure this small is not a platform. It is a reflex, and reflexes are what teams actually keep.

The takeaway

The comparison ends where the context begins, and the context is the part only the reader has. The sections gave the data: the same input, the two answers, the section-level verdicts. gitpulse is Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact., available via npx @wuchunjie/gitpulse, source at https://github.com/wuchunjie00/gitpulse. The reader who is all in one ecosystem gets the ecosystem's tool, and the reader who spans ecosystems gets the tool that spans with them, and the two are the parts the data above supports. The choice is the reader's, and the reader's is the part the comparison respects, because the respects is what the ad does not.

Top comments (0)