DEV Community

ke jia
ke jia

Posted on

The Git Metrics I Trust and the Ones That Lie (With Numbers)

Every git metric has a failure mode, and the failure mode is the part that decides whether the metric is worth reading or worth ignoring, and the list below is the split: the metrics I trust, the ones that lie, and the specific numbers that taught me which is which. Commit count is the classic liar, because it measures activity, not work, and a team that writes ten small commits a day will always outrank a team that writes three large ones, and the ranking tells you nothing about the code. The metrics that hold up are the ones that measure structure over time instead of volume, and they are the ones this list is built around.
The trusted list starts with the hotspot persistence, the file that stays hot across consecutive periods, which is the difference between a busy file and a failing one, and the number behind it is the three-month threshold I use, because one month is noise and three months is a pattern. The weekend commit ratio is second, and it is the one metric that has changed how I run a standup, because it is the only number that measures cost instead of output, and the cost is the thing the status report never captures. The liars get the same treatment in the section below: the metric, why it misleads, the number that showed me, and the replacement metric that measures the same thing honestly.

Lists are the format the reader can scan, and the scan is the feature, because the feature is the entry that the reader stops on, and the stops-on is the one that was the answer. The entries below are the ten that earned the slot, and the earned is the part that the filler does not have, because the filler is the entry that pads the count, and the pad is what the reader sees, and the sees is what the trust loses. The context is gitpulse, Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact., and the context is the part that makes the entries the specific ones instead of the generic ones, because the generic entry is the one the reader has read before, and the read-before is the one that does not stop the scroll.

The Hotspot Table, Column by Column

The hotspot table is the center of the report, and the columns are worth reading in order, because the order is the reading. The first column is the file, and the file is the unit of the measurement, which is the unit the team thinks in, because the team argues about files instead of commits. The second column is the change count in the period, and the count is the energy measurement, and the energy is what the file is spending. The third column is the distinct contributors, and the contributors are the coordination cost, because the file that five people touch is the file that five people wait on. The fourth column is the revert share, and the revert share is the understanding cost, because the reverts are the changes that did not stick, and the not-sticking is the signal that the area is not understood. The four columns together are the diagnosis, and the diagnosis is the part the single-column view, the change count alone, does not give, because the count without the context is the number that argues for itself.

The 30-Second Report

Run the tool with no arguments in a repository and you get the current month, summarized: the hot files, the commit shape, the branch behavior, the top contributors by impact rather than commit count. Thirty seconds from typing the command to having a picture of the repository that the log would take an hour to assemble by hand. The design goal was a report you would actually read, which means short enough to fit on a screen and dense enough that each line earns its place. I use it as a standing habit: first command of the week in any repository I am active in, the same way some developers start with status. The habit is the point. A report that takes ten minutes to generate gets run once a quarter. A report that takes thirty seconds gets run every Monday. The frequency is the feature, because the value of repository analytics is in the delta — what changed since last week — and the delta is only visible if you look often enough to see it move.

What Measuring Lots of Repos Taught Me

After running the same set of measurements across a large number of repositories — my own projects, client work, and public repositories from a range of stacks and team sizes — a few patterns held up consistently. Hotspots concentrate: a small fraction of files accounts for the majority of churn in almost every codebase. Commit patterns correlate with team size more than with team health: small teams look erratic, large teams look steady, and the middle is where the interesting data is. And the branch ratio is more stable over time than anyone expects — teams do not really change their shipping process, they change the people. The tool was designed to answer one question per repository. Measuring hundreds of repositories let the data answer a question I had not thought to ask: the shape of a codebase is more determined by its size and its people than by its technology, and the technology shows up mostly in the noise. The design decisions that followed — the default period, the impact weighting, the ratio as a headline metric — are all consequences of those patterns.

Reading a Report You Did Not Author

The report is a tool for the person who is not the committer, and the non-author is the larger audience: the tech lead reading the team's repository, the new developer reading the project they just joined, the reviewer reading the area before the review. The reading skill for the non-author is the same three passes the author uses, and the passes are the structure. First pass is the shape: the total volume, the active files, the contributor count, and the shape is the one-paragraph summary of the project's state. Second pass is the concentration: the hotspots, the areas where the energy is, and the concentration is the map of where the project is spending itself. Third pass is the change: the period over period delta, and the delta is the direction, and the direction is the part the summary needs, because the summary without the direction is the photograph, and the photograph is what last month's summary was. The three passes take the length of the report, and the report is the part that does not need the author present.

The Contributor Impact View

The contributor view is the people axis of the report, and the axis is the part the team reads with the most attention, because the axis is the one that names the person. The view measures the contribution per person over the period: the commit count, the file reach, the area concentration. The file reach is the interesting column, because the file reach is the breadth of the knowledge, and the breadth is what the onboarding needs, when the question is who knows the auth module, the answer is the person whose reach covers it. The area concentration is the second column, and the concentration is the depth, and the depth is what the refactor assignment needs, when the decision is who rewrites the hotspot, the answer is the person whose concentration is in the hotspot. The two columns together are the knowledge map, and the knowledge map is the onboarding answer and the refactor answer and the bus-factor answer, and the three answers are the three meetings the map replaces.

Pairing gitpulse With dotguard in One Pipeline

A repository has two kinds of health: structural and security, and they are measured by different tools. One tells you whether the code is organized and maintained — hotspots under control, contributors distributed, branches reviewed. The other tells you whether the repository is leaking — secrets in configs, tokens in history, credentials in compose files. Running both in the same weekly pass is a complete repository health check in under a minute, and the outputs are complementary: a hotspot in a config file that the secret scanner also flags is a refactoring task with a security deadline. Two small CLIs, no shared infrastructure, one habit. That is the whole architecture of the pipeline, and it is the kind of architecture that survives because nothing in it needs to be maintained. The weekly pass becomes the meeting the team does not have to schedule: the file is the agenda, the findings are the action items, and the rotation of the week is the follow-up. Infrastructure this small is not a platform. It is a reflex, and reflexes are what teams actually keep.

Branch Strategy: The PR-versus-Direct-Push Ratio

How does your team actually ship? The honest answer is in the history, not in the process document. The analytics measure the ratio of changes that arrive via merge — through review — versus changes that land on the branch directly. The number is not moral; direct push is fine for docs, dependencies, and solo work. But the shape of it tells you how much review your code actually gets. A repository where the large majority of commits are direct pushes has a code review process that exists in the wiki, not in the history. The compare flag lets you look at the ratio between two branches, which answers the practical question: is the integration branch cleaner than the trunk, or did the branching experiment produce more direct landings than the mainline? The ratio is a process fact, and process facts are the kind of thing you cannot get from a meeting. The meeting tells you what the process is supposed to be. The history tells you what it is.

When Analytics Are the Wrong Tool

Honesty section: repository analytics are not for every repository. A new project with two weeks of history has no patterns to analyze — the hotspots are noise, the commit shape is just one person working, and the branch ratio is undefined. A solo project where you are the only contributor and you already know the codebase has little to tell you. And a repository where the team will use the output as a performance signal will get a distorted version of the truth, because the data was never collected for that purpose. The right use is a diagnostic: a repository you are inheriting, a repository that feels slower than it should, a repository you are about to present to stakeholders. Use it as a stethoscope, not as a scoreboard. The stethoscope tells you where to listen; the scoreboard tells you who to blame, and the blame is never in the data. Knowing which instrument you are holding is the whole skill, and the wrong instrument, used with confidence, is worse than no instrument at all.

Commit Patterns: The Weekday, the Weekend, and the Hour

The commit pattern section reads the when, and the when is the part of the data that the status report never captures, because the status captures the what and the what is what the developer says happened. The weekday distribution is the baseline, and the baseline is the expected shape of the work. The weekend ratio is the cost signal, and the cost signal is the fraction of the commits that landed outside the working week, and the fraction is the number that the standup should see, because the standup is the meeting where the schedule gets fixed. The hour distribution is the third read, and the hour is the part that shows the late-night cluster, and the late-night cluster is the deadline signature, and the deadline signature is the pattern that repeats until the deadline stops being the deadline. The three reads are the same data three ways, and the three ways are the difference between the activity report and the cost report, and the cost report is the one the team can act on.

Reading a Quarter: The Long View

The monthly view is for the pulse; the quarterly view is for the story. Extend the period to a quarter and you can see the arc of a release: the hotspot that built up through the quarter, the contributor who carried the middle two months, the week where the commit pattern broke. That view is where hiring and attrition show up. A new contributor ramping is visible as a growing impact curve. A person leaving is visible as a curve going flat, often weeks before the announcement. And a release that was supposed to be a sprint shows up as a commit pattern that never recovered. None of this requires a dashboard or a database. It is version control history, which you already have, read by a tool that knows which questions to ask. The quarter is the right window for most organizational questions, because a month is noise and a year is archaeology. A quarter is a story with a beginning and an end, and the report reads it for you.

CSV Out: Health Reports Without a Dashboard

The tool can export its findings as CSV, and that one flag is the answer to the dashboard question. The standard objection to CLI analytics is that nobody will read the output in a terminal. Fair. So the output becomes a file: a health report, attached to a message, emailed to the team on the first of the month, dropped into a spreadsheet that the tech lead already maintains. No server, no subscription, no infrastructure project. The data leaves your machine only when you decide it should, and the format is one that every tool you already have can open. For most teams, the right analytics infrastructure is a file and a habit, not a platform. The file is the report; the habit is the monthly export and the five-minute read. Everything else — the dashboards, the integrations, the subscriptions — is what you add when the file-and-habit version stops being enough, which for most small teams is never, and for the teams where it is, the CSV is the import format they will thank you for.

The takeaway

The ten are the ten, and the ten is the part the count promised. The entries are the working set, and the working-set is the part the reader keeps, because the keeps is the bookmark, and the bookmark is the part the list buys. gitpulse is Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact., one command away with npx @wuchunjie/gitpulse, and the one-command is the part that makes the ten entries the ten commands instead of the ten readings. The repository is https://github.com/wuchunjie00/gitpulse, and the repository is the part the missing eleventh goes to, because the eleventh is the part the reader finds, and the finds is what the list that keeps working does.

Top comments (0)