DEV Community

ke jia
ke jia

Posted on

CSV Export, Spreadsheet, Standup: Turning Git Analytics Into a Conversation

Analytics that die in the terminal are measurements. Analytics that reach the standup are conversations, and the gap between the two is a CSV file. The export flag writes the report to a spreadsheet format, and that one flag changes what the data is for. In the terminal the numbers answer the question that prompted the run. In the spreadsheet they sit next to each other across months, and the questions change: not what happened this week, but what has been happening for a quarter, and which of the patterns in that series is worth a conversation instead of a fix.
The workflow that makes it stick is deliberately small. The export goes into a shared sheet with one row per month, the columns are the metrics that survive comparison across time, and the standup gets one slide that is the last three rows of that sheet. The question the slide answers is not a performance question; it is a routing question: where did the project's energy go this month, and does that match where we said it would go? When the answer is no, the standup becomes a thirty-second discussion of the delta instead of a status report nobody reads. The data is not a scorecard and should never be one; it is the shared map that lets a team argue about the territory instead of about whose memory is correct. The export is the bridge, and the conversation is what the bridge is for.

The walkthrough assumes the tool is already one command away. gitpulse is Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact. Run npx @wuchunjie/gitpulse and the help output is the table of contents for everything below. The sections follow the order a first run actually goes, not the order the README presents it in, because the first run is where the questions happen and the README is written for the second run. Where a step can fail, the failure is named, because the failure is the part the walkthrough exists to cover, and the covered failure is the difference between the tutorial and the experiment. By the end you have run the full path once, and the once is what makes the next run the fast one.

Branch Strategy as a Number

The branch strategy is a process fact, and the process fact becomes a number when the report counts the pull-request commits against the direct-push commits over the period. The ratio is the strategy: a team that merges through pull requests shows the ratio in one direction, a team that pushes to main shows it in the other, and the mixed team shows the split, and the split is where the policy lives, because the policy is usually by area instead of by rule, and the area is what the ratio reveals when the report breaks the split down by directory. The number is not the verdict. The verdict is the conversation the number starts, and the conversation is the one that used to take the meeting and now takes the slide, because the slide is the last three periods of the ratio, and the three periods show the trend, and the trend is what the policy argument is about. The number ends the fight by making the fight about the trend instead of the memory.

What Measuring Lots of Repos Taught Me

After running the same set of measurements across a large number of repositories — my own projects, client work, and public repositories from a range of stacks and team sizes — a few patterns held up consistently. Hotspots concentrate: a small fraction of files accounts for the majority of churn in almost every codebase. Commit patterns correlate with team size more than with team health: small teams look erratic, large teams look steady, and the middle is where the interesting data is. And the branch ratio is more stable over time than anyone expects — teams do not really change their shipping process, they change the people. The tool was designed to answer one question per repository. Measuring hundreds of repositories let the data answer a question I had not thought to ask: the shape of a codebase is more determined by its size and its people than by its technology, and the technology shows up mostly in the noise. The design decisions that followed — the default period, the impact weighting, the ratio as a headline metric — are all consequences of those patterns.

When Analytics Are the Wrong Tool

Honesty section: repository analytics are not for every repository. A new project with two weeks of history has no patterns to analyze — the hotspots are noise, the commit shape is just one person working, and the branch ratio is undefined. A solo project where you are the only contributor and you already know the codebase has little to tell you. And a repository where the team will use the output as a performance signal will get a distorted version of the truth, because the data was never collected for that purpose. The right use is a diagnostic: a repository you are inheriting, a repository that feels slower than it should, a repository you are about to present to stakeholders. Use it as a stethoscope, not as a scoreboard. The stethoscope tells you where to listen; the scoreboard tells you who to blame, and the blame is never in the data. Knowing which instrument you are holding is the whole skill, and the wrong instrument, used with confidence, is worse than no instrument at all.

Reading a Quarter: The Long View

The monthly view is for the pulse; the quarterly view is for the story. Extend the period to a quarter and you can see the arc of a release: the hotspot that built up through the quarter, the contributor who carried the middle two months, the week where the commit pattern broke. That view is where hiring and attrition show up. A new contributor ramping is visible as a growing impact curve. A person leaving is visible as a curve going flat, often weeks before the announcement. And a release that was supposed to be a sprint shows up as a commit pattern that never recovered. None of this requires a dashboard or a database. It is version control history, which you already have, read by a tool that knows which questions to ask. The quarter is the right window for most organizational questions, because a month is noise and a year is archaeology. A quarter is a story with a beginning and an end, and the report reads it for you.

The Hotspot Table, Column by Column

The hotspot table is the center of the report, and the columns are worth reading in order, because the order is the reading. The first column is the file, and the file is the unit of the measurement, which is the unit the team thinks in, because the team argues about files instead of commits. The second column is the change count in the period, and the count is the energy measurement, and the energy is what the file is spending. The third column is the distinct contributors, and the contributors are the coordination cost, because the file that five people touch is the file that five people wait on. The fourth column is the revert share, and the revert share is the understanding cost, because the reverts are the changes that did not stick, and the not-sticking is the signal that the area is not understood. The four columns together are the diagnosis, and the diagnosis is the part the single-column view, the change count alone, does not give, because the count without the context is the number that argues for itself.

The 30-Second Report

Run the tool with no arguments in a repository and you get the current month, summarized: the hot files, the commit shape, the branch behavior, the top contributors by impact rather than commit count. Thirty seconds from typing the command to having a picture of the repository that the log would take an hour to assemble by hand. The design goal was a report you would actually read, which means short enough to fit on a screen and dense enough that each line earns its place. I use it as a standing habit: first command of the week in any repository I am active in, the same way some developers start with status. The habit is the point. A report that takes ten minutes to generate gets run once a quarter. A report that takes thirty seconds gets run every Monday. The frequency is the feature, because the value of repository analytics is in the delta — what changed since last week — and the delta is only visible if you look often enough to see it move.

The Compare Range: Two Branches, One Answer

The compare flag takes a range and answers the question the branch list cannot: what is the relationship between the two branches, and what does the relationship cost? The main-to-develop range is the classic one, and the report is the difference: the commits on develop that main does not have, the distribution of the commits by person, the files that are hot on both sides. The hot-on-both is the collision warning, and the collision warning is the part that saves the bad merge, because the bad merge is the merge where both branches changed the same lines, and the same lines are what the report flags before the merge attempt. The range is also the feature review input, when the question is what did the feature branch change and who will need to know, and the answer is the file list with the contributor column, and the file list is the review scope, and the review scope is what the range computes instead of the human who has to walk the diff. The one command is the answer the meeting used to build by hand.

Commit Patterns: The Weekday, the Weekend, and the Hour

The commit pattern section reads the when, and the when is the part of the data that the status report never captures, because the status captures the what and the what is what the developer says happened. The weekday distribution is the baseline, and the baseline is the expected shape of the work. The weekend ratio is the cost signal, and the cost signal is the fraction of the commits that landed outside the working week, and the fraction is the number that the standup should see, because the standup is the meeting where the schedule gets fixed. The hour distribution is the third read, and the hour is the part that shows the late-night cluster, and the late-night cluster is the deadline signature, and the deadline signature is the pattern that repeats until the deadline stops being the deadline. The three reads are the same data three ways, and the three ways are the difference between the activity report and the cost report, and the cost report is the one the team can act on.

Commit Patterns and the Burnout Signal

Commits have a rhythm, and the rhythm is a health metric. The analytics look at when commits happen — by hour, by day of week, by streak — and the pattern is more honest than any survey. A team that commits steadily on weekdays is one kind of team. A team whose commits spike on Friday nights and Saturday mornings is another, and the difference is visible in the data without asking a single person how they are doing. I am not saying commit timing equals wellbeing; it is a signal, not a verdict. But it is a signal that a manager who only looks at velocity will never see, and it is exactly the kind of information that is cheap to collect and expensive to guess at. The right use is the trend, not the snapshot: one busy weekend is a fact, four busy weekends in a row is a pattern, and the pattern is the conversation worth having, had with data instead of with hunches, before the hunches become resignations.

What git log Can't Tell You

The log answers one question: who changed what, and when. It is a great question, and it is not the question. The question that actually predicts trouble is: where is the codebase accumulating damage? A file that changes forty-seven times a month is a hotspot no matter how clean each individual commit looks. A team whose commits cluster on late weekend hours is a team that is running hot no matter how green the pipeline is. The log shows you the events; the analytics show you the patterns underneath the events. That difference is the difference between a log and an analysis, and it is the reason a tool that reads your entire history and summarizes the shape of it is a different category of instrument from the version control system itself. The events are facts; the shape is the information. This article is about the shape, and about how to get it in thirty seconds instead of an afternoon of log-reading.

Branch Strategy: The PR-versus-Direct-Push Ratio

How does your team actually ship? The honest answer is in the history, not in the process document. The analytics measure the ratio of changes that arrive via merge — through review — versus changes that land on the branch directly. The number is not moral; direct push is fine for docs, dependencies, and solo work. But the shape of it tells you how much review your code actually gets. A repository where the large majority of commits are direct pushes has a code review process that exists in the wiki, not in the history. The compare flag lets you look at the ratio between two branches, which answers the practical question: is the integration branch cleaner than the trunk, or did the branching experiment produce more direct landings than the mainline? The ratio is a process fact, and process facts are the kind of thing you cannot get from a meeting. The meeting tells you what the process is supposed to be. The history tells you what it is.

The takeaway

The walkthrough is done, and the done is the state where the next run is the fast one. gitpulse is the tool that made the path the short one: Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact. The install line, for the next reader who is starting now, is npx @wuchunjie/gitpulse, and the repository is https://github.com/wuchunjie00/gitpulse. The steps above are the ones that work, and the work is the part that the version keeps working, because the keeps is the part the maintenance is. If a step failed in your run, the failure is the thing to report, because the report is what the next run reads, and the reads is what the fix becomes.

Top comments (0)