DEV Community

ke jia
ke jia

Posted on

gitpulse for a New Repo: The First Report You Should Read (and the Part to Ignore)

Run git analytics on a repository with a month of history and the report will look thin. No hotspots worth worrying about, a scatter of commits, a contributor section with your name in it once. The instinct is to conclude the tool is useless on young repos. That is backwards. The first report is a baseline: this is what quiet looks like in this project, measured with the yardstick you will still be using in six months. When the history is deep, you are comparing against something instead of against your feelings about the codebase.
Read the structure in that first run, not the numbers. The file list shows how the codebase is organized and which files already carry the most changes, even at low counts. The branch section tells you whether work goes through pull requests or straight to the main branch, which is a process fact, not a performance fact. Ignore anything that looks like a verdict: a file with three commits in its first month is not a technical debt magnet, it is just a file. Re-run the same command monthly and the tool becomes a time series instead of a snapshot. That is when the readings stop being trivia. The first report is the zero point on the ruler, and the rest of this piece is how to read a ruler you are about to use for a long time.

This is the hands-on section, and it is written as the run goes, which means the commands are in the order they are typed and the output is the output that came back, including the parts that look like errors and are not. gitpulse is Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact. The install line is npx @wuchunjie/gitpulse, and the repository is https://github.com/wuchunjie00/gitpulse if you want to read the source before you run it, because the reading is the option, not the requirement, and the requirement is the run. Each step below is short enough that a copy-paste session can follow it without losing the thread, and the thread is the part that the long tutorial loses, and the loses is what sends the reader to a different tab.

The Hotspot Table, Column by Column

The hotspot table is the center of the report, and the columns are worth reading in order, because the order is the reading. The first column is the file, and the file is the unit of the measurement, which is the unit the team thinks in, because the team argues about files instead of commits. The second column is the change count in the period, and the count is the energy measurement, and the energy is what the file is spending. The third column is the distinct contributors, and the contributors are the coordination cost, because the file that five people touch is the file that five people wait on. The fourth column is the revert share, and the revert share is the understanding cost, because the reverts are the changes that did not stick, and the not-sticking is the signal that the area is not understood. The four columns together are the diagnosis, and the diagnosis is the part the single-column view, the change count alone, does not give, because the count without the context is the number that argues for itself.

Reading a Quarter: The Long View

The monthly view is for the pulse; the quarterly view is for the story. Extend the period to a quarter and you can see the arc of a release: the hotspot that built up through the quarter, the contributor who carried the middle two months, the week where the commit pattern broke. That view is where hiring and attrition show up. A new contributor ramping is visible as a growing impact curve. A person leaving is visible as a curve going flat, often weeks before the announcement. And a release that was supposed to be a sprint shows up as a commit pattern that never recovered. None of this requires a dashboard or a database. It is version control history, which you already have, read by a tool that knows which questions to ask. The quarter is the right window for most organizational questions, because a month is noise and a year is archaeology. A quarter is a story with a beginning and an end, and the report reads it for you.

Commit Patterns and the Burnout Signal

Commits have a rhythm, and the rhythm is a health metric. The analytics look at when commits happen — by hour, by day of week, by streak — and the pattern is more honest than any survey. A team that commits steadily on weekdays is one kind of team. A team whose commits spike on Friday nights and Saturday mornings is another, and the difference is visible in the data without asking a single person how they are doing. I am not saying commit timing equals wellbeing; it is a signal, not a verdict. But it is a signal that a manager who only looks at velocity will never see, and it is exactly the kind of information that is cheap to collect and expensive to guess at. The right use is the trend, not the snapshot: one busy weekend is a fact, four busy weekends in a row is a pattern, and the pattern is the conversation worth having, had with data instead of with hunches, before the hunches become resignations.

What Measuring Lots of Repos Taught Me

After running the same set of measurements across a large number of repositories — my own projects, client work, and public repositories from a range of stacks and team sizes — a few patterns held up consistently. Hotspots concentrate: a small fraction of files accounts for the majority of churn in almost every codebase. Commit patterns correlate with team size more than with team health: small teams look erratic, large teams look steady, and the middle is where the interesting data is. And the branch ratio is more stable over time than anyone expects — teams do not really change their shipping process, they change the people. The tool was designed to answer one question per repository. Measuring hundreds of repositories let the data answer a question I had not thought to ask: the shape of a codebase is more determined by its size and its people than by its technology, and the technology shows up mostly in the noise. The design decisions that followed — the default period, the impact weighting, the ratio as a headline metric — are all consequences of those patterns.

Reading a Report You Did Not Author

The report is a tool for the person who is not the committer, and the non-author is the larger audience: the tech lead reading the team's repository, the new developer reading the project they just joined, the reviewer reading the area before the review. The reading skill for the non-author is the same three passes the author uses, and the passes are the structure. First pass is the shape: the total volume, the active files, the contributor count, and the shape is the one-paragraph summary of the project's state. Second pass is the concentration: the hotspots, the areas where the energy is, and the concentration is the map of where the project is spending itself. Third pass is the change: the period over period delta, and the delta is the direction, and the direction is the part the summary needs, because the summary without the direction is the photograph, and the photograph is what last month's summary was. The three passes take the length of the report, and the report is the part that does not need the author present.

Branch Strategy as a Number

The branch strategy is a process fact, and the process fact becomes a number when the report counts the pull-request commits against the direct-push commits over the period. The ratio is the strategy: a team that merges through pull requests shows the ratio in one direction, a team that pushes to main shows it in the other, and the mixed team shows the split, and the split is where the policy lives, because the policy is usually by area instead of by rule, and the area is what the ratio reveals when the report breaks the split down by directory. The number is not the verdict. The verdict is the conversation the number starts, and the conversation is the one that used to take the meeting and now takes the slide, because the slide is the last three periods of the ratio, and the three periods show the trend, and the trend is what the policy argument is about. The number ends the fight by making the fight about the trend instead of the memory.

Onboarding a New Repo in 60 Seconds

The worst moment in a developer's life is opening a repository they have never seen: no documentation, a multi-year history, and a codebase that looks the same in every folder. The analytics turn that moment into a sixty-second orientation. The hotspots tell you where the action is — start reading there, not at the README. The contributor impact tells you who to ask when the code does not make sense. The branch strategy tells you how changes actually get in, which is the one process fact that documentation never gets right. Inherited projects are the common case, not the exception: new job, new team, acquired codebase, open-source contribution. A tool that compresses the first hour into the first minute pays for itself on the first use. The sixty seconds buy you something rarer than time: a map, so the first day is spent building context instead of stumbling through directories hoping the important file announces itself.

Commit Patterns: The Weekday, the Weekend, and the Hour

The commit pattern section reads the when, and the when is the part of the data that the status report never captures, because the status captures the what and the what is what the developer says happened. The weekday distribution is the baseline, and the baseline is the expected shape of the work. The weekend ratio is the cost signal, and the cost signal is the fraction of the commits that landed outside the working week, and the fraction is the number that the standup should see, because the standup is the meeting where the schedule gets fixed. The hour distribution is the third read, and the hour is the part that shows the late-night cluster, and the late-night cluster is the deadline signature, and the deadline signature is the pattern that repeats until the deadline stops being the deadline. The three reads are the same data three ways, and the three ways are the difference between the activity report and the cost report, and the cost report is the one the team can act on.

The Period Flag: Why Since When Is the Whole Interface

The analytics question is always a question about a window, and the period flag is the window. The month view answers the what-is-happening-now question, the quarter view answers the what-has-been-tending question, and the year view answers the what-is-this-project question. The same repository reads completely differently in the three windows, and the different reads are the feature, because the feature is the time axis, and the time axis is what the log command does not give you without the date math and the merge-commit handling that the one-liner gets wrong on the third try. The default period is the choice the tool makes for you, and the default is the quarter, because the quarter is the window where the pattern shows and the noise is still low. The flag is the escape hatch for the specific question, and the specific question is the one the standup raised, and the standup question is the one the report should answer in one command instead of one afternoon.

What git log Can't Tell You

The log answers one question: who changed what, and when. It is a great question, and it is not the question. The question that actually predicts trouble is: where is the codebase accumulating damage? A file that changes forty-seven times a month is a hotspot no matter how clean each individual commit looks. A team whose commits cluster on late weekend hours is a team that is running hot no matter how green the pipeline is. The log shows you the events; the analytics show you the patterns underneath the events. That difference is the difference between a log and an analysis, and it is the reason a tool that reads your entire history and summarizes the shape of it is a different category of instrument from the version control system itself. The events are facts; the shape is the information. This article is about the shape, and about how to get it in thirty seconds instead of an afternoon of log-reading.

The 2>/dev/null Bug: How I Learned to Test Where Users Are

A recent release was a fix for a bug that had been silently present for three months: a shell redirect that is fine on Linux and macOS and quietly wrong on Windows, where stderr handling behaves differently. The tool ran, produced output, and the bug only showed up as missing data for a subset of users on a subset of platforms. The lesson is not that I made a mistake. The lesson is that a CLI tool's test matrix has to include the platform where your users are, not the platform where you are. Three months of silent wrongness is longer than most bugs survive in a well-tested product, and the fix was one line. But finding it required a user report, not a test. That gap — between the platforms you test on and the platforms your users run on — is the most expensive gap in a cross-platform CLI, and it is the one that never shows up in your own usage because your machine is the one platform that always works. The fix was a line. The lesson is a test matrix.

The takeaway

The run is the proof, and the proof is the part the tutorial is. gitpulse gives you Git analytics in your terminal: file hotspots, commit patterns, branch strategy, contributor impact. and the giving is one command: npx @wuchunjie/gitpulse. The repository, https://github.com/wuchunjie00/gitpulse, is where the source lives and the issues go, and the goes is the part that the stuck reader uses, because the stuck is the part the tutorial cannot see from here. The path above is the one that was run and the run was clean, and the clean is the part that the next run inherits, and the inherits is what the tutorial buys for the reader who follows it to the end.

Top comments (0)