DEV Community

Cover image for I Built an Agent Skill to Make Management Happy
László Szabó
László Szabó

Posted on

I Built an Agent Skill to Make Management Happy

It didn't make me write better code.

It didn't make me deliver features faster.

It didn't make the product more reliable.

It made me commit more often.

And according to the KPI dashboard, that meant I had improved.


A while ago, I worked at a company where individual Git commit counts were tracked as an engineering performance metric.

One day, I was told that my number of commits was lower than some of my coworkers'.

I needed to improve it.

There was one small problem.

At the time, I wasn't working purely as an individual contributor.

I was simultaneously doing things like:

  • software architecture
  • technical decision-making
  • mentoring
  • code reviews
  • team leadership
  • debugging difficult problems
  • coordinating work across products
  • and, occasionally, actually writing code

Most of those things produce exactly zero Git commits.

But the dashboard didn't know that.

The dashboard knew one thing:

commits = productivity
Enter fullscreen mode Exit fullscreen mode

So I decided to improve my productivity.

At least according to the dashboard.

I built an agent skill

My normal development style was fairly simple.

I would work on a coherent piece of functionality, finish it, review the changes, and commit it.

Sometimes that meant one commit.

Sometimes a few.

I never thought:

"How many commits should this work generate so my KPI looks good?"

Apparently, I should have.

So I created an agent skill called:

crazy-commiting
Enter fullscreen mode Exit fullscreen mode

Its job was simple.

Take my current changes and split them into the maximum number of reasonable, coherent commits.

Not fake commits.

Not whitespace changes.

Not:

fix
fix2
final fix
final fix 2
really final
Enter fullscreen mode Exit fullscreen mode

Every commit still had to make sense independently.

The agent would inspect my changes, determine which parts could logically stand on their own, stage them separately, and create proper commit messages.

Instead of doing this:

Implement user synchronization
Enter fullscreen mode Exit fullscreen mode

I might now get something closer to:

Add synchronization configuration model
Add repository method for loading remote users
Implement user mapping logic
Add synchronization service
Add synchronization validation
Add synchronization error handling
Add synchronization tests
Enter fullscreen mode Exit fullscreen mode

Same feature.

Same code.

Same amount of engineering work.

Much better KPI.

My workflow effectively became:

finish work
↓
review changes
↓
"do the crazy commiting"
↓
agent generates many clean commits
↓
push
Enter fullscreen mode Exit fullscreen mode

And something interesting happened.

Management was happy

A few weeks later, my numbers had improved.

Management noticed.

The KPI dashboard was happy.

Therefore, apparently, my performance had improved.

Except nothing meaningful had changed.

I hadn't become a better engineer.

I hadn't started delivering dramatically more value.

I hadn't suddenly become more productive.

I had simply optimized my behavior for the measurement system.

Same engineer.

Same work.

Same code.

More commits.

Better performance according to the dashboard.

That was probably the clearest demonstration of Goodhart's Law I've personally experienced:

When a measure becomes a target, it stops being a good measure.

The funny part is that I didn't even need to game the system dishonestly.

I automated the completely legitimate act of making smaller commits.

The metric did the rest.


A commit has no unit of value

The fundamental problem with commit count is that a commit doesn't represent a standardized amount of work.

These are both one commit:

Fix typo in error message
Enter fullscreen mode Exit fullscreen mode

and:

Migrate authentication subsystem to new identity provider
Enter fullscreen mode Exit fullscreen mode

One might represent thirty seconds of work.

The other might represent three weeks of investigation, architecture discussions, implementation, testing, deployment planning, and coordination.

Git doesn't care.

The counter goes:

1
1
Enter fullscreen mode Exit fullscreen mode

You can also take exactly the same change and represent it as:

1 commit
Enter fullscreen mode Exit fullscreen mode

or:

4 commits
Enter fullscreen mode Exit fullscreen mode

or:

17 commits
Enter fullscreen mode Exit fullscreen mode

without changing the final state of the repository at all.

That's not a great property for something you're using to evaluate human performance.


It gets worse for senior engineers

Commit count becomes even less meaningful as engineers become more senior.

Imagine a staff engineer spends an afternoon helping three teams avoid a bad architectural decision.

The outcome could be:

commits: 0
Enter fullscreen mode Exit fullscreen mode

Another engineer spends that afternoon implementing the agreed solution.

Their outcome might be:

commits: 6
Enter fullscreen mode Exit fullscreen mode

Who contributed more?

That's not even the right question.

They contributed in different ways.

A technical lead might spend time:

  • reviewing pull requests
  • mentoring engineers
  • designing systems
  • investigating production incidents
  • coordinating migrations
  • reducing technical risk
  • challenging unnecessary complexity
  • helping someone else solve a problem
  • preventing bad code from being written in the first place

The best architectural decision of the week might result in:

0 lines added
0 commits
0 pull requests
Enter fullscreen mode Exit fullscreen mode

because someone realized that a new service didn't need to exist.

That's still engineering.

Sometimes the most valuable code is the code you successfully convince the team not to write.


The metric changes the behavior

This is the part that worries me more than the accuracy of the dashboard.

People adapt to incentives.

If engineers know commits are measured, eventually they start optimizing commits.

If lines of code are measured, they'll produce more lines.

If story points are measured, suddenly every task becomes an eight.

If pull requests are measured, features get fragmented into unnecessary PRs.

If tickets closed are measured, tickets get smaller.

You haven't necessarily improved productivity.

You've improved people's ability to produce the thing being counted.

My agent skill just made the feedback loop extremely obvious.

I effectively automated KPI optimization.

And the dashboard couldn't tell the difference between:

engineer became more productive
Enter fullscreen mode Exit fullscreen mode

and:

engineer learned how the measurement works
Enter fullscreen mode Exit fullscreen mode

That is a serious problem if the number is being used for performance management.


Activity metrics aren't necessarily useless

This doesn't mean commit data has zero value.

A sudden change in activity can be useful context.

If someone normally contributes regularly and disappears from the repository for three weeks, it might be worth asking:

"What are you working on?"

Maybe they're stuck.

Maybe they're doing architecture work.

Maybe they're mentoring a new engineer.

Maybe they're investigating production issues.

Maybe their work is happening in another repository.

Maybe they're disengaged.

The commit graph can help start the conversation.

The problem begins when we skip the conversation and let the graph become the conclusion.

signal → question → context → understanding
Enter fullscreen mode Exit fullscreen mode

is useful.

signal → performance score
Enter fullscreen mode Exit fullscreen mode

is dangerous.


What should we measure instead?

There isn't one magical replacement metric.

That's kind of the point.

Software engineering is a complex system.

Frameworks such as DORA focus much more on how effectively software moves through an organization: delivery speed, reliability, recovery, and the performance of the delivery system.

The SPACE framework explicitly treats developer productivity as multidimensional, covering things like satisfaction, performance, activity, communication, collaboration, efficiency, and flow.

Notice the important difference.

They aren't trying to find:

developer_productivity = some_counter
Enter fullscreen mode Exit fullscreen mode

because that number probably doesn't exist.

For individuals, I'd much rather have conversations around things like:

  • What did you help deliver?
  • What difficult problems did you solve?
  • What risks did you reduce?
  • What did you improve?
  • Who did you unblock?
  • What knowledge did you spread?
  • Did your decisions make the team faster?
  • Did the system become easier or harder to maintain?
  • Did the product become better?

Those questions are messier than a dashboard.

They're also much closer to the work we're actually trying to evaluate.


AI makes this problem even more interesting

There's another reason I've been thinking about this lately.

AI coding agents make activity metrics even easier to inflate.

Generating:

  • commits
  • pull requests
  • lines of code
  • tests
  • documentation
  • tickets

is becoming increasingly cheap.

If an organization measures engineering productivity primarily through visible repository activity, agents can generate enormous amounts of apparently productive activity.

That doesn't necessarily mean more value is being delivered.

In my case, the agent wasn't even writing more code.

It was simply reorganizing the same code into a shape that made the KPI dashboard happier.

We're entering an era where optimizing superficial engineering metrics can literally be delegated to an agent.

That should probably make us reconsider the metrics.


The weirdest part

I originally created crazy-commiting partly because I thought the situation was absurd.

But it worked.

The number went up.

The dashboard showed improvement.

Management was happy.

And that might be the best test of a bad KPI I've found so far:

If I can improve the metric significantly with an agent without improving the product, the team, or the engineering outcome, what exactly is the metric measuring?

Probably not productivity.


This is part of my Professional Antipatterns series, where I write about strange, counterproductive, and occasionally absurd things I've encountered in software engineering — without naming the companies or people involved.

The original post is also available on my blog:

https://blog.lezli01.is-a.dev/blog/professional-antipatterns-commit-count/

If you've encountered similarly questionable engineering KPIs, I'd genuinely like to hear the story.

Because somewhere out there, a dashboard is probably congratulating someone for generating their 47th commit today.

Top comments (0)