Apparently knowing the entire codebase isn’t the same as knowing what actually happened.
Coding agents have reached a slightly ridiculous level.
I can point one at a repo and say:
add rate limiting to this endpoint
and it will happily wander through files I forgot existed, figure out how everything connects, write the implementation, update the tests and hand me a diff.
Great.
Then I give it a real production bug:
some users can’t finish checkout
And suddenly we’re both detectives.
Which users?
Not sure.
What happens when they try?
Not sure.
Any errors?
Probably.
Can you reproduce it?
Of course not.
Wonderful.
The agent knows 40,000 lines of my codebase.
It just doesn’t know what happened 30 seconds ago.
GitHub knows the code. It doesn’t know the crime scene.
This is the part I hadn’t really thought about until recently.
We’ve spent an absurd amount of effort making AI understand codebases.
And it’s working.
Give a coding agent your repository and it can understand architecture, follow dependencies, inspect git history, find relevant functions and make surprisingly decent changes.
But production bugs have this annoying habit of containing information that isn’t in the repository.
Take a checkout bug.
The code might look completely reasonable.
The tests pass.
The agent reads it.
Also reasonable.
Meanwhile, one actual user did this:
Opened checkout
Changed quantity
Applied coupon
Removed item
Clicked Pay
Request timed out
Clicked Pay again
JavaScript error
Left
Oh.
That’s useful.
None of it was in GitHub.
I think we’ve been giving coding agents half the story
A repository tells an agent how the application is supposed to behave.
Production tells it how the application actually behaved.
Those are not always the same thing.
Anyone who has shipped software knows this because users possess a supernatural ability to find states that nobody considered.
Your test:
Add item → Checkout → Pay
Your user:
Add item
Remove item
Add it again
Open another tab
Come back 17 minutes later
Change Wi-Fi networks
Double-click Pay
Somehow summon a state that should not exist
Then they close the tab.
No bug report.
No explanation.
Just vibes.
And a conversion rate that’s mysteriously worse today.
This is how I ended up looking at HeronSignal
I initially thought the interesting part was session replay.
It records what happened around real user sessions, alongside things like errors, network activity and performance.
Useful.
But session replay isn’t exactly a new concept.
Then I noticed what they were doing with that data.
Instead of treating production context as something only a developer reads, they’re making it available to the agent too.
That’s much more interesting.
Because now the conversation with the coding agent changes.
Instead of:
Users say checkout is broken. Please investigate.
you can effectively give it:
_Here’s the affected session.
Here’s what the user did.
Here’s the error.
Here’s the failed request.
Here’s the stack trace.
Here’s the relevant deployment.
And here’s the repository._
Now go investigate.
That’s a completely different starting point.
And then I found this button.
This was the point where HeronSignal stopped looking like another observability product to me.
After investigating an interaction, HeronAgent can do this:
Yep.
It can go from the production interaction to the repository, investigate the relevant code, prepare a fix and open the PR.
The first time I saw this my reaction was basically:
Wait. We’re letting the monitoring tool code now?
Kind of.
But there’s an important detail.
It doesn’t merge anything.
You still get the diff.
You still review it.
You still decide whether the AI has brilliantly fixed the problem or confidently created a completely new one.
I appreciate that distinction.
This is where things get weird
Think about what just happened.
Monitoring tools traditionally end here:
Something broke.
Good luck.
Okay, maybe that’s unfair.
They give you a stack trace too.
Something broke.
Here's a stack trace.
Good luck.
We got better dashboards.
Better logs.
Better alerts.
Better session replay.
But the fundamental contract stayed roughly the same:
The tool gives you evidence. You investigate.
An agent changes that contract.
The evidence can now be consumed by something capable of actually doing part of the investigation.
And potentially acting on it.
That’s not really “better monitoring.”
That’s monitoring becoming part of the development loop.
The PR is actually the boring part
Opening a pull request sounds like the headline feature.
I don’t think it is.
We already have agents that can write code.
Give Cursor, Claude Code or another capable coding agent enough repository context and creating a PR isn’t particularly shocking anymore.
The interesting part is everything that happened before the PR.
The agent didn’t start with a Jira ticket somebody wrote three days later.
It started with what happened in production.
That’s the shift.
We spent the last couple of years connecting AI to our code.
Now we’re starting to connect it to the consequences of our code.
That feels much bigger.
Your next bug report might not be written by a human
This is the part I keep thinking about.
A traditional bug report is basically a human trying to reconstruct production:
User says checkout didn’t work.
Then somebody asks for browser.
Then screenshots.
Then reproduction steps.
Then logs.
Then timestamps.
Then someone tries to figure out which deployment was live.
It’s archaeology.
But production already witnessed the event.
Why are we asking humans to reconstruct it?
A system watching the session could theoretically assemble something much better:
Checkout failed for 23 sessions.
Started after deployment #1842.
21/23 affected sessions followed the same sequence:
Apply coupon
→ change quantity
→ submit payment
Associated error:
TypeError in checkout.ts:241
Business impact:
18 users abandoned checkout.
Relevant files:
checkout.ts
cart.ts
coupon.ts
That is a better bug report than I have written in my entire career.
And nobody had to write it.
Then the agent reads it.
Now things start connecting.
Production finds the issue.
Production provides the evidence.
The agent investigates.
The repository provides implementation context.
The agent proposes the change.
Developer reviews.
We might be building software backwards soon
Right now the normal loop looks like:
Human notices problem
↓
Human gathers evidence
↓
Human explains problem to AI
↓
AI investigates code
↓
AI proposes fix
There’s a very obvious redundant component in there.
Us.
Not the decision-making part.
The courier part.
The part where we spend 45 minutes collecting information from five tools just so we can paste it into the sixth.
That is exactly the kind of work machines should eat.
A much more interesting loop is:
Production notices problem
↓
Agent gets evidence
↓
Agent investigates
↓
Human gets proposed fix
↓
Human decides
That doesn’t remove developers.
It removes detective work that developers didn’t particularly enjoy doing anyway.
There is one thing I definitely don’t want
Fully autonomous production fixes.
Please don’t.
If my checkout breaks at 3 AM, I want AI to investigate it.
I want it to find the suspicious commit.
I want it to reproduce what it can.
I want it to prepare the fix.
I would absolutely love to wake up to:
“This broke overnight. Here’s why. PR is ready.”
I do not want to wake up to:
“This broke overnight. I changed 14 files and deployed it. Sleep well.”
Different product.
Very different anxiety level.
The PR boundary is actually a pretty sensible place to stop.
The real upgrade isn’t AI writing more code
We’re already getting very good at that.
The next useful upgrade is AI understanding why the code it already wrote failed when it met reality.
That’s why the HeronSignal idea caught my attention.
Not because it has AI.
Everything has AI now.
Not because it can generate code.
That’s becoming normal too.
It’s because it connects two worlds that have mostly been separate:
the repository
and
what users actually experienced.
Once agents can see both, debugging gets a lot more interesting.
And maybe one day this:
“Can you reproduce it?”
finally dies.
I won’t miss it.



Top comments (0)