DEV Community

Cover image for Production Reliability vs Code Review
Parsa Mohammadi
Parsa Mohammadi

Posted on Originally published at tomosu.ai AI-assisted

Production Reliability vs Code Review

Code review is one of the most established ways engineers evaluate software changes.

But code review and production reliability are not the same thing.

A reviewer can determine that an implementation is correct while still not knowing exactly how the change will behave across production dependencies, traffic, configuration, and existing system behavior.

The two questions are different:

Code review: Is the implementation correct?

Production reliability: What does this change mean for the running system?

What code review is good at

Code review can catch problems such as:

  • Incorrect logic
  • Bad API usage
  • Security issues
  • Poor error handling
  • Maintainability problems
  • Unclear implementation choices
  • Violations of project conventions

It is primarily focused on the implementation.

What production reliability adds

A production reliability assessment adds context around the change.

It asks:

  • What components are affected?
  • Which dependencies are involved?
  • How is this code used in production?
  • What testing covers the change?
  • Has this area caused incidents before?
  • How large could the blast radius be?
  • How easy is the change to roll back?

A reviewer may understand the code perfectly and still lack some of this context.

Example: a shared database library

Imagine a small change to a shared database library.

The diff is straightforward.

The code review looks good.

But the library is used by 40 services, several of those services handle customer transactions, and the affected query path has caused performance problems before.

The implementation may still be correct.

The production context simply makes the change more important to investigate.

Blast radius matters

The number of changed lines is not the same thing as production impact.

A small change in a shared component can have a large blast radius.

A large change in an isolated internal service may have a much smaller one.

That is why dependency and production usage are important parts of change assessment.

What about AI generated code?

AI assisted development makes this distinction more important.

AI can increase the amount of code produced and the number of changes entering review.

Code review still matters.

But the reviewer may need more context to understand what a generated change could affect.

Production reliability can provide another layer of evidence around the change.

Code review and observability

Observability answers another question: what is happening after the system is running?

Production reliability connects change context with production evidence before and around deployment.

They work together rather than replacing each other.

A practical workflow

Code change
    ↓
Code review
    ↓
Production context
    ↓
Dependencies + testing + history
    ↓
Deployment decision
Enter fullscreen mode Exit fullscreen mode

The goal is not to replace code review.

It is to make the decision around the change more informed.

Where PRI fits

The Production Reliability Index (PRI) is designed to summarize multiple signals around a software change so engineers can identify where additional investigation may be useful.

The score is not a substitute for reviewing the underlying findings.

It is a way to bring the signals into one place.

The main idea

Code review asks whether the implementation makes sense.

Production reliability asks what the implementation could mean in production.

You need both perspectives when the goal is to understand a change before it reaches users.

https://tomosu.ai/start

Top comments (0)