DEV Community

Cover image for What Is Production Reliability?
Parsa Mohammadi
Parsa Mohammadi

Posted on Originally published at tomosu.ai AI-assisted

What Is Production Reliability?

A change can pass tests, pass code review, and still cause problems in production.

That is because production reliability is about more than whether code works in isolation.

It is about how a change behaves when it meets the real system: real traffic, real dependencies, real data, real configuration, and real users.

Production reliability is the ability of a software change to behave correctly and consistently under real production conditions.

Why production reliability is different from test results

Tests tell you what happened under the conditions you tested.

Production introduces conditions that are difficult to reproduce completely:

  • Different traffic patterns
  • Larger datasets
  • External dependency failures
  • Configuration differences
  • Concurrent requests
  • Unexpected inputs
  • Existing production behavior

A change can therefore have strong test evidence while still having meaningful uncertainty around production behavior.

What should you look at?

A useful assessment starts with the change itself and then adds context.

1. Change scope

What files, services, APIs, databases, or infrastructure are affected?

The size of a diff is useful context, but it does not tell you the full impact.

2. Dependencies

What depends on the changed component?

A small change to a shared library can have a much larger production impact than a large change inside an isolated service.

3. Testing and verification

What evidence exists that the important behavior works?

Look at unit tests, integration tests, end to end tests, regression tests, and failure cases.

4. Production behavior

How is the affected component actually used?

Traffic, data volume, runtime behavior, and external integrations can all change the practical impact of a change.

5. Change and incident history

Has this area caused incidents, rollbacks, or repeated regressions before?

History does not determine whether a change is safe, but it is useful context.

6. Deployment and rollback

What happens if the change behaves badly?

Can it be rolled back quickly? Can it be deployed gradually? Is a feature flag available?

Production reliability vs code review

Code review asks whether the implementation makes sense.

Production reliability asks what the implementation means for the running system.

You generally want both.

Read more: https://tomosu.ai/blogs/production-reliability-vs-code-review.html

Production reliability vs observability

Observability tells you what is happening in a running system.

Production reliability adds change context: what changed, what it affects, and what could happen when it reaches production.

Observability is therefore one useful input into reliability assessment, not a replacement for it.

Read more: https://tomosu.ai/blogs/production-reliability-vs-observability.html

A practical model

A useful way to think about the assessment is:

Change
  ↓
Evidence
  ↓
Uncertainty
  ↓
Action
Enter fullscreen mode Exit fullscreen mode

The goal is not to eliminate uncertainty before every deployment.

The goal is to find changes where the potential production impact and remaining uncertainty justify additional attention.

Where PRI fits

Tomosu's Production Reliability Index (PRI) brings multiple reliability signals together, including:

  • Fragility
  • Drift
  • Governance Compliance
  • Runtime Signals
  • Code Volatility
  • Deployment Velocity
  • Escalation

The number is useful as a summary, but the underlying evidence matters more than the number by itself.

The main idea

Production reliability is not a property of a diff alone.

It is a property of a change in context.

Change + dependencies + testing + production behavior + history + deployment conditions give engineers a much better picture of what may happen after a change ships.

Run a Production Reliability Index assessment:

https://tomosu.ai/start

Top comments (0)