DEV Community

Cover image for How to Assess Production Reliability Before Deployment
Parsa Mohammadi
Parsa Mohammadi

Posted on Originally published at tomosu.ai

How to Assess Production Reliability Before Deployment

Most teams assess a change using code review and tests.

That catches a lot. But it doesn't answer everything.

A change can pass the test suite and still create problems in production because of its dependencies, how the affected code is used, recent incidents, or the size of its potential blast radius.

So the question before deployment is not just:

Does this code work?

It is:

How is this change likely to behave in production?

Start with the change itself

First, understand what the change actually touches.

Look at:

  • Files and components changed
  • APIs modified
  • Database changes
  • Configuration changes
  • Authentication or authorization logic
  • Shared libraries
  • Customer facing functionality
  • Infrastructure changes

A small diff isn't necessarily a small production change.

Look at dependencies

Next, ask what depends on the code being changed.

Change
  ↓
Dependencies
  ↓
Affected Components
  ↓
Production Usage
Enter fullscreen mode Exit fullscreen mode

The further a change can propagate, the more context you need before deploying it.

Check the testing evidence

Tests provide evidence about a change. They don't prove that every production scenario will behave correctly.

Look at unit, integration, end to end, regression, and failure tests.

Then ask:

What important production behavior is still unverified?

Look at production behavior

Production context can change how a change should be evaluated.

Consider:

  • Real customer traffic
  • Large datasets
  • Concurrent requests
  • Production configuration
  • External service failures
  • Unusual inputs

Check incident and change history

Look at previous incidents, rollbacks, related fixes, change frequency, and repeated failures.

History doesn't decide whether a change is safe.

It provides context.

Consider code volatility

A component that changes constantly can require different attention from one that has remained stable for months.

Look at how often it changes and whether previous changes caused problems.

Assess deployment and rollback conditions

Ask:

  • Can the change be rolled back quickly?
  • Can it be deployed gradually?
  • Is a feature flag available?
  • Are database changes reversible?
  • Can the affected functionality be disabled?

A simple framework

Change
  ↓
Evidence
  ↓
Uncertainty
  ↓
Action
Enter fullscreen mode Exit fullscreen mode

Understand the change.

Collect evidence.

Identify what remains uncertain.

Then decide what additional action makes sense.

Pre deployment checklist

  • [ ] What changed?
  • [ ] What components are affected?
  • [ ] What dependencies are involved?
  • [ ] Is the important behavior tested?
  • [ ] Has this area caused incidents before?
  • [ ] How large could the blast radius be?
  • [ ] Can the change be rolled back?
  • [ ] What monitoring will be available after deployment?

Where PRI fits

Tomosu's Production Reliability Index (PRI) brings multiple reliability signals together to give engineers a change level signal before deployment.

The goal is not to replace engineering judgment.

It is to surface useful evidence earlier.

Run a PRI assessment:

https://tomosu.ai/start

Top comments (0)