DEV Community

Said Olano
Said Olano

Posted on

Good Practices in Escalations and How to Manage Them

An escalation is not a failure. It is a mechanism that moves a problem to someone with more authority, more context, or more capacity. Done well, escalations protect customers, shorten incidents, and keep teams from burning out on problems they cannot solve alone. Done badly, they become a blame ritual, a queue of frustrated people, or a way for everyone to avoid owning a problem.

This article describes good practices for managing escalations in software and operations: how to define them, how to route them, how to keep them moving, and how to learn from them. It is written for engineering managers, support leads, and on-call engineers. It includes a small Java model that applies an escalation policy to open issues and flags the ones that need attention.

Define what an escalation is before you need one

Most escalation problems start before the first escalation. Teams that cannot agree on what counts as an escalation will argue during every incident.

A clear definition answers three questions:

  • What triggers an escalation? Examples include a missed response time, customer impact above a threshold, a security concern, or a request for a decision the current owner cannot make.
  • Who receives it at each level? The first level is usually the owning team. Higher levels might be a team lead, an engineering manager, a director, or an on-call incident commander.
  • What is expected at each level? An escalation to the next level should come with a clear statement of the problem, what has been tried, and what decision is needed.

Write this down in one page and make it visible. People escalate more confidently when they know the path and the criteria.

Separate severity from escalation level

A common mistake is to treat escalation as a reward for urgency or a punishment for failure. Severity describes the impact of an issue. Escalation level describes who needs to be involved. They are related, but they are not the same.

A severe issue handled well by the owning team may never leave that team. A minor issue that the team cannot resolve, perhaps because it needs access to another system, may escalate quickly. Keeping these concepts separate makes the process fairer and more predictable.

Make the route explicit and time-bound

Escalations stall when nobody knows who should act next or how long they have to act. Two elements fix most of this:

An owner at every level. Each escalation level should name a person or role, not a team channel that everyone assumes someone else is watching.

A time limit at every level. If the first owner does not acknowledge or make progress within a defined window, the issue moves up automatically. The time limit should depend on severity. A critical customer outage might escalate after fifteen minutes; a low-priority request might wait a day.

Automatic escalation removes the awkward decision of whether to bother a manager. It also means that nobody has to remember to check the queue.

Write escalations so the receiver can act

The quality of an escalation determines how quickly it is resolved. A useful escalation message contains:

  1. The problem in one sentence, including the affected customer, system, or process.
  2. The impact, stated in measurable terms when possible: how many users, how much revenue, how long.
  3. What has been tried, with results.
  4. What decision or help is needed, stated as a specific request.
  5. The next checkpoint, so the receiver knows when they will hear more.

An escalation that says "the payments service is slow, please help" forces the receiver to start from zero. One that says "payments p95 latency is 4 seconds for 30 percent of checkouts since 14:10; we restarted the cache with no change; we need approval to fail over to the secondary region" can be acted on immediately.

Keep the owner accountable while the escalation is active

When an issue moves up, a common failure is that the original owner stops working on it. The receiving manager assumes the team is handling it, and the team assumes the manager has it. The result is that nobody is working on the problem.

Good practice is to define two roles explicitly: the owner, who keeps driving the work, and the decision maker, who unblocks it. These can be different people. The owner should continue to post updates until the issue is closed.

Communicate on a fixed rhythm

Silence is the most damaging feature of a long escalation. Stakeholders who do not hear from the team fill the gap with worst-case assumptions and call the people doing the work.

Set an update cadence based on severity and state it when the escalation starts: every thirty minutes for an active outage, every day for a customer dispute, every week for a long-running risk. Then keep to it, even when the update says "no change, still investigating." That line is information.

Close the loop and learn

Every significant escalation should end with a short review. The purpose is not to assign blame. It is to answer three questions:

  • Did the escalation route and time limits work as designed?
  • Was the escalation message clear enough for the receiver to act?
  • What could have resolved the issue at a lower level, and what would make that possible next time?

Track simple measures across escalations: how many reached each level, how long each level held the issue, and how many were resolved without leaving the owning team. A rising number of escalations to the top level usually means the lower levels lack authority, information, or skills, not that the people there are doing something wrong.

A small model for escalation policy

The following Java model applies an escalation policy to a set of open issues. Each level has a time limit, and an issue that has been waiting longer than its level allows moves to the next one. The model returns the issues that need attention and the level they should be at.

import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
import java.util.List;

public class EscalationPolicy {

    public enum Severity { LOW, MEDIUM, HIGH, CRITICAL }

    public record Level(String name, String owner, Duration maxWait) {

        public Level {
            if (maxWait.isNegative() || maxWait.isZero()) {
                throw new IllegalArgumentException("maxWait must be positive for level " + name);
            }
        }
    }

    public record Issue(String id, Severity severity, Instant lastUpdate, int currentLevel) {}

    public record Action(String issueId, int fromLevel, int toLevel, String ownerToNotify, String reason) {}

    private final List<Level> levels;

    public EscalationPolicy(List<Level> levels) {
        if (levels.isEmpty()) {
            throw new IllegalArgumentException("policy needs at least one level");
        }
        this.levels = List.copyOf(levels);
    }

    /** Severity shortens the wait at each level: critical issues move up faster. */
    private Duration effectiveWait(Level level, Severity severity) {
        return switch (severity) {
            case CRITICAL -> level.maxWait().dividedBy(4);
            case HIGH -> level.maxWait().dividedBy(2);
            case MEDIUM, LOW -> level.maxWait();
        };
    }

    public List<Action> evaluate(List<Issue> issues, Instant now) {
        List<Action> actions = new ArrayList<>();
        for (Issue issue : issues) {
            int level = issue.currentLevel();
            if (level < 0 || level >= levels.size()) {
                throw new IllegalArgumentException("invalid level for issue " + issue.id());
            }
            Duration waited = Duration.between(issue.lastUpdate(), now);
            Duration allowed = effectiveWait(levels.get(level), issue.severity());

            if (waited.compareTo(allowed) > 0 && level + 1 < levels.size()) {
                Level next = levels.get(level + 1);
                actions.add(new Action(
                        issue.id(),
                        level,
                        level + 1,
                        next.owner(),
                        "no update for " + waited.toMinutes() + " min at " + levels.get(level).name()
                                + " (limit " + allowed.toMinutes() + " min)"));
            } else if (waited.compareTo(allowed) > 0) {
                actions.add(new Action(
                        issue.id(),
                        level,
                        level,
                        levels.get(level).owner(),
                        "already at top level and stalled; decision needed"));
            }
        }
        return actions;
    }

    public static void main(String[] args) {
        EscalationPolicy policy = new EscalationPolicy(List.of(
                new Level("team", "on-call engineer", Duration.ofHours(1)),
                new Level("lead", "team lead", Duration.ofHours(2)),
                new Level("manager", "engineering manager", Duration.ofHours(4))));

        Instant now = Instant.parse("2026-10-10T12:00:00Z");
        List<Issue> issues = List.of(
                new Issue("INC-101", Severity.CRITICAL, now.minus(Duration.ofMinutes(40)), 0),
                new Issue("INC-102", Severity.MEDIUM, now.minus(Duration.ofHours(3)), 1),
                new Issue("INC-103", Severity.LOW, now.minus(Duration.ofMinutes(30)), 0));

        for (Action a : policy.evaluate(issues, now)) {
            System.out.printf("%s: level %d -> %d, notify %s (%s)%n",
                    a.issueId(), a.fromLevel(), a.toLevel(), a.ownerToNotify(), a.reason());
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

With the sample data, the critical issue INC-101 has waited 40 minutes at the team level. Because critical issues use a quarter of the limit, its allowed wait is 15 minutes, so it moves to the lead. INC-102 has waited three hours at the lead level, whose limit is two hours, so it moves to the manager. INC-103 is within its limit and does not appear. The useful property here is that the policy is explicit: anyone can read why an issue moved, and the reason is recorded in the action.

The model is intentionally simple. A real policy would also consider business hours, on-call schedules, and whether the escalation was acknowledged. Those details matter, but they should be added deliberately, not left implicit.

Common mistakes

Several patterns consistently make escalations worse.

Escalating as a way to hand off. If an escalation means "this is no longer my problem," the owner stops working on it. Escalate to get help, and keep ownership until the problem is resolved.

Escalating without context. A receiver who has to ask basic questions adds delay. Include what has been tried.

Punishing escalation. If people are criticized for escalating, they will wait too long. Make escalation a normal, expected part of the process.

Letting the top level become the default. When everything goes to the manager, the managers become the bottleneck, and the lower levels never build the skills they need.

Skipping the review. Without a review, the same escalations repeat.

Practical guidance for engineering leaders

Before the next escalation, check:

  1. Is there a written definition of what triggers each level, and do people know it?
  2. Does every level have a named owner and a time limit?
  3. Does an escalation message include the impact, what has been tried, and the specific decision needed?
  4. Is there a fixed update rhythm for active escalations?
  5. Do we review escalations to improve the lower levels, not only to find who was late?

If you cannot answer yes to most of these, the escalation process depends on individual heroics, and that will not scale.

Key takeaways

Good escalation practice combines clear criteria, named owners, time limits, and messages that let the receiver act. Keep ownership with the team that does the work, communicate on a fixed rhythm, and review each significant escalation to strengthen the levels below. Escalation is how an organization gets help to the right place at the right time, and it works best when it is designed, written down, and treated as a normal part of the work.

Top comments (0)