DEV Community

Wilson Santos
Wilson Santos

Posted on Originally published at draugr.dev

What to fix first when everything is critical

Run any scanner across a real system and you will get the same shape of result, a wall of
findings, most of them stamped "high" or "critical." Those labels come from
CVSS, the Common Vulnerability Scoring System, a 0 to 10 scale
published with most vulnerabilities. The word is a band on that scale, and the bands are defined by
FIRST in the CVSS v3.1 specification:

The label The score What it took to earn it
Critical 9.0 to 10.0 reachable over a network, by somebody with no account, needing nothing from a user, and it takes confidentiality, integrity and availability together
High 7.0 to 8.9 most of the above, with one condition attached. An account is needed, or the attack is not straightforward, or one of the three impacts is spared
Medium 4.0 to 6.9 several conditions at once, or a partial impact. The bulk of what a scanner returns
Low 0.1 to 3.9 local access, or heavy prerequisites, or an impact that is real and narrow

That is a useful thing to know and it is not a statement about you. The specification is explicit
about it: a base score reflects severity "assuming the reasonable worst case impact across
deployed environments"
. So AV:N/PR:N, the pair that makes a 9.8, is a claim that somebody's
deployment exposes this to the internet without authentication. It is not a claim that yours does.

So the score did its job: it told you how severe each issue is in the abstract. What it did not tell
you is the only thing you need on Monday morning: which one do I fix first?

Severity is not priority. And the gap between them is where security programs quietly fail.

Why severity can't rank your work

CVSS scores a vulnerability in isolation, as if every place it appears were identical. But
the same CVE is not the same risk everywhere:

The same CVE, sitting on What it actually is
an internet-facing service with no auth a live attack path
an internal tool three network hops away, behind a policy a someday-maybe
the component that takes the platform with it when it falls over an emergency, whatever the score says

Feed all of that into one severity number and you get the two failure modes every team knows:
either you chase everything marked critical and burn out, or you stop believing the labels
and route around the gate. Both end in the same place, with the urgent thing sitting in a
backlog next to a hundred things that look exactly as scary.

The two questions a scanner can't answer

To rank findings you need two pieces of context that no scanner can compute, because they
aren't properties of the code:

How exposed is it? Reachable from the internet, or namespace-scoped behind a network policy? This is likelihood, meaning how reachable the weakness is at all.

How much would it cost? A platform-wide outage, or a dev tool nobody would notice for a week? This is impact, meaning what it costs when it goes wrong.

Real risk is the product of the two, layered on top of raw severity. A scanner sees none of it.

But your app description does

You already know this context. You know which services face the world and
which are buried. You know which ones are load-bearing. That knowledge is stable, small, and
exactly the kind of thing Draugr already captures
in the app descriptor.

So it becomes two more attributes on a component you already describe, its exposure and its
business criticality, and prioritization falls out of the description you wrote once.

components:
  - name: ingress-gateway
    exposure: public          # reachable from the internet
    criticality: critical     # failure = platform-wide outage
  - name: dev-dashboard
    exposure: restricted      # internal, network-scoped
    criticality: supporting   # nobody pages at 3am for this
Enter fullscreen mode Exit fullscreen mode

Same finding, same severity, two answers, because the components differ.

Now the same CVE resolves two different ways, so on the gateway it is fix-now and top of the list.
On the dashboard it is track-and-move-on. Severity was identical and priority was not, because
priority combines the finding's severity with the
component's exposure and criticality into a single, ordered answer.

The third input, whether anyone is exploiting it

Exposure and criticality describe your side. There's one more signal, and it comes from the
world: exploitability.

A CVE on CISA's KEV catalog is known to be exploited in the wild, not
theoretically exploitable, and used against real systems. That escalates a finding whatever
its CVSS score says, because a modest score being exploited today outranks a 9.8 nobody has ever
weaponized. A high EPSS probability, the likelihood a CVE gets exploited
in the next thirty days, bumps a finding up a band on the same logic.

The severity did not change and neither did the code. What changed is that one of them is being used, which is a fact about the world rather than about your repository, and it is the input almost nobody feeds in.

Both are opt-in and read from feeds you provide, so the ranking stays reproducible, because the same
inputs give the same order, and you can point at where each escalation came from.

An order is still a list, until it names the work

Ranking gets the right thing to the top and shortens nothing: four hundred findings ranked are
still four hundred rows, and the first twelve are often one library. So the last step groups them
by the thing that fixes them:

$ draugr scan . --view actions

WHAT TO DO  10 actions clear 508 findings
  P1  Update python:3.8-slim
      control images · 465 findings · upstream · CVE-2026-42010 +464
  P1  Upgrade jquery 1.8.3
      control sca · 16 findings · web/package-lock.json:10 · and 2 more · CVE-2020-11023 +15
  P1  Upgrade Jinja2 2.10
      control sca · 6 findings · app/requirements.txt:5 · CVE-2019-10906 +5
  ⋯
Enter fullscreen mode Exit fullscreen mode

That is the demo project, which carries 1,091 findings. Ten actions clear 508 of them, and the
first row is one base image carrying 465 on its own.

A row is a remediation, not a finding. Six vulnerabilities in one library are one upgrade; the
same misconfiguration in three Dockerfiles is one habit. Listing those separately makes the
repetitive work crowd out everything else, which is how a ranked list still ends up unread.

Actions are ordered by the worst thing each one clears, and only then by how many. An action
clearing a single P1 outranks one clearing forty P4s, because volume is not a reason to do
something first and a P1 is not something to trade away for a bigger number.

upstream on that first row means the team does not build that image, so the action is to take a
newer one rather than to upgrade a package nobody there can reach. Advice you cannot act on, at the
top of a list called fix first, teaches people the list is not worth reading, and that is a fact
about a contract rather than something a scanner can see. It comes from the descriptor, like
everything else here.

Grouping is opt-in while that annotation spreads: --view actions turns it on, and the default
still lists one finding per row. The report files always carry findings separately either way, so
an auditor reading the SARIF sees one record per finding whichever way the console was asked to
show them.

What that changes

For a developer

The output stops being a 400-row dashboard and becomes a short list: these three, in this order, today. Clear action, not triage archeology.

For the business

Consistent and defensible, because the same rules apply to every service. "What is our real exposure right now?" has an answer, and the urgent gets a faster response because it is no longer buried under look-alikes.

For your coding assistant

The difference between a list and an order. Point one at Draugr and it answers from the same ranking your pipeline uses. A scanner can hand it a pile of findings; only your description says which one to fix.

This is what Draugr does. Not another dashboard that shows you more, but a gate that tells you
less, being the few things that matter, ranked by your own context. Describe your app, and
the description doesn't just decide which scanners run. It decides what you fix first.

Top comments (0)