security SLA tracking
developer security metrics
measure DevSecOps performance
engineering KPI security
gamify vulnerability management
DevSecOps KPIs
engineering security metrics
MTTR security metrics
mean time to remediate
security SLA breach rate
developer SLA compliance
secure coding KPIs
measuring security hygiene
developer performance metrics
ethical security metrics
DevSecOps operational pitfalls
toxic engineering metrics
perverse incentives software development
fair developer metrics
secure developer culture
Integrating Security SLAs into Developer Metrics Ethical and Operational Pitfalls
Back to blog
The Rise and Fall of MTTR as a Developer Security Metric
A quick correction: "MTTR" in DORA is not vulnerability remediation
Applying an average to squad performance is flawed
The Ethical and Operational Pitfalls of Raw Speed Metrics
- Rushing risky code and "Band-Aid" fixes
- The illusion of security: measuring patching, not verification
- Ticket gaming and time manipulation
- Alert fatigue, burnout, and a toxic culture A Healthier Alternative: SLA Breach Rate by Squad Why it is better What "risk-adjusted" looks like in practice How to Safely Gamify Vulnerability Management with InstaSLA
- Standardize risk-based deadlines
- Group findings with Fix Campaigns to keep it fair
- Leaderboards built on reliability, not speed (with guardrails)
- Executive visibility and contextual exceptions Implementing a Balanced DevSecOps Measurement Program Conclusion Sources Integrating Security SLAs into Developer Metrics: Ethical and Operational Pitfalls In the race to ship software faster, engineering organizations have adopted automated security scanning, shifting security left into the IDE and the CI/CD pipeline. But as tools multiply and alerts surge, engineering managers face a hard question: how do you actually measure DevSecOps performance?
The knee-jerk answer from leadership has been to borrow operations metrics. Mean Time to Remediate (MTTR) became the default engineering KPI for security, because it is a single number that fits neatly on a board slide.
The problem is that when you tie performance reviews, bonuses, or team standing to a raw speed metric, you create perverse incentives. Developers under pressure to hit arbitrary deadlines rush risky fixes, game the ticketing system, or burn out.
This article looks at why raw speed is a poor way to grade squads, what the 2026 data says about how remediation really works, and why "SLA Breach Rate by Squad" is a fairer and safer alternative, including how to use it in a platform like InstaSLA without recreating the same problems.
The Rise and Fall of MTTR as a Developer Security Metric
Mean Time to Remediate measures the average elapsed time between discovering a vulnerability and confirming its resolution. It is easy to see why executives like it. If the MTTR for critical vulnerabilities drops from 30 days to 14, the program looks like a success.
A quick correction: "MTTR" in DORA is not vulnerability remediation
Security teams often borrow credibility for MTTR from the DORA (DevOps Research and Assessment) metrics. That borrowing is shakier than it looks.
In DORA, the "R" historically stood for recover (or restore service), meaning how fast you get back to normal after a production failure. It was never a measure of how fast you patch vulnerabilities. Vulnerability MTTR is a different metric that happens to share an acronym.
DORA has since retired the term entirely. In 2023 the metric was renamed and redefined as failed deployment recovery time, scoped to failures caused by a software change, because a single average blended together failures with very different causes.
DORA's model has kept evolving. The 2024 research added deployment rework rate as a fifth metric, and the 2025 report dropped the elite/high/medium/low tiers in favor of seven team archetypes that blend delivery metrics with human factors such as burnout and friction.
So if you say "MTTR is one of the foundational DORA metrics" in a board deck, be precise about what you mean. What you are measuring is remediation time, and it deserves its own scrutiny.
Applying an average to squad performance is flawed
Security work is not widget manufacturing. A library bump takes five minutes. A deeply embedded architectural flaw, or an XSS issue that requires significant logic changes, can take weeks of refactoring and testing. Grade developers on the average time to close tickets and you ignore the context, complexity, and risk behind the work.
There is also a bigger-picture reason to be careful with averages. The 2026 numbers show the whole industry struggling with remediation speed for reasons that have little to do with individual effort:
Verizon's 2026 DBIR found that vulnerability exploitation is now the top initial access vector, at 31% of breaches, while the median time to patch known exploited vulnerabilities rose from 32 to 43 days.
Only 26% of CISA KEV vulnerabilities were fully remediated in 2025, down from 38%, and the median organization faced 16 KEV vulnerabilities, up from 11.
Even the best-performing organizations remediate only about 30 to 40% of KEV instances in the first week, a ceiling that has barely moved regardless of tooling or maturity.
Veracode's 2026 State of Software Security report found that 82% of organizations carry security debt (flaws unresolved for more than a year) and 60% carry critical security debt.
If the industry's own median is measured in weeks or months, a squad-level MTTR target is mostly measuring backlog size, dependency depth, and change-approval friction, not developer diligence.
The Ethical and Operational Pitfalls of Raw Speed Metrics
Announce that this quarter's developer security metric is "reduce MTTR," and you invite several predictable failure modes. Goodhart's law explains all of them: when a measure becomes a target, it stops being a good measure. DORA itself lists setting metrics as a goal as a top pitfall, and it cites Goodhart's law by name.
- Rushing risky code and "Band-Aid" fixes The most dangerous consequence of speed-based KPIs is worse code. When the clock is ticking, the goal quietly shifts from "secure the application" to "stop the clock."
Imagine a static analysis tool flags a SQL injection issue. The proper fix is to move the data-access layer to parameterized queries across several microservices, roughly a five-day job. A developer judged on a 48-hour remediation target will be tempted to add a localized input-sanitization filter instead. The ticket closes, the dashboard turns green, and the root cause remains. You have traded a metric win for long-term security debt.
- The illusion of security: measuring patching, not verification Teams often equate applying a patch with remediating a risk. A developer bumps a vulnerable package, marks the ticket resolved, and the clock stops. But did the fix work? Did it break a downstream service? Did it actually close the attack path?
Without mandatory automated retesting, your MTTR measures when someone claimed a fix, not when the risk went away. The lesson for SLA design: define what "remediated" means (verified by a rescan, not just a merged pull request) before you start the clock.
- Ticket gaming and time manipulation Engineers are problem solvers, and if you give them a flawed system to be measured against, they will engineer around it. Common patterns:
Cherry-picking: grabbing easy, low-severity findings to drag the average down while hard, critical ones sit in the backlog.
Clock manipulation: starting the clock at triage or assignment instead of discovery. A critical finding can sit in a queue for two weeks and still show a 24-hour MTTR.
Premature closure: marking findings "Risk Accepted" or "False Positive" without security review just to remove them from the calculation.
The clock-manipulation problem is well understood at the policy level. CISA's BOD 26-04 starts the remediation timeline at whichever comes first: CISA adding the vulnerability to the KEV catalog, or the agency identifying it on an asset and reporting it to the CDM dashboard. Your internal SLAs should be just as explicit about when the clock starts, and it should never be a discretionary event.
- Alert fatigue, burnout, and a toxic culture Raw speed metrics also hurt morale. Modern tooling generates an enormous volume of findings, and racing a ticking clock across hundreds of them produces fatigue. Cycode's survey of security leaders, reported by ReversingLabs, found that 85% of CISOs said alert fatigue had strained the relationship between security and development teams, and that enterprise AppSec teams now juggle an average of 49 security tools.
That creates an adversarial dynamic: security is seen as generating time bombs, and developers are punished for not defusing them fast enough. DORA's own research has moved in the same direction, treating burnout and friction as first-class signals rather than side effects.
A Healthier Alternative: SLA Breach Rate by Squad
If MTTR is a poor squad-level grade, how do you hold teams accountable? Change the question from "how fast did you do it?" to "did you meet the agreed-upon standard for this risk?"
An SLA Breach Rate is the percentage of vulnerabilities that remained open past their assigned, risk-adjusted deadline.
Why it is better
It respects complexity and risk. A mature SLA framework sets different deadlines by severity, exposure, and exploitability. The industry is moving this way. CISA's BOD 26-04, issued June 10, 2026 to replace BOD 22-01 and BOD 19-02, drops CVSS-based timelines in favor of a four-variable model: asset exposure, KEV status, exploit automation, and technical impact. Depending on the combination, the deadline is as short as 3 days, or 14 or 60 days, or as late as the next scheduled upgrade. (It binds federal civilian agencies directly, but it is a useful public template for private-sector SLAs.)
It encourages thoroughness, not just speed. If a medium-severity issue has a 30-day deadline, there is no prize for a sloppy two-day fix. The goal is a correct fix before the deadline, which lets security work fit into normal sprint planning.
It promotes consistency. A squad that resolves 95% of its findings inside SLA shows predictable hygiene. That is more valuable than a squad with a wildly fluctuating average that alternates between heroics and dropped balls.
What "risk-adjusted" looks like in practice
Two reference points show how regulated environments now think about this:
CISA BOD 26-04: the highest-risk tier (publicly exposed, KEV-listed, automatable, high impact) carries a three-day deadline. According to CISA's own analysis of a large federal agency, only about 1% of vulnerability instances land in that tier, which is the point: most findings do not need an emergency clock.
FedRAMP: the traditional 30/90/180-day windows for high, moderate, and low findings are giving way to timeframes driven by potential agency impact, internet reachability, and exploitability. FedRAMP has said that mandatory adoption of its new Vulnerability Detection and Response rules takes effect December 7, 2026, in alignment with BOD 26-04. Anything not remediated or mitigated within 192 days of evaluation must be formally categorized as an accepted vulnerability.
Note that the older claim that BOD 22-01 sets a 30-day deadline for medium-severity flaws was never accurate. BOD 22-01 covered only KEV entries, and it has now been superseded. If you cite a compliance basis for your SLA tiers, cite the current one.
How to Safely Gamify Vulnerability Management with InstaSLA
Gamification works when the rules are fair and the incentives match business goals. Platforms like InstaSLA are built to move organizations from ticket-chasing toward structured SLA management. Here is how engineering managers can use one to run a healthy program.
Standardize risk-based deadlines
Define SLA timers by severity, exploitability, and asset context, and let the platform calculate the due date automatically when a finding is discovered. Every squad then operates under the same transparent rules, and the "clock starts at triage" trick disappears.Group findings with Fix Campaigns to keep it fair
Nothing ruins gamification faster than a rigged game. If one dependency update triggers 500 alerts across a squad's repositories, penalizing them for 500 separate breaches is demoralizing. Fix Campaigns group identical or related findings so the squad is measured on completing one campaign within the SLA, which reflects the real engineering effort.Leaderboards built on reliability, not speed (with guardrails)
Show the SLA Breach Rate across squads instead of a "fastest closer" ranking:
Squad Alpha: 99% SLA compliance
Squad Bravo: 94% SLA compliance
Squad Charlie: 82% SLA compliance
Because the outcome is binary (breached or not), squads are pushed to plan sprints, budget time for security debt, and ship verified fixes before the deadline.
A note of caution, though: even a good metric can be turned into a bad league table. DORA's guidance says its metrics are meant to be applied at the application or service level, and warns that comparing very different applications can mislead. Its 2023 report also cautioned that league tables produce unhealthy comparisons. Some guardrails that help:
Compare a squad against its own trend before comparing it to other squads. The squad that improved most is often not the one at the top.
Segment by portfolio. A squad owning a legacy monolith with deep transitive dependencies is not comparable to one owning a greenfield service.
Pair breach rate with a counterweight, such as reopen rate, verified-fix rate, or time-in-queue before triage, so the SLA cannot be met by cutting corners.
Watch for SLA-specific gaming, such as lobbying for longer deadlines, splitting findings, or leaning on exceptions. The metric should be reviewed and adjusted like any other.
- Executive visibility and contextual exceptions Sometimes breaching an SLA is the right business decision, for example when a fix would take a revenue-critical system offline on its busiest day. If an exception is requested and approved by leadership, it should not count against the squad's breach rate. Exceptions need an owner, a rationale, compensating controls, and an expiry date, so they are a managed decision rather than a loophole. FedRAMP's 192-day accepted-vulnerability rule shows the same principle at scale: long-lived risk is allowed, but it must be named, documented, and reported.
Implementing a Balanced DevSecOps Measurement Program
Moving from raw speed to SLA compliance is a cultural shift, and it needs active management.
Step 1: Audit your current metrics. Do you rely on MTTR? Look for signs of gaming: a high reopen rate, spikes of "false positive" closures near reporting dates, or a suspicious gap between discovery time and ticket creation.
Step 2: Define clear, attainable SLAs. Base SLAs on operational reality and current frameworks, not wishful thinking. If your critical-finding remediation today takes 60 days, mandating 24 hours overnight guarantees a 100% breach rate and immediate demoralization. Given that the industry median for known exploited vulnerabilities is currently 43 days, start with achievable baselines and tighten them as automation improves. Reserve the shortest windows for the small share of findings that are internet-facing, exploited in the wild, and high-impact.
Step 3: Decouple security metrics from punitive action. The goal is to find bottlenecks, not to punish people. If Squad Charlie has a high breach rate, the response is an investigation:
Is the squad carrying too much feature work?
Do they lack security training or tooling?
Is the architecture so tightly coupled that simple fixes become complex?
Treat the metric as a diagnostic, and keep it out of individual performance reviews.
Step 4: Celebrate consistent hygiene. Use all-hands meetings or engineering newsletters to highlight squads with sustained compliance, and share what they do differently. Public recognition of thoroughness over frantic speed sets the cultural tone.
Conclusion
As development accelerates, and as AI-assisted attackers shrink the window between disclosure and exploitation, measuring DevSecOps performance well matters more than ever. But grading developers on raw MTTR is a flawed strategy. It encourages Band-Aid fixes, metric manipulation, and burnout, and it leans on a DORA lineage that does not actually apply to vulnerability remediation.
The better path is metrics that reward consistency and reliability against risk-adjusted deadlines: fair clocks that start at discovery, groupings that reflect real engineering effort, exceptions that are governed rather than hidden, and dashboards used to support teams, not rank them. Track SLA Breach Rate by Squad with a tool like InstaSLA, add guardrails against Goodhart-style gaming, and security becomes a predictable part of high-quality engineering rather than a race against the clock.
Sources
CISA, BOD 26-04: Prioritizing Security Updates Based on Risk and implementation guidance (June 10, 2026)
Patrowl, BOD 26-04: Patch Less, but Better
DORA, A history of DORA's software delivery metrics and DORA's software delivery performance metrics
re:cinq, What are DORA metrics? The 5 metrics, benchmarks and AI
Tenable, Key findings from the Verizon DBIR 2026
Symmetry Systems, 8 unsurprising findings from the 2026 Verizon DBIR
Cyber Insurance News, Verizon 2026 DBIR: vulnerability exploitation is now the #1 breach vector
Veracode, 2026 State of Software Security report
FedRAMP, Public Notice on BOD 26-04 and Vulnerability Evaluation and Reporting rules
ReversingLabs, AppSec alert fatigue: 4 ways to reduce burnout (citing Cycode's survey)
zbmowrey, DORA metrics: what the research actually says (secondary source for the DORA 2023 "league tables" caution)Preventing "Alert Relocation": The Missing Step in Shift-Left Security
Top comments (0)