In 1953, the Eisenhower administration commissioned a study of the structural integrity of the U.S. highway bridge system. The study found significant variation in design standards, inspection practices, and maintenance quality across states. By 1968, the Federal Aid Highway Act had established national standards for bridge design and inspection — standards that applied regardless of which state a bridge was in, who built it, or who funded its maintenance.
The logic was straightforward: bridges are part of a national transportation network. The failure of any bridge in that network has consequences that extend beyond the jurisdiction that owns it. A national network of critical infrastructure requires national standards for the components that constitute it.
The software systems that now constitute the U.S. digital infrastructure — the banking networks, the healthcare platforms, the election management systems, the energy grid control software — have the same property. They are components of a national network. Their failure has consequences that extend beyond the organisations that own them. And they are governed, where they are governed at all, by a patchwork of sector-specific regulatory frameworks that were written before the systems they regulate had become critical infrastructure.
The DORA Four Key Metrics are not a magic solution to this governance gap. They are four numbers, derived from a decade of empirical research, that measure the most predictive indicators of software delivery performance. They are sector-neutral, implementation-independent, and available to any organisation with access to its own pipeline and incident data. They are, in short, the most plausible candidate for a national software quality benchmark that currently exists. This post makes the case for treating them as one.
What a National Software Quality Benchmark Would Address
The current regulatory landscape for software quality in critical infrastructure is characterised by three structural gaps:
Gap 1 — Measurement Without Benchmark. Most regulatory frameworks for operational resilience — OCC SR 21-3 for financial services, Joint Commission standards for healthcare, NERC CIP for energy — require organisations to measure operational performance and demonstrate continuous improvement. None specifies what "good" looks like in quantitative terms. An organisation that deploys software once per year with a 60% change failure rate may technically satisfy "continuous improvement" requirements if it deploys slightly more frequently than it did the previous year.
Gap 2 — Sector Isolation. The software delivery capabilities of a financial institution, a hospital system, and an energy utility are assessed by entirely different regulatory frameworks with no common measurement vocabulary. A bank that is a High DORA performer and a hospital that is a Low DORA performer cannot be compared — not because comparison is inappropriate, but because the regulatory frameworks that govern them do not share a unit of measurement.
Gap 3 — Outcome Without Process. Regulatory frameworks for operational resilience measure outcomes: recovery time objectives, recovery point objectives, incident frequency. They do not measure the process quality that predicts whether those outcomes will be achieved before the next incident. The DORA Four Key Metrics are predictive of resilience outcomes — organisations in the Elite performer cohort have measurably better incident outcomes than Low performers. A regulatory framework that only measures outcomes misses the process quality signal that predicts future performance.
The Empirical Foundation
The case for DORA metrics as a national benchmark rests on three properties that distinguish them from other candidate metrics.
Property 1 — Longitudinal Empirical Validation. The DORA research programme has collected performance data from tens of thousands of organisations across more than a decade. The relationship between DORA performance cohort and business outcomes (including resilience outcomes) is not a theoretical claim — it is a statistically robust empirical finding. Nicole Forsgren's research methodology (described in Accelerate) uses structural equation modelling to establish causal relationships, not just correlations.
Property 2 — Sector Neutrality. The DORA Four are defined at the delivery process level, not at the application domain level. Deployment Frequency is the same metric for a payment processing API and a clinical decision support system. This sector neutrality is essential for a benchmark that would apply across the eleven critical infrastructure sectors that are now dependent on software systems.
Property 3 — Self-Measurability. Unlike some quality metrics that require external audit to measure reliably, the DORA Four can be measured by any organisation with access to its own CI/CD pipeline data and incident management records. This reduces the compliance overhead of measurement and makes the benchmark accessible to organisations that lack the resources for dedicated quality measurement programmes.
────────────────────────────────────────────────────────────────────────────
DORA PERFORMER COHORT DEFINITIONS (2023 State of DevOps Report)
DEPLOYMENT FREQUENCY:
Elite: Multiple times per day
High: Once per week to once per month
Medium: Once per month to once every 6 months
Low: Fewer than once every 6 months
LEAD TIME FOR CHANGES:
Elite: Less than 1 hour
High: Between 1 day and 1 week
Medium: Between 1 month and 6 months
Low: More than 6 months
CHANGE FAILURE RATE:
Elite: 0–15%
High: 16–30%
Medium: 16–30% (same range as High in recent research)
Low: 46–60%
MEAN TIME TO RESTORE:
Elite: Less than 1 hour
High: Less than 1 day
Medium: Between 1 day and 1 week
Low: Between 1 week and 1 month
────────────────────────────────────────────────────────────────────────────
RESILIENCE OUTCOME CORRELATION (Accelerate research findings):
Elite + High performers vs. Low performers:
→ 2× more likely to exceed profitability goals
→ 2× more likely to exceed productivity goals
→ 24× less time lost to unplanned work
→ 3× lower change failure rate
→ 24× faster recovery from incidents
────────────────────────────────────────────────────────────────────────────
What a Policy Recommendation Would Look Like
A national software quality benchmark based on DORA methodology would have three components: a measurement requirement, a benchmark target, and an improvement mandate.
────────────────────────────────────────────────────────────────────────────
PROPOSED NATIONAL SOFTWARE QUALITY BENCHMARK FRAMEWORK
COMPONENT 1: MEASUREMENT REQUIREMENT
Regulatory applicability: Operators of critical infrastructure
systems (CISA-designated sectors) with > $1B in annual revenue
or > 1M users dependent on the system.
Requirement: Annual self-assessment of DORA Four Key Metrics
for all software systems supporting critical business services.
Assessment methodology: DORA Quick Check or equivalent pipeline-
derived measurement (not self-report by programme managers).
Reporting: Annual submission to sector regulator;
methodology documented and available for examination.
COMPONENT 2: BENCHMARK TARGETS (by system criticality tier)
Tier 0 (life-critical systems — healthcare, nuclear):
→ Deployment Frequency: ≥ monthly (at minimum; more is better)
→ Lead Time: ≤ 1 week
→ Change Failure Rate: ≤ 15%
→ MTTR: ≤ 1 day
Tier 1 (critical infrastructure systems — financial, energy, election):
→ Deployment Frequency: ≥ weekly
→ Lead Time: ≤ 1 day
→ Change Failure Rate: ≤ 30%
→ MTTR: ≤ 1 day
Note: Targets do not require Elite performance — they establish
a minimum standard below which regulatory scrutiny increases.
The targets correspond approximately to the High performer cohort,
which represents achievable performance for well-managed enterprises.
COMPONENT 3: IMPROVEMENT MANDATE
Organisations in the Low performer cohort must:
→ Submit a remediation plan within 90 days of measurement
→ Demonstrate measurable improvement within 12 months
→ Re-measurement annually; sustained Low performance triggers
enhanced regulatory oversight (not fines — examination priority)
────────────────────────────────────────────────────────────────────────────
How DORA Maps to Existing Regulatory Frameworks
The proposed benchmark does not require new regulatory authority. It requires that existing operational resilience frameworks — which already have quantitative measurement requirements — adopt DORA metrics as the measurement vocabulary for software delivery quality.
────────────────────────────────────────────────────────────────────────────
DORA → EXISTING REGULATORY FRAMEWORK ALIGNMENT
OCC SR 21-3 (Financial Services Operational Resilience):
Requirement: Continuous monitoring of operational performance
DORA mapping: Deployment Frequency and MTTR are direct measures
of the delivery and recovery performance SR 21-3
requires organisations to demonstrate.
Gap filled: SR 21-3 requires measurement; it doesn't specify
what to measure. DORA provides the measurement vocabulary.
NERC CIP (Energy Sector Reliability Standards):
Requirement: Configuration change management (CIP-010)
DORA mapping: Change Failure Rate is the quality measure for
the changes that CIP-010 requires to be managed.
A utility with a 50% CFR on bulk power system
changes is clearly not managing changes effectively,
regardless of CIP-010 process compliance.
Gap filled: CIP-010 measures process; CFR measures outcome quality.
HIPAA Security Rule (Healthcare):
Requirement: Contingency planning; emergency mode operations
DORA mapping: MTTR is the operational metric for contingency plan
effectiveness. Lead Time for Changes measures the
organisation's ability to deploy security patches
in response to identified vulnerabilities.
Gap filled: HIPAA measures plan existence; DORA measures plan quality.
Joint Commission (Healthcare IT Standards):
Requirement: Performance improvement processes for clinical systems
DORA mapping: Deployment Frequency + Lead Time measure the
organisation's ability to deliver improvements.
CFR measures whether those improvements degrade
existing performance.
Gap filled: Joint Commission standards measure process; DORA measures
the outcome quality of the delivery process itself.
CISA Election Security Guidance:
Requirement: Incident response planning; resilience testing
DORA mapping: MTTR is the primary operational metric for election
infrastructure resilience. CFR measures whether
changes to election systems are being managed safely.
Gap filled: CISA guidance is process-focused; DORA adds outcome
measurement to the resilience assessment.
────────────────────────────────────────────────────────────────────────────
The International Comparison
The United States is not alone in recognising the gap between operational resilience regulation and software delivery quality measurement. Two international examples demonstrate that the bridge from resilience requirements to delivery benchmarks is navigable.
UK Government Digital Service (GDS) Performance Framework. The UK GDS established performance metrics for government digital services that include deployment frequency and service availability. GDS-assessed services are publicly rated on a dashboard that includes these delivery quality indicators. The framework demonstrates that government digital services can be assessed against quantitative delivery benchmarks in a way that is transparent and actionable.
EU Digital Operational Resilience Act (DORA — the regulation). The EU regulation of the same name (different acronym: Digital Operational Resilience Act) requires financial institutions to test and demonstrate ICT resilience. Article 24 of DORA mandates threat-led penetration testing. Article 25 requires ICT risk management frameworks with quantitative performance indicators. The EU DORA does not use the DORA research metrics by name — but its substantive requirements for quantitative ICT performance measurement are consistent with what the DORA Four would measure.
A Phased Policy Adoption Pathway
────────────────────────────────────────────────────────────────────────────
PHASED POLICY ADOPTION: FROM VOLUNTARY TO MANDATORY
PHASE 1 (Years 1–2): Voluntary Adoption with Recognition
Action: NIST publishes DORA-based software quality guidance
as a voluntary framework supplement to the Cybersecurity
Framework (CSF) and AI RMF.
Incentive: Organisations demonstrating High/Elite DORA performance
receive streamlined examination treatment under
sector-specific regulatory frameworks.
Outcome: Builds the measurement infrastructure and industry
vocabulary before mandatory measurement begins.
PHASE 2 (Years 2–4): Mandatory Measurement, Voluntary Benchmarking
Action: OCC, CMS, NERC, and EAC require DORA measurement for
Tier 1 and Tier 0 systems under their oversight.
Benchmark targets are published but not yet enforced.
Incentive: Measurement data builds the national baseline that
makes benchmark targets evidence-based rather than
arbitrarily imposed.
Outcome: National visibility into software delivery quality
distribution across critical infrastructure sectors.
PHASE 3 (Years 4–6): Mandatory Measurement and Minimum Benchmarks
Action: Low performer status (below proposed minimum targets)
triggers enhanced oversight and remediation requirements.
Targets calibrated from Phase 2 measurement data.
Precedent: This approach mirrors how NERC CIP standards evolved —
voluntary guidelines (1998–2006) → mandatory standards
(2008–present) as the industry built measurement
infrastructure.
────────────────────────────────────────────────────────────────────────────
Objections and Responses
Objection: "DORA metrics were designed for technology companies, not regulated enterprises."
Response: The DORA research programme has surveyed organisations across all industries since 2014. The finding that regulated enterprises cluster in Low and Medium performer cohorts is precisely the evidence that motivates the policy recommendation — not a reason to exempt them from measurement. Regulated enterprises have the most to gain from improvement and the most to lose from continued Low performance.
Objection: "Regulated enterprises can't achieve Elite DORA performance because of compliance overhead."
Response: The proposed benchmark targets High performer cohort levels, not Elite. High performance (weekly deployments, < 1-week lead time, < 30% CFR, < 1-day MTTR) is achievable in regulated environments with the SRE practices documented throughout this series. The Beyond DORA post specifically addresses how regulated enterprise constraints (CAB cycles, compliance change freezes) are accommodated in the RE-adjusted DORA framework.
Objection: "Self-measurement is unreliable — organisations will game the metrics."
Response: Pipeline-derived DORA metrics are difficult to game because they are derived from objective system logs (deployment timestamps, incident creation/resolution timestamps) rather than from qualitative self-assessment. The measurement methodology is specified — not self-report — and examination teams can validate it against actual pipeline data.
Common Antipatterns in the Policy Conversation
The Compliance-as-Quality antipattern → Arguing that organisations already in compliance with sector-specific regulations do not need additional quality benchmarks. HIPAA compliance does not measure deployment frequency. PCI-DSS compliance does not measure MTTR. CIP-010 compliance does not measure change failure rate. Compliance measures process adherence; DORA measures delivery outcome quality. These are orthogonal.
The Sector-Exceptionalism antipattern → Arguing that a specific sector is too unique for DORA metrics to apply. Financial services organisations make this argument (unique regulatory constraints); healthcare organisations make it (unique patient safety obligations); energy operators make it (unique OT/IT complexity). In each case, the DORA metrics apply at the software delivery layer — which all three sectors have. The sector-specific constraints affect the target values, not the applicability of the measurement framework.
The Small-Organisation Exemption antipattern → Proposing mandatory benchmarking without proportionality provisions. The proposed framework includes a revenue/user threshold that exempts small operators. A community bank with $50M in assets does not have the same systemic risk profile as a globally systemically important financial institution. The benchmark should apply proportionally to the organisations whose failure has national consequences.
Five Action Items for This Week
Measure your organisation's DORA performance cohort using the Quick Check tool at dora.dev. Even if the measurement is not yet being used for regulatory purposes, knowing where your organisation sits in the distribution is the prerequisite for any improvement initiative. The Quick Check takes fifteen minutes and provides the self-assessment baseline that a formal measurement programme would produce.
Map your organisation's existing regulatory reporting obligations to the DORA Four. For each report you submit to sector regulators: which of the DORA Four does it most closely measure? Where are the gaps? The mapping exercise produces the argument for why DORA measurement would reduce regulatory reporting overhead (by consolidating multiple sector-specific metrics into a common framework) while improving measurement quality.
Submit a comment to NIST during the next CSF or AI RMF public comment period recommending the inclusion of software delivery quality metrics. The NIST public comment process is the most accessible pathway for practitioners to influence federal guidance. A well-reasoned comment from a practitioner with operational experience in regulated environments has influence that policy generalists cannot replicate.
Present your organisation's DORA metrics to your compliance or risk function with the regulatory alignment table from this post. The argument that DORA metrics already measure what SR 21-3, HIPAA, NERC CIP, and Joint Commission require — in a more quantitative form than the requirements themselves specify — is the compliance-function adoption argument. If compliance adopts DORA measurement, it becomes organisationally durable.
Contribute your organisation's performance data to the annual DORA State of DevOps survey. The survey is anonymous; the aggregate data is what drives the research. Every additional regulated enterprise respondent improves the statistical validity of the regulated enterprise cohort analysis — and that analysis is what makes the policy argument empirically grounded rather than theoretical.
"The United States has federal standards for the structural integrity of every bridge on the interstate highway system. It does not have federal standards for the delivery quality of the software systems that now carry the economy across a different kind of infrastructure. The DORA metrics are not perfect — no four numbers can fully characterise the delivery capability of a complex organisation. But they are the most empirically validated, sector-neutral, self-measurable quality benchmark the software delivery field currently has. The question is not whether they are perfect. The question is whether the absence of any benchmark at all is a better policy position."
Top comments (0)