DEV Community

Logan Foster
Logan Foster

Posted on

Most AI Governance Programs Track Activity. Here's How to Track Whether They're Working.


The most common failure mode in mature AI governance programs isn't bad policy — it's policy that exists but doesn't do anything.

The programs that fall into this trap share a characteristic: they measure inputs (how many policies were written, how many trainings were completed, how many systems were reviewed) rather than outcomes (whether the governance controls are actually reducing risk).

Activity metrics are not effectiveness metrics. Here's the difference, and how to build a measurement approach that actually tells you whether your program is working.


Why Activity Metrics Are Seductive and Wrong

Activity metrics feel rigorous. "We completed AI impact assessments on 100% of high-risk systems this year" is a concrete, verifiable statement. It sounds like evidence of a functioning program.

The problem is what it doesn't tell you: whether the assessments found anything, whether the findings were acted on, whether the systems are actually safer as a result of the assessments existing.

A governance program that produces completed assessments nobody reads is generating documents, not reducing risk. A training program with 98% completion rates but no behavior change is producing certificates, not capability.

The organizations with the most mature AI governance programs have moved from activity measurement to effectiveness measurement. Here's what that looks like.


The Six Metrics That Actually Matter

1. Risk finding rate and resolution rate

Track: of the AI impact assessments completed in the period, what percentage identified at least one material risk? Of the risks identified, what percentage have been resolved, mitigated, or accepted with documented rationale?

If your impact assessments are consistently finding zero issues, one of two things is true: your AI systems are genuinely low-risk, or your assessment process isn't calibrated to find real risks. The former is possible; the latter is more common.

A healthy finding rate signals that assessments are rigorous enough to surface real issues. A healthy resolution rate signals that the governance process has teeth — findings lead to action, not just documentation.

2. Assessment currency rate

Track: of AI systems currently in production, what percentage have a current impact assessment (completed within the required review cycle for their risk tier)?

This metric degrades over time as new systems are deployed, existing systems change, and review cycles pass without reassessment. A program that starts at 100% currency and drops to 60% over twelve months has a maintenance problem — either the intake process is missing new systems, or the reassessment cadence isn't being followed.

3. Human oversight failure rate

For AI systems with required human oversight steps — review before action, approval before deployment, escalation for edge cases — track how often those steps are being skipped or bypassed.

This requires logging the oversight step, not just documenting that it should exist. Organizations that don't have technical controls enforcing oversight requirements (and rely purely on process) often find significant bypass rates when they actually measure.

4. Third-party compliance rate

Track: of AI vendor contracts subject to your third-party AI requirements (incident notification, no-training-on-customer-data, audit access rights, model update notification), what percentage include the required provisions?

A governance policy that says "all AI vendor contracts must include X" is meaningful only if X is actually in the contracts. The gap between policy and contracts is one of the most common findings in AI governance audits.

5. AI incident rate and time to detection

Track AI-related incidents: outputs that caused harm, near-misses, bias complaints, security incidents involving AI systems. Both the rate (how many incidents per period, normalized by system count or transaction volume) and the detection timeline (how long between incident occurrence and detection) are meaningful.

Incident rate alone is ambiguous — a high incident rate might mean AI systems are behaving badly or might mean detection capability is good. Time to detection disambiguates: a program that detects incidents quickly is more effective than one that detects them slowly, regardless of volume.

6. Shadow AI policy bypass rate

Track: how many unauthorized AI tools are identified per period through network monitoring, help desk tickets, or employee surveys? What's the trend?

A declining trend suggests the official AI program is meeting employee needs and the bypass incentive is decreasing. A flat or increasing trend is a signal that the official program isn't keeping up with demand.


Building the Measurement Infrastructure

These metrics require instrumentation. Most of them aren't available from existing systems without deliberate design.

Assessment tracking needs to be centralized — a system of record for AI impact assessments, not a collection of individual documents. The system needs to record not just that an assessment was completed, but what it found and what was done about it.

Human oversight steps need technical controls, not just process descriptions. If the oversight step isn't logged by a system, you can't measure whether it's happening.

Vendor contract data needs to be extractable. This usually means either a contract management system with AI provision tagging or a periodic manual audit of vendor contracts against requirements.

Incident data needs a consistent intake process — a defined channel for reporting AI-related incidents, with categorization that allows analysis.

None of this is technically complex. All of it requires intentional design. The organizations that have it typically built it during the governance program implementation, not as an afterthought.


Reporting It Upward

The metric set above is designed to produce a one-page governance scorecard that's meaningful to a board or executive committee. Not the detail — the summary:

  • Coverage: are all high-risk systems being assessed and current?
  • Quality: are assessments finding real issues and resolving them?
  • Controls: are oversight mechanisms actually operating?
  • Contracts: are vendor requirements actually in contracts?
  • Incidents: is the program detecting and responding to issues?

A governance program that can answer these five questions with current data is demonstrably effective. One that can't is a policy exercise, not an operational program.


Further reading:

Top comments (0)