In traditional software engineering, pre-release reviews focus on performance, unit test coverage, and security vulnerability scans. If the code passes its tests and the database handles load, the service is ready for release.
With generative AI, testing code syntax is not enough. You must evaluate societal impact, potential harms, fairness, and system autonomy.
Under ISO/IEC 42001 (Clause 6.1.4 and Clause 8.4) and the companion guidance in ISO/IEC 42005, organizations must conduct AI System Impact Assessments (AIIA).
An impact assessment is not a generic risk review. It is a structured evaluation that determines whether an AI system could cause harm, how autonomous it is permitted to be, and what human oversight must be in place.
Here is how we operationalized AI impact assessments, how the process worked in production, and what to watch out for.
The Idea: A Two-Stage Assessment Lifecycle
Requiring a 50-page formal risk review for every simple internal chatbot paralyzes product velocity. Conversely, letting high-impact agents deploy without oversight creates enterprise risk.
We structured our impact assessment process as a two-stage lifecycle:
[ New Agent / AI Feature Proposed ]
|
v
[ Stage 1: Fast Screening ]
Quick check on intended use, data sensitivity, and autonomy.
|
+--------+--------+
| |
Low Risk Elevated Risk
| |
v v
[ Standard Review ] [ Stage 2: Full Impact Assessment ]
Auto-approved with - Autonomy & Human-in-the-Loop gates
baseline guardrails - Failure mode & severity analysis
- Bias, fairness & environmental impact
- Multi-party sign-off (Legal, Security, Ops)
1. Stage 1: Quick Screening
A rapid check answers four baseline questions:
- Does the system make autonomous financial, legal, or safety decisions?
- Does it process special category data (health data, personal biometrics, financial records)?
- Does it interact with external users or children?
- Does it execute write operations against external enterprise systems?
If the answer to all four is no, the agent is categorized as low risk and can proceed through standard development with baseline platform guardrails.
2. Stage 2: Comprehensive Impact Assessment
If any answer is yes, the system enters a full assessment covering four specific dimensions:
- Autonomy and Human Oversight: Is the agent purely advisory (recommending suggestions to an engineer), human-in-the-loop (requiring an explicit button click before executing an action), or fully autonomous?
- Failure Modes & Severity: What happens if the model hallucinates? Does it produce an awkward phrase, or does it trigger an unauthorized financial transfer?
- Transparency & Disclosures: How will end users know they are interacting with an AI system? What disclaimers and system boundaries are displayed?
- Fallback & Kill Switches: If the agent exhibits unexpected behavior in production, how do operators immediately disable its tools or take it offline without breaking the host application?
How It Worked Well
- Clear Boundaries on Autonomous Tools: Evaluating autonomy during the assessment prevented accidental over-permissioning. High-impact tools (such as deleting database tables or sending external emails) were required to have explicit human-in-the-loop approval cards built into the user interface.
- Standardized Severity vs. Likelihood Matrix: Engineering and product teams scored risks using objective severity and likelihood definitions. This eliminated guesswork and allowed teams to focus on concrete mitigations (such as prompt guardrails, output validation, or tighter tool allowlists).
- Audit Readiness: Each completed assessment is stored as an immutable versioned record. When compliance auditors reviewed our platform, we demonstrated a documented, repeatable trail of impact assessments across all active agents.
- Actionable Kill Switches: Because the assessment explicitly requires a tested kill switch, every production agent has the capability to be decoupled from its external tools or placed into read-only mode with a single operational flag.
What to Watch Out For
- Re-Assessment Triggers: An impact assessment is not a one-time event. If an agent changes its system prompt, swaps its underlying foundation model, or connects to a new database tool, its risk profile changes. Establish automated re-assessment triggers whenever core configurations or tools are updated.
- Vague Intended Purpose Statements: If an application team defines their agent's purpose as "Help users be more productive," it is impossible to evaluate risk accurately. Require teams to specify precise operational domains, intended user personas, and explicit out-of-scope tasks.
- Ignoring Environmental and Compute Impact: ISO 42001 encourages evaluating the resource consumption of AI systems. A reasoning model that consumes 20,000 tokens per query for a simple lookup is not just financially wasteful; it carries an unnecessary carbon and compute footprint. Assessments should verify that model selection matches task complexity.
- Keeping Assessments Visible: Do not lock completed assessments inside siloed legal drives. Keep impact ratings and operational boundaries visible in the internal developer portal so on-call engineers know the exact constraints of each production agent.
Top comments (0)