Execution Integrity | Engineering AI for Partial Execution, Recovery and Control | R.A.H.S.I. Framework™
An agent updates a customer record, opens a service ticket and requests a downstream adjustment.
The first two actions commit.
The last request times out.
Did the adjustment fail, commit without confirmation or remain pending?
The agent’s answer cannot establish the transaction’s state.
A demo proves the path that finished. Architecture must also control the path that did not.
Microsoft recommends analyzing each step of a critical workload flow, its dependencies, failure modes and blast radius.
For an enterprise agent, the analysis spans:
- unsuitable model output
- stale grounding
- changed permissions
- tool failures
- incorrect step ordering
- lost workflow state
The timeout makes an automatic retry hazardous.
Azure’s Retry pattern warns that repeating a non-idempotent operation can cause duplicate side effects.
A lost confirmation is not proof of a failed write.
The R.A.H.S.I. Execution-Integrity Progression
1. Detect and Contain
Identify uncertain steps.
Gate further actions and isolate failing dependencies.
The system should distinguish between:
- confirmed success
- confirmed failure
- pending execution
- ambiguous execution
- execution requiring reconciliation
Continuing a workflow while the state of a consequential operation is unknown can extend the blast radius beyond the original failure.
2. Preserve State
Retain:
- completed steps
- pending work
- execution identifiers
- timestamps
- tool responses
- outcome evidence
- authorization context
- correlation information
A workflow checkpoint enables resumption.
It cannot prove that an external system committed a transaction.
That distinction matters.
Persisting orchestration state answers:
Where was the workflow?
It does not necessarily answer:
What happened inside the external system?
3. Recover
Reconcile before retrying.
Retries should be bounded and applied only where the failure mode is understood.
For transient and safely repeatable operations, retry may be appropriate.
For operations capable of producing irreversible or duplicate side effects, the system may first need to determine whether the previous request actually committed.
Where business rules permit, recovery can include:
- retry
- reconciliation
- compensation
- alternate execution paths
- human approval
- escalation
- controlled termination
Retry is not recovery when transaction state is unknown.
4. Reconstruct and Reassure
Correlate decisions and effects across services.
Reconstruct:
- what the agent decided
- what actions were requested
- what systems received those requests
- which operations completed
- which outcomes remain uncertain
- which compensating actions occurred
- who authorized subsequent execution
Before consequential work resumes, verify both state and authorization.
Recovery without state verification risks duplication.
Recovery without authorization verification risks executing a technically valid action that is no longer permitted.
Compensation Is Not Always Rollback
Saga and Compensating Transaction patterns can provide structures for coordinating recovery across distributed operations.
But compensation does not necessarily mean reversing a database transaction.
A completed business operation may no longer be technically reversible.
Compensation might instead require another explicit business action.
For example:
Original Action
↓
Customer adjustment submitted
↓
Execution outcome becomes uncertain
↓
Reconciliation
↓
Adjustment confirmed
↓
Business compensation required
↓
Create corrective adjustment
↓
Approval + audit evidence
The compensating action itself becomes another governed operation.
It requires its own:
- authorization
- execution evidence
- failure handling
- monitoring
- audit trail
Monitoring Is Necessary — But It Is Not Recovery
Microsoft Foundry evaluation and monitoring can help surface:
- unexpected model behavior
- quality degradation
- execution failures
- latency
- abnormal patterns
- operational problems
But a dashboard cannot reconcile an ambiguous commit.
It cannot determine whether repeating a non-idempotent transaction is safe.
And it cannot independently authorize another consequential write.
Observability tells us something happened.
Execution integrity determines what we are allowed to do about it.
Partial Success Is Still a System State
One of the most dangerous assumptions in agent architecture is that a workflow has only two outcomes:
SUCCESS
or
FAILURE
Enterprise execution is rarely that simple.
A more realistic state model may resemble:
NOT_STARTED
↓
IN_PROGRESS
↓
PARTIALLY_COMMITTED
↓
OUTCOME_UNCERTAIN
↓
RECONCILING
↓
RECOVERED
↓
COMPENSATED
↓
VERIFIED
Partial success can leave inconsistent records while appearing close to completion.
That is precisely why the failure path must be designed as deliberately as the happy path.
Execution Integrity
Execution integrity means establishing what happened, what remains uncertain and who can decide what happens next.
That requires more than prompt quality.
It requires architecture capable of controlling:
State → Evidence → Reconciliation → Recovery → Authorization → Resumption
The enterprise agent should not simply ask:
Can I continue?
The architecture should be able to establish:
Do we know enough about the previous execution state to continue safely?
That is a very different engineering standard.
R.A.H.S.I. Framework™
🛡️ Need implementation, not just insights?
Design reconciliation, compensation and recovery gates into your agent workflow.
🛡️ Article Link | Read the full article
🛡️ Let’s Connect | Work with Aakash Rahsi

aakashrahsi.online
Top comments (0)