DEV Community

Cover image for Engineering Agentic Systems for Financial Workflows: Harnessing LLM Non-Determinism While Guaranteeing Deterministic Execution
Aparna Pradhan
Aparna Pradhan

Posted on

Engineering Agentic Systems for Financial Workflows: Harnessing LLM Non-Determinism While Guaranteeing Deterministic Execution

The integration of Large Language Models (LLMs) into financial technology introduces a fundamental engineering challenge: financial systems require absolute mathematical determinism, whereas LLMs are intrinsically probabilistic and non-deterministic engines.

When financial engineering teams attempt to use LLMs as direct calculation engines, failure is inevitable. LLMs are notoriously ill-suited for arithmetic, ledger balancing, interest rate calculations, tax computations, FX conversions, and exact reconciliation rules. However, attempting to eliminate non-determinism entirely by restricting enterprise automation to hardcoded rule engines leaves institutions incapable of processing the unstructured, ambiguous, and multi-source realities of modern commerce.

The core architectural breakthrough lies in shifting the paradigm: the objective is not to make the LLM compute financial numbers, but to utilize LLM non-determinism as a cognitive coordinator that interprets unstructured ambiguity, forms hypotheses, and emits typed specifications for deterministic engines to execute.


1. The Core Selection Test: Where Non-Determinism Adds Value

To build a reliable agentic financial system, architects must apply a rigorous evaluation framework before assigning any task to an LLM.

                                  +-----------------------------------+
                                  |   Incoming Financial Case / Event |
                                  +-----------------------------------+
                                                    |
                                                    v
                                  +-----------------------------------+
                                  |     Is the input structured,      |
                                  |     calculation pre-specified,    |
                                  |   & path fully deterministic?     |
                                  +-----------------------------------+
                                         /                     \
                                    YES /                       \ NO
                                       v                         v
                       +-------------------------------+   +-------------------------------+
                       + Deterministic Code Execution  +   + LLM Cognitive Coordinator     +
                       | - Double-entry accounting     |   | - Context & intent parsing    |
                       | - Tax & interest formulas     |   | - Dynamic path selection      |
                       | - State transitions & rules   |   | - Hypothesis generation       |
                       | - Transaction authorization   |   | - Semantic normalization      |
                       +-------------------------------+   +-------------------------------+
                                                                         |
                                                                         v
                                                           +-------------------------------+
                                                           | Emits Typed Execution Plan    |
                                                           +-------------------------------+
                                                                         |
                                                                         v
                                                           +-------------------------------+
                                                           | Deterministic Tool Validation |
                                                           +-------------------------------+
Enter fullscreen mode Exit fullscreen mode

When Non-Determinism is Essential

A financial workflow is an optimal candidate for an LLM when it satisfies the following six criteria:

  1. Input Variability: The task originates from unstructured or semi-structured sources—such as natural language emails, PDF invoices, free-text remittance notes, customer support tickets, or inconsistent source schemas.
  2. Path Uncertainty: The required sequence of API calls or database queries cannot be pre-calculated before inspecting the case details.
  3. Semantic Ambiguity: Contextual evaluation is required to interpret ambiguous business terminology (e.g., distinguishing whether a billing variation is a "temporary migration overlap," a "contractual discount," or an "unauthorized discrepancy").
  4. Cross-Source Synthesis: Resolving the case requires correlating structured general ledger (GL) entries with unstructured documentation, such as procurement contracts, customer threads, or external vendor policies.
  5. Multiple Plausible Hypotheses: The system must generate and systematically test competing candidate explanations before evidence rules them out.
  6. Human-Tailored Explanation: The final output requires translating verified numerical results into contextual business narratives tailored to specific stakeholders (e.g., CFOs, collections managers, or end customers).

Tasks That Must Remain 100% Deterministic

LLMs must never operate as the primary execution engine or system of record for:

  • Double-entry bookkeeping and trial balance validation.
  • Arithmetic, interest, tax, and FX conversions.
  • Revenue and EBITDA formula calculations.
  • Database writes, payment authorizations, and spending limit enforcement.
  • State machine transitions, access control, and transaction permissions.

2. The 80/20 Layered Technical Architecture

To achieve sub-second execution safety and zero-hallucination guarantees, financial architectures employ an 80/20 hybrid layered model.

+-----------------------------------------------------------------------------------+
|                            LAYER 2: DYNAMIC REASONING (20%)                      |
|   - LLM Cognitive Coordinator (Amazon Bedrock / Foundation Models)                |
|   - Parses intent, chooses tools, formulates hypotheses, synthesizes evidence     |
+-----------------------------------------------------------------------------------+
                                         |
                                         v
+-----------------------------------------------------------------------------------+
|                        LAYER 1: PRECONFIGURED SKILLS (80%)                         |
|   - Hardcoded, deterministic, idempotent Python/Java code functions              |
|   - variance_analysis(), invoice_match(), aging_analysis(), claim_validate()      |
+-----------------------------------------------------------------------------------+
                                         |
                                         v
+-----------------------------------------------------------------------------------+
|                        LAYER 3: CUSTOM CALCULATION SANDBOX                        |
|   - Pydantic-validated Execution Specs (e.g., CalculationSpec JSON)               |
|   - Isolated execution environment for ad-hoc projections & scenario modeling     |
+-----------------------------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Layer 1: Preconfigured Financial Skills (80% of System Logic)

This layer consists of deterministic code functions with strictly typed Pydantic schemas, idempotent execution guarantees, unit test coverage, and audit logging. Key skills include:

  • variance_analysis()
  • price_volume_mix()
  • invoice_match()
  • payment_candidate_match()
  • aging_analysis()
  • claim_validate()

Layer 2: Dynamic Reasoning Layer (20% of System Logic)

The LLM sits in this layer as the orchestration engine. It evaluates incoming unstructured events, determines which Layer 1 skills to invoke, identifies missing context, evaluates returned evidence, and constructs candidate resolution proposals.

Layer 3: Custom Calculation Sandbox

For complex, ad-hoc scenario modeling that falls outside preconfigured skills, the LLM must not execute arbitrary code or shell scripts. Instead, it outputs a constrained, typed CalculationSpec:

{
  "template": "custom_projection",
  "inputs": ["monthly_revenue", "headcount_cost"],
  "horizon_months": 6,
  "transformations": [
    {"type": "growth", "rate": 0.03},
    {"type": "add_cost", "start_month": 3, "amount": 15000}
  ]
}
Enter fullscreen mode Exit fullscreen mode

This specification is validated against Pydantic schemas and policy constraints before execution in an isolated Python runtime, returning a fully deterministic, auditable execution trace.


3. High-Value Implementation Case Studies

Case Study A: Payment & Reconciliation Exception Resolution

A common operational bottleneck occurs in cash application when incoming payments cannot be matched directly to open receivables.

The Problem Context

  • Bank Stream Event: ACME TECH SERVICES 142,750 USD
  • Internal Records: Invoice A ($142,750), Invoice B ($71,375), Credit Note ($71,375).
  • Customer Communication: "We settled both open balances after applying the agreed adjustment."
  • Rule Engine Failure: Traditional string-matching or exact-amount lookups fail because the transaction descriptor does not reference a single invoice number.
+------------------+     +-----------------------+     +------------------------+
| Bank Transaction | --> | Unstructured Metadata | --> | LLM Agent Verification |
| $142,750         |     | Customer Email Notes  |     | Parses & Synthesizes   |
+------------------+     +-----------------------+     +------------------------+
                                                                   |
                                                                   v
+------------------+     +-----------------------+     +------------------------+
| Final Execution  | <-- | Human Approval Gate   | <-- | Proposed Resolution    |
| Ledger Settlement|     | (If risk > threshold) |     | Invoice A Allocation   |
+------------------+     +-----------------------+     +------------------------+
Enter fullscreen mode Exit fullscreen mode

Orchestrated Division of Labor

  • LLM Role: Interprets unstructured remittance text, parses email threads, correlates cross-source records, formulates allocation hypotheses, and drafts resolution proposals.
  • Deterministic Code Role: Queries ledger database, calculates residuals, executes exact-match mathematical verifications, enforces double-allocation prevention, and gates ledger postings behind policy thresholds or human approvals.

Case Study B: Serverless Real-Time KYC Architecture on AWS

Modernizing Know Your Customer (KYC) workflows requires transitioning from legacy, batch-oriented monolithic architectures (which average 3–5 days per onboarding case) to event-driven, real-time agentic pipelines.

                           AWS CLOUD ARCHITECTURE

 +-------------------+      +-------------------------------------------------+
 | Customer Requests | ---> | Amazon MSK (Managed Streaming for Apache Kafka) |
 +-------------------+      +-------------------------------------------------+
                                                     |
                                                     v
                            +-------------------------------------------------+
                            | AWS Lambda (Asynchronous Event Listeners)       |
                            +-------------------------------------------------+
                                                     |
                                                     v
 +-----------------------------------------------------------------------------------+
 |                    AMAZON BEDROCK AGENTCORE RUNTIME ENVIRONMENT                   |
 |                                                                                   |
 |                      +---------------------------------------+                    |
 |                      | KYC Orchestration Supervisor Agent    |                    |
 |                      +---------------------------------------+                    |
 |                                          |                                        |
 |        +-----------------+---------------+-----------------+----------------+     |
 |        |                 |               |                 |                |     |
 |        v                 v               v                 v                v     |
 |  +-----------+    +------------+   +-----------+    +------------+   +------------+ |
 |  | Identity  |    | Document   |   | Fraud     |    | Compliance |   | Customer   | |
 |  | Verification   | Analysis   |   | Detection |    | & Risk     |   | Experience | |
 |  | Sub-Agent |    | Sub-Agent  |   | Sub-Agent |    | Sub-Agent  |   | Sub-Agent  | |
 |  +-----------+    +------------+   +-----------+    +------------+   +------------+ |
 |        |                 |               |                 |                |     |
 |        +-----------------+---------------+-----------------+----------------+     |
 |                                          |                                        |
 |                                          v                                        |
 |                      +---------------------------------------+                    |
 |                      | AgentCore Shared Memory & Context Store|                   |
 |                      +---------------------------------------+                    |
 +-----------------------------------------------------------------------------------+
                                            |
                         +------------------+------------------+
                         |                                     |
                         v                                     v
 +-----------------------------------------------+ +---------------------------------+
 | Vector Knowledge Base                         | | Real-Time State Store           |
 | - Amazon OpenSearch Serverless (Embeddings)   | | - Amazon DynamoDB             |
 | - Amazon S3 (Regulatory Docs & Policies)      | | - Sub-millisecond Risk Lookups |
 +-----------------------------------------------+ +---------------------------------+
                         |                                     |
                         +------------------+------------------+
                                            |
                                            v
 +-----------------------------------------------------------------------------------+
 |               AGENTCORE GATEWAY & ON-PREMISES ENTERPRISE INTEGRATION              |
 | - AWS Direct Connect / VPN                                                        |
 | - OpenAPI Specs targeting Core Banking, Case Management, and Risk Systems         |
 +-----------------------------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Architectural Breakdown

  1. Event Streaming Backbone: Amazon Managed Streaming for Apache Kafka (Amazon MSK) ingests inbound customer requests, document uploads, and third-party verification events asynchronously.
  2. Orchestration Runtime: Amazon Bedrock AgentCore hosts the runtime environment, providing native session state persistence, context management, and multi-agent coordination.
  3. Supervisor & Domain-Specific Sub-Agents: The KYC Orchestration Supervisor Agent dynamically constructs execution plans across five specialized sub-agents:
    • Identity Verification Sub-Agent: Cross-references customer identities against global watchlists, PEP, and sanctions databases.
    • Document Analysis Sub-Agent: Performs OCR on identity documents, assesses image quality, translates multi-language documents, and detects document forgery.
    • Fraud Detection Sub-Agent: Conducts behavioral analysis, evaluates IP duplicate applications, and performs semantic similarity searches over historical fraud vectors.
    • Compliance & Risk Sub-Agent: Interprets jurisdiction-specific regulations (e.g., BSA, USA PATRIOT Act, EU AMLD, MAS guidelines) and generates verifiable audit attestations.
    • Customer Experience Sub-Agent: Identifies onboarding friction points and optimizes applicant communication.
  4. Retrieval-Augmented Generation (RAG): Amazon OpenSearch Serverless provides vector search over institutional policies stored in Amazon S3, while Amazon DynamoDB provides sub-millisecond status lookups.
  5. Dynamic Confidence Routing: The Supervisor Agent routes outcomes based on sub-agent confidence scores:
    • High Confidence (>95%): Automated real-time approval.
    • Medium Confidence (75%–95%): Triggers step-up verification workflows.
    • Low Confidence (<75%): Escalates to human compliance reviewers with an auto-generated decision packet.

Performance & Operational ROI

  • Validation Latency: Reduced from 3–5 days to under 5 minutes.
  • Compliance Efficiency: Automated routing enables compliance specialists to handle up to 4x their previous caseload by focusing exclusively on complex, escalated exceptions.

4. Strategic Portfolio Analysis of Financial Agentic Workflows

When prioritizing agentic implementations across financial operations, workflows are evaluated by business impact, money-handling risks, and architectural feasibility:

Workflow Rank & Title Core LLM Functionality Deterministic System Functionality Business Impact & Primary Metrics
Rank 1: Payment Reconciliation Exception Agent Unstructured remittance parsing, cross-source entity resolution, hypothesis generation. Residual arithmetic, candidate matching, double-allocation constraints, ledger write gates. High ROI: Dramatically increases automated cash application rate; cuts analyst resolution time per exception.
Rank 2: Accounts Receivable (AR) Case Handling Agent Communication analysis, dispute classification, automated contextual response drafting. Aging calculation, outstanding balance computation, task scheduling, policy enforcement. Direct Cash Impact: Decreases Days Sales Outstanding (DSO) and reduces manual collector overhead.
Rank 3: Variance-to-Decision FP&A Agent Translating business queries into investigation paths, hypothesis generation over EBITDA drivers. Bridge calculations, price-volume-mix formulas, period/currency conversions. Decision Velocity: Compresses monthly variance analysis cycles from days to minutes.
Rank 4: Cash-Forecast Exception Investigator Identifying operational drivers behind forecast deviations, qualitative impact synthesis. Cash runway modeling, scenario impact quantification, threshold monitoring. Risk Reduction: Accelerates early warning detection for liquidity risks.
Auxiliary: Management Intent to Analysis Plan Parsing vague prompts (e.g., "Can we afford 3 hires if sales slow?") into typed AnalysisRequest objects. Executing financial scenario projections, cash constraint validations. UX Transformation: Bridges non-technical business intent with complex modeling engines.
Auxiliary: Case Escalation Packet Generator Synthesizing investigated evidence, failed hypotheses, and citations into compact decision summaries. Evidence package bundling, permission checks, audit logging. Handoff Efficiency: Eliminates context reconstruction time during human escalations.

5. Architectural Principles for Engineering Agentic Finance

  1. Enforce Rigid Separation Between Reasoning and Calculation: Never permit an LLM to directly calculate monetary figures or perform ledger state transitions. The LLM generates structured investigation plans; code executes the operations and validates the math.
  2. Implement Claim Schema Verification: Before presenting an LLM-synthesized narrative to a user or reviewer, run a verifier component that validates every numerical claim against underlying source database records.
  3. Design for Fallback and Human Escalation: Establish explicit confidence thresholds. High-risk or low-confidence actions must emit standardized decision packets for human approval before execution.
  4. Scope Tasks Around Complete Units of Work: Build agents around bounded operational problems that feature a structured trigger, an unstructured middle, deterministic tool availability, a safe action boundary, and clear evaluation metrics.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.