Part 5 of 8 of Building the Bank from the Top Down. The main paper makes the argument; this part holds the detail.
C3. Case in point: the Personal Finance Agent
Applying the method of Appendix C2 to one outcome, the Personal Finance Agent (PFA), market A's first product, makes the difference between core-led and outcome-led transformation concrete. Same bank, same ambition, same legacy estate: core-led, customers wait for the foundations; outcome-led, the first release can reach them in quarters rather than years. The principle: don't modernise the core as a prerequisite for AI; invest in it where resilience, regulation, economics or a named customer outcome requires it.
C3.1 The customer need, and why they would choose the bank
The customer's words: "Help me feel in control of my money, across every account I have, without having to manage it."
Take a salaried customer earning €3,420 a month net, building a €4,000 emergency buffer and saving for a home in 44 months. The question that matters to them is not "what is my balance?" It is: "If our household income drops 20%, what happens to the home?" Answering it well aims at the upper tiers of Bain's pyramid: reducing anxiety, providing hope, a sense of control.
Trust alone will not win this customer. A general-purpose assistant is free and already on their phone. The bank's agent has to offer things that assistant cannot easily match, and each one is a claim for the first release to test:
| Advantage | What the customer gets | How the first release tests it | Kill threshold at month 9 (illustrative) |
|---|---|---|---|
| Deeper financial context | Answers built on every consented account, goal and commitment, with pending, booked and reversed items read correctly | Answer accuracy against a general assistant on the same questions | Not ahead of a free general assistant on a fixed test set of customer questions |
| Trusted execution | An agent that can act, within limits, not only advise | Share of recommendations the customer completes | Fewer than 1 in 5 recommendations completed |
| Cross-provider orchestration | One view and one plan across banks, pensions and insurers | Providers linked per active user | Median active user links no provider beyond the bank |
| Transparent evidence | Every number traceable to the data and claims behind it | Customer rating of explanation quality | Rated no better than the general assistant's answers |
| Controllable autonomy | The customer sets what the agent may do alone, and can take it back | Share of users who raise or lower autonomy | Fewer than 1 in 10 active users ever change a setting |
| Liability protection | The bank compensates if its agent errs within the mandate | Trust score; dispute and reimbursement rates | Trust score not above the general assistant's; disputes above the bank's card-payment rate |
These are kill criteria, not reporting metrics. The clock starts with the first read-only release: review at month 6, decide at month 9. The deciding test is control and trust, the last three rows, because that is where a free general assistant is weakest. If the PFA misses those thresholds at month 9, that is the signal in Appendix C6 to shift weight and budget from market A to market B. The thresholds are illustrative; the board sets the real ones with the PFA investment envelope (Appendix C7.1).
C3.2 Core-led delivery
Few banks run this sequence in pure form; most modernise selected domains alongside digital journeys. Core-led here means the variant where customer value is gated by the sequence: stabilise the core, harmonise the data in a warehouse, build APIs, then put an agent on top.
- Value arrives last. Customers see nothing until the foundation steps are done; on typical plans that is two to three years (a working estimate).
- The view is narrow. The agent sees one bank's data. The customer's other banks, pensions and insurers are invisible.
- Risk is concentrated. Large core refactoring brings regressions, outages and compliance breaks.
- The window may close. By launch, a challenger may already hold the customer's conversation.
C3.3 Outcome-led delivery
The outcome-led programme starts from the customer outcome and works backwards. It builds three apex components first:
- Personal context graph. Models the customer's goals and constraints (income, dependants, home target, buffer) rather than account tables.
- Purpose-bound consent router. Turns each approval into a narrow, time-limited, revocable claim, such as "income above €3,000, for mortgage affordability only."
- Semantic control plane. AI reads legacy and open-finance data where it sits, through a controlled chain: raw data, semantic mapping, canonical financial concepts, evidence and provenance, verified claims, then agent reasoning. Fields where a misreading is costly (pending items, reversals, currencies, interest, product definitions) get canonical definitions first; the rest are interpreted on demand.
Around these sits an agent control architecture: identity for people and agents, delegated authority, policy, independent verification, transaction limits and idempotency, audit, monitoring, human escalation and recovery. Autonomy climbs a seven-rung ladder. Each step is gated by risk, reversibility, value and confidence, not by customer permission alone. The Technical annex sets out the full stack, a liability matrix and failure scenarios.
The ladder becomes policy once each class of action has a ceiling. The rungs are 1 observe, 2 explain, 3 recommend, 4 prepare, 5 ask approval, 6 execute within limits and 7 execute autonomously. The ceilings below are a starting position for the board's autonomy policy, not a legal view:
| Risk class | Example | Reversible? | Highest rung today |
|---|---|---|---|
| Information only | Spending summary, cash-flow forecast | No money moves | 2: explain |
| Guidance on regulated products | Savings, credit or investment suggestion | No money moves | 3: recommend, under advice rules |
| Transfers between the customer's own accounts | Sweep €100 to savings | Yes | 6: execute within limits |
| Payments to known payees | Recurring utility bill | Partly | 5: ask approval; 6 once PSR standards on delegated payments settle |
| New payees or one-off payments | First payment to a new merchant | No | 5: ask approval |
| Mandate changes | Change a standing order or direct debit | Partly | 4: prepare |
| Credit and investment decisions | Loan application, fund purchase | No | 4: prepare; the lender or adviser decides |
Rung 7 is out of scope for payments until the PSR technical standards settle (Technical annex, T3).
Full detail: Technical annex
The legacy core stays in place. Core investment continues where resilience, regulation, economics (cost, capacity, vendor end-of-life) or a named customer outcome requires it. It is no longer the gate for customer value.
Five architecture principles. Both products run on one platform, so these hold for market A and market B alike:
- A control plane, not a new warehouse. Map only the high-risk fields canonically; interpret the rest on demand, with provenance (Technical annex, T6 and T7).
- Consent as a product. The PSR permission dashboard is the customer's control surface for every agent, the bank's and outside ones, not a compliance screen. It is where the bank can keep the relationship even when the conversation happens elsewhere.
- One agent control plane. Identity for people and agents, policy, independent verifiers, limits and idempotency, append-only audit, escalation and recovery are shared services, used by the PFA and by the claim service alike.
- Conversation-first, app for confirmation. The agent lives where the customer talks: in the app, in messaging and in system assistants. High-stakes steps deep-link into the bank's app, which already holds strong authentication.
- Observable and stoppable from day one. Every claim, recommendation and action is traceable, reversible where possible, and can be switched off per action class. The board sees the resulting metrics (C7.2).

Core-led versus outcome-led delivery of the PFA · timings are working estimates
C3.4 Map 2: what to build, buy, rent and contain
Placing each PFA component on a Wardley map turns the architecture into an investment decision. Placements use the test in Appendix B1.2. The bank builds what is uncharted and customer-facing, buys what is maturing, rents what is commodity, and contains the core.

Wardley Map 2 · PFA components by layer and evolution · from the FOB-analysis workbook
| Component | Stage | Evidence for the placement | Action | Bain element served |
|---|---|---|---|---|
| Financial autonomy and confidence | Genesis | No mass-market product delivers it; no agreed measure of financial confidence | Own: this is the North Star | Provides hope, self-actualisation |
| Explained answers | Genesis | No standard for evidence-linked financial answers; AI Act explainability practice still forming | Build: every answer traceable to evidence, never a bare model output | Reduces anxiety, informs |
| Graduated autonomy | Genesis | Authentication for delegated payments is unsettled under the PSR (Appendix B2.2) | Build: clear guardrails for what the agent may do alone | Reduces effort, avoids hassles |
| Personal context graph | Custom | Budgeting apps and aggregators build their own; no product models household goals across providers | Build: model goals and constraints, not products | Integrates, organises |
| Consent router | Custom | PSR permission dashboards are required but not yet built; purpose-bound consent is not sold as a product | Build: a likely moat, because it sits where the PSR dashboard lands | Reduces risk, reduces anxiety |
| Revocable memory | Custom | Retrieval tooling is productised; memory the customer can edit and revoke is not | Build: the customer edits what the agent remembers | Organises |
| Maths verifiers, credential standards, saga orchestration | Product | Several vendors and open standards exist (W3C Verifiable Credentials, EU Digital Identity Wallet). The standard is a product; the bank's claim-issuing service (Appendix C1.5) is not | Buy or adopt standards | Quality, simplifies |
| LLM engine, open finance APIs, cloud | Commodity | Sold per token, per call or per hour by several interchangeable suppliers | Rent: do not train foundation models or build bespoke connectors | Connects |
| Core ledger | Commodity | Several vendors sell core banking; ledgering adds no difference the customer sees | Contain: invest where resilience, regulation, economics or a named outcome requires it | Reduces risk |
C3.5 Side by side
| Core-led | Outcome-led | |
|---|---|---|
| Starting point | The core system | The customer outcome |
| Data | Harmonise everything first, in a warehouse | Scoped data read through a semantic control plane |
| Scope | The bank's own accounts | Every provider the customer consents to |
| First customer value | Years (working estimate) | Quarters (hypothesis, see C3.6) |
| Where the money goes | Mostly on core and data foundations | Mostly on context, consent, control and experience |
| Delivery risk | High: core refactoring, regressions, outages | Contained: core largely untouched, new layer isolated |
| New risks to manage | Few new ones | Model errors, explainability, agent liability (Technical annex) |
C3.6 Where the time goes
Outcome-led delivery does not skip controls. It takes some dependencies off the critical path, shrinks others and runs the rest in parallel. The first table answers the question a sceptical CTO will ask: which dependencies go away?
| Dependency | Core-led | Outcome-led | What changes |
|---|---|---|---|
| Core replacement or refactoring | Early, before the agent | Only where a named outcome depends on it | Off the critical path |
| Full API modernisation | Before the agent | Only the APIs the first journeys use | Scoped |
| Canonical data model | Before the agent | Canonical concepts for costly fields only | Off the critical path, not eliminated |
| Data cleansing | Broad programme | Targeted semantic translation of the fields used | Scoped |
| Legacy data access | Multi-year harmonisation | Read-only connectors to two or three systems | Scoped |
| Consent | After the APIs | First capability, built as a product | Moves to the start |
| Risk controls | A later platform phase | Embedded in each capability, per release | Earlier and smaller |
The second table shows the activities that remain, and which of them are still long poles.
| Activity | Core-led | Outcome-led | What changes |
|---|---|---|---|
| Customer discovery | After the foundations | First, then continuous | The outcome sets the scope |
| Agent MVP | Last | Early and read-only (observe, explain) | Starts at the lowest-risk rung |
| Security assessment | Once, at the end | Per release, on a smaller surface | Scoped, not skipped |
| Model validation | Late | Early and continuous | Long pole: not faster, run in parallel |
| Legal and privacy review | Late | Early, on the consent design | Run in parallel |
| Operational readiness | Big-bang | Per release | Smaller increments |
| Production deployment | One large release | Quarterly releases | Smaller batches |
The remaining long poles are model validation, security and the semantics of the fields the agent uses. That is why the first release should be read-only: it proves value without touching payment authentication. Each bank should test these timings against its own portfolio.
C3.7 The economics to prove
The paper does not claim a business case; each bank must build its own. The model below lists the drivers and the evidence each needs.
The board's test is one number: net economic value per active agent customer, per year.
\text{Net value} = \underbrace{(R + P + W + S + K)}_{\text{value}} - \underbrace{(I + E + D + G + L + H)}_{\text{cost}}
Value is retention (R), product penetration (P), deposits and share of wallet (W), servicing cost avoided (S) and risk reduction (K). Cost is inference (I), engineering and integration (E), data (D), governance and compliance (G), liability and fraud (L) and human oversight (H).
| Side | Driver | Evidence the bank must supply |
|---|---|---|
| Value | Retention uplift among agent users | Churn of pilot users against a matched control group |
| Value | Product penetration: savings, lending, investment, insurance | Conversion from agent recommendations |
| Value | Deposits held and share of wallet | Balance movement across providers the customer links |
| Value | Lower servicing cost | Contact-centre and branch demand avoided |
| Cost | Inference and infrastructure | Cost per active user per month at pilot scale |
| Cost | Engineering, data and integration | Team cost of the first three releases |
| Cost | Model governance, validation and compliance | Validation effort per model change |
| Cost | Liability, fraud and human oversight | Error, dispute and escalation rates in the pilot |
An illustrative example. Every number below is an assumption chosen to show how the model works. None is an estimate for any bank.
| Side | Driver | Illustrative assumption | € per active user per year |
|---|---|---|---|
| Value | Retention (R) | Churn falls from 8% to 6% on €900 annual revenue per customer | 18 |
| Value | Product penetration (P) | 5% of users add one product earning €400 a year | 20 |
| Value | Deposits and share of wallet (W) | €1,000 more average balance at a 1.5% net margin | 15 |
| Value | Servicing cost avoided (S) | Two fewer contacts a year at €4 each | 8 |
| Value | Risk reduction (K) | Fewer arrears as buffers grow | 4 |
| Value | Total value | 65 | |
| Cost | Inference (I) | €0.50 per user per month | 6 |
| Cost | Engineering and integration (E) | Three releases, amortised over the user base | 12 |
| Cost | Data (D) | Semantic control plane and data feeds | 3 |
| Cost | Governance and compliance (G) | Model validation, AI Act and advice controls | 6 |
| Cost | Liability and fraud (L) | Errors, disputes and reimbursements | 3 |
| Cost | Human oversight (H) | One escalation a year at €5 | 5 |
| Cost | Total cost | 35 | |
| Net | Net value per active user | 30 |
On these inputs the case survives a halving of the retention and penetration effects (net value €11). It fails if that happens and engineering cost also doubles (net value −€1). So the pilot must measure retention and product take-up first; they carry most of the value.
The first release is designed to produce these numbers. The Scale gate in Appendix C7 should depend on them.
Previous: Appendix C1–C2: One platform, two products, and the method · Next: Appendix C4–C7: Product lines, guardrails and roadmap
Top comments (0)