Anthropic open sourced its reference agents for investment banking, equity research and fund administration. Ten of them, Apache-2.0, sitting at 36,000 stars.
The agents are interesting. The thing I would actually steal is five lines near the top of the README.
Capability is the easy half
Anyone can write an agent prompt that says what to do. "You are an expert financial analyst. Build a DCF model." That takes a minute, and it is the part every agent repo shows you.
Here is what these prompts also say, verbatim:
They do not make investment recommendations, execute transactions, bind risk, post to a ledger, or approve onboarding; every output is staged for human sign-off.
Five refusals. Read them again and notice they are not generic safety hedging. Each one names a specific action in a specific workflow, and each one is the exact point where drafting becomes deciding.
The KYC agent parses onboarding documents and flags gaps. It does not approve onboarding. The reconciliation agent traces a break to its root cause. It does not post to the ledger. The model builder builds the DCF. It does not tell you the company is undervalued.
Each of those lines is a conversation somebody had
That list did not come from a prompt engineer being careful. It came from compliance officers at regulated firms drawing a line, and somebody writing the line down where the model can see it.
That is why it is worth reading even if you will never open a comps table. Most of us write agent prompts by accumulating capabilities until the demo works. Almost nobody writes the boundary first, because the boundary is invisible until the agent crosses it, and by then you are explaining an incident rather than a design.
The pattern generalises cleanly. A support agent that drafts refunds but does not issue them. A deploy agent that prepares the release but does not promote to production. An HR agent that screens applications but does not reject candidates. In every case the valuable sentence is the second half, and in every case it is the half that gets left out.
Try it on something you already built
Take an agent prompt you have shipped. Write down the five actions it must never take, phrased as concretely as the list above. Not "be careful with user data" but "does not delete records" and "does not send email to customers".
If that list is hard to write, that is the finding. It usually means the boundary was never decided, only assumed, and an assumed boundary is one the model has no way to respect.
Then keep the list under version control and diff it when you change the prompt, because refusals get quietly dropped during a rewrite far more often than capabilities do. A plain text diff between prompt versions catches it in seconds.
Fuller write-up of the repo, including what running these agents actually costs once you count the data subscriptions: Claude for Financial Services.
Top comments (0)