Sentinel Vault is one of the Runs on Atlassian apps that review Confluence content with AI, through the Forge LLMs API, with no API key to paste and no external data egress declared. That's the whole pitch in one sentence. Now say a request to install it has just landed in your queue, and you're the admin or the security person who has to sign off on the AI feature. Then it's also the sentence I'd tell you to distrust. At least until somebody shows you the code behind it.
So this post shows you the code behind it. We read Sentinel Vault at commit e649d9f on 1 October 2026. We ran Atlassian's own eligibility check against the production build, and we read Atlassian's Forge LLMs pages the same day. We also ran one live AI review on our test site, which runs the development build. What follows is what holds, who pays for it, and how the cost ceiling actually behaves. Then come the four places where it's weaker than the marketing line. If you're deciding whether to switch the AI review on, I'd say the weak places matter more than the strong ones.
Why Runs on Atlassian apps matter to a Confluence admin
Most AI features in third-party apps work the same way. You paste an API key for OpenAI or Anthropic into a settings screen. From then on, the app sends your page text to that provider. The provider's terms now apply to your content. Your security team has to review a second vendor, and somebody has to own the key.
Runs on Atlassian is Atlassian's badge for apps that don't do that. Atlassian's page on the program sets out three requirements: "Apps exclusively use Atlassian-hosted compute and storage. Apps support data residency that matches data residency provided by the host Atlassian app. Customers can control external data egress (for example, analytics and logs) via admin controls." It also says it plainly. "Your app must not egress data, with the exception of egress for analytics purposes."
You don't apply for it. In Atlassian's words, "The Runs on Atlassian badge is automatically applied to eligible apps on the Atlassian Marketplace." That matters for trust. The badge is decided by Atlassian from the deployed app, not from what the vendor writes on its listing.
The Sentinel Vault listing page renders that badge today. The Marketplace's public REST API has no field for it, so we checked the listing page itself. It carries a "Runs on atlassian badge" label in its metadata. The listing shows version 6.5.0, published 30 September 2026.
There's one caveat on Atlassian's own page that I'd like you to read before you treat the badge as a security guarantee: "While controls that limit external data egress are in place, these controls do not prevent misuse of access granted to the app during installation or abuse of the app runtime." The badge tells you where the data can go. It doesn't tell you the app uses its permissions well. That second question is what the rest of this post is about.
How the Forge LLMs API keeps content inside Atlassian
What lets an AI feature live inside that badge is the Forge LLMs API. Atlassian announced it as generally available on 29 July 2026: "Today, we're announcing the Forge LLMs API is now generally available for all developers." The same announcement has the part you'll care about as an admin: "Apps that use the Forge LLMs API can qualify for Runs on Atlassian because app data stays contained within Atlassian cloud the entire time."
The API reference says it again from the developer's side: "The app retains its Runs on Atlassian eligibility after the module is added." Atlassian's main page on the API puts it more simply still. "Apps using this API are badged as Runs on Atlassian."
Atlassian also says it filters what goes through: "Requests to Forge LLMs undergo the same moderation checks as Atlassian first‑party AI and Rovo features. High‑risk messages (per the Acceptable Use Policy) are blocked."
In Sentinel Vault, the declaration is four lines in manifest.yml. It names one module, sentinel-vault-llm, and one model family, claude. There's no remotes section and no external permissions block. We grepped the manifest for both and got nothing back. Beyond its Confluence scopes and a content-styles entry, it asks for read-only JSM Assets scopes (used to import classification levels) and app storage. None of them is an external fetch permission.
What data egress the app does not do
The cleanest proof isn't the manifest, though. It's Atlassian's command-line check. Forge ships a forge eligibility command that tells a developer whether a deployed version qualifies. We ran it against both environments of the same app on 1 October:
$ forge eligibility -e production
The version of your app [6.5.0] that's deployed to [production] is eligible for the Runs on Atlassian program.
$ forge eligibility -e development
The version of your app [8.98.0] that's deployed to [development] is not eligible for the Runs on Atlassian program.
- App is using a webtrigger module that can egress data
Same app, two builds, two answers. The difference is the useful bit. The development build carries a web trigger we use for our own test harness. A web trigger that can return arbitrary data counts as possible egress. The production deploy script strips that module out before it deploys (scripts/deploy-prod.sh runs strip-dev-modules.mjs and then runs the eligibility check). The stripped build passes.
You can see the same care in a newer feature. Since mid-September the app has a configuration REST API, and it was built deliberately as a static web trigger. The comment in the manifest says why: "a static trigger cannot egress anything the manifest did not spell out, which keeps 'Runs on Atlassian'." The static trigger was in the design from the first commit. That's the kind of decision I like to see a vendor make up front. The implementation commit also closed an internal review's findings, one being that the roster, AI settings and rule text are "never mirrored" out through it.
Atlassian-hosted LLMs: no API key to hand over
Because the model runs on Atlassian's side, there's nothing for you to configure in terms of credentials. The AI section of Sentinel Vault's default configuration has a switch, a model, thresholds and a budget. It has no key field. The comment at the top of the app's LLM client explains the architecture in two lines. The admin screen gives you the consequence in one: "Token usage is billed to this app's Forge account, so AI is off by default and limited to Claude Haiku."
That last sentence is the interesting one, so let me take it apart.
Off by default, proven from the defaults
Marketing pages say "off by default" easily. We wanted it from the code. The default configuration in logic.js sets the validation engine's master switch to enabled: false. AI has its own switch, ai: { enabled: false, ... }, which is independent of the master switch and also off. Then every place that could start an AI review checks ai.enabled === true before doing anything. That's the manual review request, the workflow gate, and both entry points in the background worker. A fresh install doesn't call the model at all until an admin turns it on.
The manual review is also admin-only. The code works out the space from the page itself instead of trusting what the caller says. Anyone else is refused with "Only an admin of this page's space can run an AI review". So a reader of a page can't spend the budget by clicking a button. An editor can, though. In most setups, requesting a move into an AI-gated state queues a review, and so does re-requesting it after an edit.
What the model actually sees
The app sends text only. Anything multimodal is flattened to text before the call. The page text is capped at 40,000 characters by default, and the answer at 4,096 output tokens. Atlassian's own limits per installation are a 200,000-token context, 100 requests a minute and 50,000 tokens a minute per model. A single page review sits comfortably inside them.
The call itself runs in a background queue, not in the request you make. The worker's comment explains why: "The Forge LLM call can exceed the 25s resolver limit, so it runs on the ai-validation-queue (120s function)." You ask for a review, and the verdict lands a little later.
What the Forge LLMs API costs, and who pays
Here's the fact that shapes everything else. Atlassian's pricing page for Forge LLMs is explicit: "No free usage allowance: The Forge LLMs API does not include a free monthly usage quota. All token usage is billed. Forge LLMs usage is charged to the developer of the Forge app and counted toward your Forge monthly bill."
The developer of the app is us. Every token your site spends on a Sentinel Vault review is a line on LeanZero's Forge bill, not on your Atlassian invoice. The first line of the app's LLM client says the same thing in code: "Token costs bill to the app vendor's Forge bill, so we enforce a Haiku-only policy (see isForgeLlmModelAllowed) at every layer."
Atlassian's pricing page puts Haiku 4.5 at 10 credits per million tokens. That works out at $1 per million input tokens and $5 per million output tokens. Sonnet 4.5 is $3 and $15. Opus 4.6 is $5 and $25. We don't have a measured bill for Sentinel Vault reviews to show you. The app doesn't surface usage yet, and we didn't read the counters on our installs. So we won't quote a cost per review. The rate card is the honest number.
The way I see it, this changes the incentive in a useful way for you. The vendor has every reason to keep the model small and the calls few, because the vendor is paying. A bring-your-own-key app has the opposite incentive: your key, your bill. If you'd like the other side of that trade laid out for a Jira app, our Forge LLM tutorial for Jira builds AI validators on the same API.
Haiku, forced at four points
The manifest declares the model family claude. Atlassian's models page lists three tiers under that family: "Forge LLMs supports Claude models across three tiers: Haiku, Sonnet, and Opus." So the platform would let the app call Opus. The narrowing to Haiku is our policy, not Atlassian's. It's enforced in code at four separate points.
The rule itself is a one-line regular expression, /haiku/i, on the model id. It's applied when the admin screen lists the models it offers, so only Haiku appears. It's applied when a configuration is saved, so a non-Haiku id is never stored. That includes saves through the new configuration REST API. It's applied in the background worker, in both of its entry points (one marked with the comment "cost backstop"). And it's applied in the chat adapter itself, which logs that a model "not allowed" was clamped.
One detail about direction, because it matters if you're scripting configuration. A non-Haiku model id isn't rejected with an error. It's quietly replaced with claude-haiku-4-5-20251001. If you set a Sonnet id through the API, the save succeeds and the app runs Haiku anyway.
Our own design note on the cost ceiling describes this as three layers. The code at today's commit has four. The worker check was already there when the note was written, and the note missed it. I'm only mentioning it so that if you compare the two, you know which one to believe.
How a failed call is handled
The client retries transient failures. That means HTTP 429, 408, any 5xx, or a matching error message. It makes up to four attempts in total, waiting 400, 800 and 1,600 milliseconds between them. There's a Math.min(2000, ...) cap in that delay calculation that can never actually bind, because the largest delay it computes is 1,600. It's dead code. Harmless, and a fair sign that nobody has had to tune it.
The Forge LLMs chat() call has no structured-output option. So the app can't ask the model for JSON the way some APIs allow. Instead it adds an instruction to the system message: "Respond with ONLY a valid JSON object. No markdown fences, no surrounding prose, no explanation outside the JSON." Then a tolerant parser recovers the answer. It strips code fences and tries a plain parse. Next it falls back to the outermost object or array, repairs unescaped quotes, and repairs a truncated answer. It returns nothing instead of throwing. Its test file passes 11 of 11 on today's code.
What happens when that still fails is the part I like. In a workflow gate, if the model's answer was cut off by the token limit, the code detects it and never grants a pass. The verdict is "AI review was cut off (too long) — please retry." A manual review doesn't check this. It keeps whatever the parser salvaged from the cut-off answer. If the answer can't be parsed at all, the app records it in the audit trail and posts no comment. The comment in the code reads "fail-closed — never fabricate". Atlassian's API reports input and output tokens on every response, and the worker adds them to a counter after each call.
AI access to Confluence without giving up control: what admins check
Sentinel Vault uses the AI in two ways. A space admin can ask for a review of a page. And since July, a workflow state can require an AI review before a page is allowed to move into it.
That second one means correcting something we say ourselves. Our Sentinel Vault product page says "The deterministic engines do the enforcing — the AI advises." That was true when it was written. Since the July release that added transition conditions, a workflow state can carry an entry condition, requireAi. When it's on, the AI's verdict can block a page from entering that state. It's off by default, and you choose the threshold (low, medium or high, default medium). But once you switch it on, the AI is enforcing, not advising. The code is the thing to trust here.
The gate reviews the pinned version of the page, the one being approved, not whatever the latest edit is. If the model call fails, the page can't be read, or the answer can't be parsed, the gate ends in a terminal failure. Never a pass. Since late September, space admins who act as approvers also count towards the quorum on an AI-gated approval.
If you haven't seen the rest of the app, our post on why you can't block a Confluence save and what Sentinel Vault does instead covers the detect-and-restore engine that does the non-AI enforcing.
The four places the cost ceiling is weaker than it looks
Our product page says the AI is "capped by a monthly token budget you set". That's true. It's also the sentence we'd push back on hardest if we were reviewing this app for our own site. All four of the following are open in the code today, and our own design note lists them.
The default budget is zero, and zero means unlimited. The configuration line reads monthlyTokenBudget: 0, // 0 = unlimited, and the admin screen starts at the same value. If you switch AI on and don't set a number, there's no ceiling.
The budget is counted per space. The counter key is built from the space key and the month, ai-usage-<space>-<month>, and kept for 120 days. A budget of a million tokens is a million tokens for each space that uses AI review, not for the site. Fifty spaces means fifty ceilings.
The budget is checked when a review is queued, not when it runs. Both the manual review and the workflow gate check the counter before they put a job on the queue. The background worker that actually calls the model has no budget check at all. We grepped it for the word and got nothing back. It only adds to the counter afterwards. So a burst of requests queued at the same moment can all pass the check before any of them has been counted. The space can end the month over its number.
And there's no screen that shows usage. We searched the whole user interface for the token counters and found zero references. The number exists in storage, but no admin can see it from inside Confluence.
When the budget is reached, the behaviour is clear and you choose it. A manual review is refused with "This space has reached its monthly AI token budget." A workflow gate fails with "This space has reached its monthly AI budget.", unless you set the state's onBudgetExhausted option to allow. Then the page passes with a warning. The default is block.
A required AI gate passes when AI is switched off
This is the one I want every admin to know before building a workflow on it. If a state requires an AI review and AI is switched off, at the site level or for that space, the condition is skipped and counts as passed. The message is "AI review isn't enabled — condition skipped."
I think that's a defensible choice. Otherwise, switching AI off would freeze every page in a gated workflow. But it means the AI switch is also a bypass switch for the gate. If you rely on the gate for something a compliance reviewer will ask about, decide who's allowed to turn AI off. Then write that down.
The space-level setting wins over the site-level one when it's set, so check both places.
Adding AI is a consent event, not a silent update
Atlassian makes this part safe from your side. From the Forge LLMs page: "Administrators will be informed (via Marketplace listing and during installation) when an app uses Forge LLMs. Adding Forge LLMs—or a new model family—to an existing app triggers a major version upgrade requiring admin approval." A vendor can't add a new model family to an app you already approved without you seeing it. It also means the narrowing to Haiku is something we chose inside the family you already consented to. A later move to a different family would come back to you as an approval.
What we have not tested, and the date to watch
There are three things we couldn't verify, and you should weigh them.
We ran one live AI review for this post, and only one. Our test site, wolfaenpak, runs the development build (8.98.0). As the eligibility output above shows, that isn't the one that carries the badge. On 1 October we switched AI review on there with Haiku, one custom rule ("No personal email addresses on shareable pages."), a style line about acronyms, the tone "Plain and direct" and a budget of 200,000 tokens. Then we ran it on a throwaway test page with a made-up email address, two acronyms nobody had spelled out, and a sentence of marketing fluff. About 12 seconds after we clicked Run AI review, the panel showed three findings. One was high, for the email address. Two were medium, for the acronyms and the tone. The job result, which the panel doesn't show, reported 446 input and 465 output tokens. Afterwards we put the settings back the way they were and deleted the page. That one run is what the two screenshots with AI switched on show. Everything else about review behaviour above, including the workflow gate, the budget refusal and the truncation handling, comes from reading the code at e649d9f, not from watching it. The production build is on our demo site, and that's where the off-by-default screenshot comes from. We didn't switch AI on there.
We have no measured token bill, for the reason already given.
And there's a date on Atlassian's models page you should know about. The app allows exactly one model id, claude-haiku-4-5-20251001. It's also the only Haiku on Atlassian's list. Its retirement date reads "Not sooner than October 15, 2026". Atlassian adds that "Where possible, Forge provides at least six months' notice for model deprecations," and "not sooner than" isn't a retirement announcement. But from reading the code, if that id were removed, the Haiku filter would leave no models to offer. And the fallback is the same id. What the app would actually do at that point is something we've inferred from the shape of the code. Nobody has tested it. If you rely on the AI gate, watch for a Sentinel Vault release that updates the model before that id goes.
Two smaller things. The code comment at the top of the LLM client still calls Forge LLMs "Preview since 2026-06-01". It's been generally available since 29 July. And the tier pricing for AI in our internal design note is marked as an open decision, so don't read anything about AI tiers into the current listing. The listing's payment model is Paid via Atlassian.
Should you switch it on?
If your security review stops at "where does the content go", Sentinel Vault's AI review clears it cleanly. You get Atlassian-hosted models, no key, no declared egress, and an eligibility check from Atlassian's own tool, not from us. If your review goes on to "who controls the spend and what happens when it fails", the answers are mostly good and partly unfinished. Failures close instead of opening. The model is pinned small. A gate's verdict comes from the version being approved. But if I were you, I'd set a budget explicitly, and remember it's per space and enforced at queue time. Accept that you can't see usage yet. And treat the AI switch as a gate bypass.
If you're weighing this against your own data security policy, our post on what an Atlassian data security policy actually needs to guard and the Forge app access rule tutorial cover the controls that sit around any app, this one included.
Originally published on leanzero.net. More Atlassian, Forge and local-AI write-ups at leanzero.net/blog, and if you're planning a migration or a Forge app, that's what we do: leanzero.net/services.



Top comments (0)