Written for: developers who built something real with Lovable, Bolt or v0 and are now being asked to make it multi-tenant.
This is Part 2 of The Last 30% — a series about the part of an app you still write yourself. Part 1 was about Supabase RLS: it protects rows, and Postgres's answer for columns — grants — is per database role rather than per JWT claim, which is where it stops fitting. This one is the same argument one altitude up. Part 1 asked what the database doesn't cover. Part 2 asks what generation doesn't produce at all.
The short version: AI app builders are excellent at code with high pattern density and local correctness — a component, a form, a schema, a page. They are weak at properties that are global: authorization, multi-tenant isolation, compensating transactions, idempotency, and audit. Generation is local. Invariants are global. That mismatch is the entire gap, and it is invisible in a working demo — which is why nobody catches it until the second customer, or the first refund.
(Disclosure: I work on Supero, which is in this category and splits the problem differently. There is a section below on where these tools beat us, because a comparison that claims a clean sweep isn't worth your time.)
Type a prompt into Lovable, Bolt or v0 and ninety seconds later you have a working app. Real React, sensible components, a layout better than most of us produce by hand, deployed to a URL you can send to someone. The reflexive engineer dismissal of that is wrong, and this post isn't it.
Why some code generates beautifully
Generation works spectacularly on things with high pattern density and local correctness:
- A React component. Thousands of examples in training, obvious success criteria, and wrong is visibly wrong.
- A CRUD form. Even more examples.
- A landing page. The model has seen a million.
- A schema for a blog. There is almost a canonical answer.
What those share is a property worth naming, because it's the whole mechanism: you can look at the output and tell whether it's right. Tight feedback loop, local correctness, and 95% is fine — because the last 5% is visible and quick to fix.
Here is the shape of handler you get. Not a strawman; this is good code:
@app.get("/invoices")
def list_invoices(db: Session = Depends(get_db)):
return (db.query(Invoice)
.order_by(Invoice.created_at.desc())
.limit(50)
.all())
Readable, typed, paginated, correct ordering. Reviewed on its own it passes. The problem only exists in relation to code that isn't on screen: there's no tenant predicate, and you cannot see that from here. You can only see it by knowing what every other handler in the app does.
Add the predicate and the local problem goes away:
@app.get("/invoices")
def list_invoices(db: Session = Depends(get_db), claims = Depends(auth)):
return (db.query(Invoice)
.filter(Invoice.tenant_id == claims.tenant_id) # <-- the whole ballgame
.order_by(Invoice.created_at.desc())
.limit(50)
.all())
Now notice what you actually bought. Not isolation — one isolated endpoint. Isolation is only true if that line is true of every query, including the one written six months from now by someone who read none of the first fifty. That's a global property, and no amount of per-prompt generation produces one.
The list of things that don't generate, and why it's always the same list
| The gap | Why generation misses it |
|---|---|
| Authorization | Not login — login is a form and a token, and generates fine. Authorization is who may see which rows and fields, under which conditions, and it's a property of your whole query surface. Globally correct, or not correct at all. |
| Multi-tenant isolation | True only if true of every query ever written against the schema. |
| Compensating transactions | The charge succeeded; the order insert failed. What reverses it? Often a surprising amount of code, exercised only in failure, with no visual signal when it's absent. |
| State machines that refuse |
refunded → paid should be a 409 and an unchanged record. The generated answer is usually an UPDATE. |
| Idempotency | The network retried. Did the customer get charged twice? Nothing in the output tells you. |
| Audit | Not "do you log" — can you answer, in six months, who accessed this record. |
Every one of those is invisible in a working demo. They're properties of the system under failure, under adversarial use, or under time. You cannot look at the screen and see that compensation is missing. That's the difference between the two columns:
On the number: "70/30" is my estimate from watching teams take generated apps to production, not a measured statistic, and I'd rather label it than have it quoted back as one. The proportion is arguable. The kind of thing in the right-hand column is not, and the kind is the point.
Credit where it's due: one of those six gets real attention
"They skip authorization" would be a tidy line for me to write and it isn't accurate, so I'm not writing it.
Lovable runs a Quick scan automatically every time the publish dialog opens: it checks your database access rules and flags tables with no row-level security and rules that let everyone through. A separate Deep scan reviews your application code — including server code that bypasses the rules your database enforces — but it does not run in that dialog; you trigger it from the Security view. Both are free on every plan, and a workspace can switch on a setting that blocks publishing while critical findings are unresolved. That is a shipped feature and it is more than most of this category ships. Lovable also creates RLS policies itself when it builds a feature that stores user data. Its own documentation is candid about the ceiling — these tools "cannot guarantee complete security" — and recommends a professional review for anything sensitive.
Two things remain true, and they're more interesting than the tidy line:
-
The policy a builder reaches for by default is per-row ownership, not tenant isolation.
user_id = auth.uid()is not a tenant boundary. RLS can express organisation membership — Part 1 opens with a policy that does, and calls it a real security boundary. What it can't do is guarantee that the org-scoped version, rather than the owner-scoped one, is what got written on every table — including the table added after the prompt that created the first fifty. And it says nothing at all about a client holding a service-role key that bypasses every policy you wrote. -
Completeness is tooled. Correctness is not. A scanner can tell you which tables have no rule — Lovable's publish scan does, and so does Supabase's Security Advisor, and you should use both. What no scanner tells you is whether the rule you wrote expresses your model: that
user_id = auth.uid()is the right predicate for an org-scoped app, that a tenant-scoped role is honoured, and that the application layer above re-derives the same decision. The disclosure behind the CVE below says precisely this about Lovable's publish check — it "will help ensure that RLS policies are enabled in all tables and notify if they aren't (but doesn't necessarily indicate if they are sufficient)". Table-level completeness is tooled. Semantic correctness across your whole query surface is the part you still own.
The CVE, stated precisely
The stakes here are documented rather than hypothetical — but the record needs quoting exactly, because the imprecise version of this story is everywhere and it's unfair.
CVE-2025-48757. CVSS 3.1 base score 9.3 CRITICAL, vector AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N — assigned by the CNA and carried in the NVD record as a Secondary metric. NVD itself has published no score: vulnStatus is Deferred, meaning NVD will not analyse the entry. Published 2025-05-30, last modified 2026-06-17. NVD's description, verbatim:
An insufficient database Row-Level Security policy in Lovable through 2025-04-15 allows remote unauthenticated attackers to read or write to arbitrary database tables of generated sites. NOTE: this is disputed by the Supplier because each individual customer of the Lovable platform accepts a responsibility over protecting the data of their application.
Five things you should take from that, in order:
-
It is bounded in the past. The affected range stops at 2025-04-15, and the entry marks every other version
unknownrather than affected. Exploitation required the project to be using an external database as its backend. Nothing here says anything about Lovable's security today, and I'm not implying it. - Provenance, because precision cuts both ways. The finding was disclosed by Matt Palmer at Replit, Inc. — a Lovable competitor, as am I. His own submitted vector was scope-unchanged and scored 8.26, not the 9.3 the CNA filed. Lovable acknowledged it in three days and shipped a patch on 2025-04-24, five weeks before public disclosure. Every one of those facts cuts against the scary version of this story, and all three are in the record.
-
NVD carries it with a
disputedtag (alongsideexclusively-hosted-service), withvulnStatus: Deferred. The supplier's position is in the quote: each customer accepts responsibility for protecting their own application's data. You can agree or disagree with that division of responsibility — I'd just rather you read it than take my summary of it. -
The widely-quoted "170+ affected apps" is the disclosing researcher's scan, not Lovable's figure, and I have not independently reproduced it. If you cite it, cite it that way. Two fields in the same entry carry more weight than any count: it is classified CWE-863, Incorrect Authorization — the taxonomy's own name for this layer — and the CISA Coordinator's SSVC assessment records
exploitation: proof-of-concept,automatable: yes,technical impact: partial. - The reason to mention a disputed, historical CVE at all is that it is the clearest public evidence of where generated apps break. Not in the components. In the layer that has to be globally true. Lovable is also the only one of the three with a public entry of this kind, and that is not evidence it is the weakest — it is evidence that someone looked, told the vendor, and wrote it down, which is more than exists for the other two. Read the asymmetry as a gap in the public record, not as a ranking.
Four tests, about an hour, useful regardless of what you build on
This is the part I'd actually do before the next planning meeting. None of it involves me or my product.
1. Be the second tenant. Create two orgs, two users. With tenant A's token, request a record id owned by tenant B:
curl -s -o /dev/null -w "%{http_code}\n" \
-H "Authorization: Bearer $TENANT_A_TOKEN" \
https://your-app.example/api/invoices/$ID_OWNED_BY_TENANT_B
403 or 404 is right. 200 is the whole post. Then do it again on a list endpoint with a filter parameter, and again on whatever search or export route exists — those are the ones that get added later and miss the predicate.
2. Break a payment in the middle. Do tests 2 and 3 in your payment provider's test mode, with test keys, against a staging tenant — the point is to see the state you're left in, not to move real money. Make the write after the charge fail, on purpose:
charge = payments.create(amount=4900, currency="usd") # succeeds
raise RuntimeError("simulated failure after the charge")
db.add(Order(id=order_id, charge_id=charge.id)) # never runs
db.commit() # never runs
Money moved; no order exists. Now go look for the code that reverses the charge. If you can't find it in under two minutes, it isn't there — and it is often a surprising amount of code, enough that teams defer it.
3. Retry a charge. Fire the same mutation twice and compare:
for i in 1 2; do
curl -s -X POST https://your-app.example/api/orders/42/capture \
-H "Authorization: Bearer $TOKEN" \
-o /dev/null -w "attempt $i -> %{http_code}\n"
done
Then check the payment provider's dashboard for two charges. Two 200s and one charge is fine (something short-circuited). Two 200s and two charges is a bug with a customer attached to it. While you're there, set an order to refunded and POST the pay transition — you want a 409 and a record that didn't move.
4. Ask the audit question. Not "do we log" — run the actual query a security reviewer will ask you for:
select actor, action, occurred_at
from audit_log
where record_type = 'invoice'
and record_id = $1
and action = 'read'
order by occurred_at desc;
If that returns zero rows because nothing ever wrote a read event, you have write-side audit only. That's worth knowing in advance of the conversation rather than during it. (Keep reading — so do we.)
The three ways teams handle the 30%
Write it yourself. Correct, and it's the part that takes most of the time — on a codebase you didn't design, whose conventions you're inferring as you go.
Ship without it. Extremely common, and genuinely fine until it isn't. Internal tools and prototypes do not need compensating transactions. The problem is that prototypes get promoted, and nobody ever schedules the sprint where the prototype becomes a product.
Put it in the platform. Don't generate the invariants — inherit them. Generate what generates well, and let the rest be a structural property of the thing your app runs on.
How we split it, including the rows I'd rather not print
Supero generates the same layer everyone generates: data model, CRUD REST API, a generated client library, OpenAPI docs, admin console, customer web UI, deployed and running. Be precise about that client library, because our own API reference is: it is not typed. The reference says in as many words that the generated client documents your fields and gives you no compile-time or editor-time checking of them — a misspelled type name is still a runtime failure. Python is the one you can actually install (supero, on PyPI, and its own classifier says beta). The JavaScript package our docs tell you to npm install is not published, so that command 404s today. Go and Java are registry stubs, and the failure mode is the one this whole post is about: ask for java and you get 202 queued and then a Python and a JavaScript artifact, with nothing in the response marking the language as unavailable. The rest is underneath, not generated:
| The invisible part | How it's provided |
|---|---|
| Tenant isolation | Structural: the tenant comes from the authenticated principal and is resolved once in the shared data layer every service calls, so a handler cannot omit it. Our own published multi-tenant checklist scores that honestly and I'll borrow its words — a developer cannot leak by forgetting, and can leak by being provisioned wrong. Two documented ways: the admin tier has standing exemptions, and anyone left in the auto-created default-tenant gets project-wide read across every other tenant in that project, whatever their role — so a real customer must never be put there. On a connected external source the binding is per table, and it's worse than configuration you can get wrong: a live table created through the console's wizard is always saved with a fixed tenant binding whatever you type into the tenant-routing panel, and a fixed binding does no per-row filtering at all. The policy layer sitting on top of all that is not deny-by-default: an entity with no access policy gets full create, read, update and delete rather than a refusal. Deny-by-default isn't a switch we ship — it's something you build, an access policy per role with default_access: "none" plus a rule per entity you want to allow — and that object policy is evaluated on writes: it is not consulted on a read, so setting an entity to "none" stops a role writing it and leaves the records readable. Separately, our platform-, domain- and project-admin roles bypass access policies outright, and that tier includes project-scoped API keys, not just human super-admins — so a key you mint for your own integration is not constrained by the policy layer. The tenant admin inside your app does not bypass. Part 1 carried the same admission. |
| Field permissions | A field a caller can't see can't be distinct-ed, grouped by or aggregated over, including via a raw pipeline — enforced on the query path, not only on the response, and asserted end to end by test_53_access_policy_e2e. Sort and filter are a different story, and the honest version is worse than the hedge I had here before: they're guarded on a live external source, and on our own storage they are not. The guard that would refuse them is deployed in monitor mode — it records that it would have denied the read, then returns the rows anyway — so an order_by on a hidden field gets you a 200 and the argmax row. That's published in full, with the request that demonstrates it, and it's Part 3 of this series: permissions that leak through ORDER BY. The response-side strip is defence in depth, not the strong half: it falls through on a policy-resolution failure, a missing session context, an empty policy and an unexpected error, and credential-shaped fields are held back by a separate hard-coded denylist rather than by the policy engine. |
| Role escalation | Custom roles carry a permission ceiling: they can't shadow a builtin, exceed the creator's base, or forge is_builtin. |
| Compensation | A saga orchestrator — step 4 fails, steps 3, 2, 1 reverse, as platform code rather than yours. Reversal depends on each operation declaring a correct inverse, and fourteen of ours are currently mis-declared. When an inverse can't be built the walker refuses to fire it and stops — which abandons the reversal of every earlier step in that saga too, not just the one. We found them because the suite that was supposed to cover them had been exercising its own fixture rather than the real manifests. Ask us which operations are affected before you build a money path on one. |
| Invalid transitions | State machines return 409 and leave the record alone. |
| Idempotency | A state-guarded transition refuses an out-of-state repeat, so a retry doesn't write the record twice. Two things it is not: we are not the payment processor — we record the charge your app made at Stripe or PayPal, so a duplicate charge is prevented by the idempotency key you send them, not by us; and the guard is read-then-write with no lock, so two genuinely concurrent retries can both pass it. Cumulative operations like partial refunds need your own key either way. |
| Audit | Write-side only. We can tell you who wrote a record. We cannot tell you who read one. And check the plan before you count on even that half: the audit log is bundled from Pro up and is a $49/mo add-on below that, so on a free or entry plan it isn't write-side-only, it's absent. |
Several of those rows are admissions, not features, and they are the ones I'd raise first rather than a complete list. Audit is on my own list of six, so leaving it out of our own table would have made this post a sales page — it's the gap I'd raise first if I were evaluating us, and test 4 above will find it. Permissive-by-default access is the other, and the test for it is not test 1. Test 1 is the one the shared data layer is built to pass, which is exactly why it won't show you this: run it and you get a clean 403. (Not a flawless pass either — see the default-tenant and admin-exemption admissions in the isolation row above.) Create a second non-admin user inside your own tenant, leave an entity with no access policy, and try to list, edit and delete a row belonging to the first user. Ours returns the row and lets you write it. I'd rather you hit both here than after signing something.
The fine print, since you'd find it anyway: connector write-back in place runs only in live read-write mode, it is implemented for Postgres (Supabase included) and MySQL, and it is validated end to end on Postgres only — our own supported-sources matrix says MySQL live write-through has not been verified end to end, so don't take my word for it in either direction. In sync mode a connector is a copy: an edit on our side returns a 200 and never reaches your database. There is a signal — we stamp an access_mode on the response — but it is per connector, not per table, and when one connector mixes modes the most permissive one wins, so that stamp can read live-rw for a write that actually landed in our local copy. We're designed for SOC 2 controls — not certified against SOC 2, ISO 27001 or PCI DSS. On HIPAA, read our compliance page rather than my summary: no Business Associate Agreement is offered, and that page tells you in as many words not to put protected health information into Supero Cloud on the strength of it. Read audit is the control a HIPAA reviewer asks about first, which is precisely the row above. Custom domains: the code is built but dark, and our published position is blunter than that — there is no self-service flow on any deployment target, no DNS automation and no certificate provisioning anywhere in the deploy path, we publish no timeline, and the $25/project/month custom-domain line in our own catalog does not enable one if you buy it. A scheduled connector sync you can configure and save, and it validates — but whether the runner actually fires it depends on the deployment, so confirm a real run in Run History before you depend on one. And our source export is one-way, in the two senses spelled out below.
The distinction that actually matters: the platform provides these; the prompt does not emit them. They aren't better-generated code. They're code written once, by people, tested adversarially, that your generated app runs on top of.
Where Lovable, Bolt and v0 beat us
Visual quality. They're better. The design-first builders produce interfaces that look considered; ours look like competent conventional software. If the interface is the product, use them.
Speed of first result. Theirs is seconds — a sign-up and a free token budget, no access request. Ours is a free signup (no card) and then under a minute to a few minutes. The deployment story is the real gap, and it deserves stating in the direction that's fair to both sides: our free plan does hand you a public preview URL, but it's reaped after about thirty minutes, and a deployment that stays up starts at $9.99/mo (listed at $19.99, currently half off) where theirs is a public URL on the free plan. That is a real disadvantage when someone is evaluating three tools on a Thursday afternoon.
Iteration feel. The prompt-see-adjust loop on a UI is tight and pleasant in a way that generation over a governed backend is not.
Code you own from second one. All three say you own the code they generate — read each one's terms, they word it differently — and their access models differ in ways worth knowing:
-
Bolt has the most straightforward path: project title → Export → Download hands you a
.zipof your project files — code, not the database or your environment variables. Its docs state no plan requirement for it, which is worth noting because the same docs do say custom domains and bringing your own Supabase are paid-only. Its GitHub integration also runs both ways: it auto-commits your changes and polls GitHub every 30 seconds for commits made outside Bolt (its docs note Bolt's copy wins a simultaneous-edit conflict). - Lovable gates full codebase download to paid plans, but Git sync is a genuine two-way sync to a repo you own on GitHub, GitLab (including a self-managed instance) or Bitbucket, on any plan, with GitHub Enterprise Cloud/Server on the Enterprise plan. It exports from Lovable; it won't import an existing repo.
- v0 doesn't need a repository at all — Vercel's docs say Git integration is optional, and projects without GitHub deploy their current code straight to Vercel. Connect one and v0 creates a private repo in an account or org you choose, and that sync is bi-directional. Publishing goes to Vercel, though the docs say you can export the code and deploy elsewhere.
Check each one's current docs before you rely on any of that. It's the fastest-moving column in the comparison and any of it may have changed since the date at the bottom of this post. None of the three is relationship-free, but the defaults are not what the lazy version of this sentence says: Bolt's default home is its own Cloud (.bolt.host plus a Bolt database on every plan; Netlify is the documented alternative, and using your own Supabase account needs a paid plan), Lovable now ships its own managed Postgres with Supabase as an optional connection, and v0 publishes to Vercel though its docs say you can export and host elsewhere.
Their code access is more open than ours in two ways, not one. Bolt's and Lovable's and v0's Git syncs all run both directions where our export is one-way — you get the code, your edits to it don't come back. And their output runs on its own where ours doesn't: the code we hand you proxies its /api/* calls back to Supero, so it needs Supero behind it — our cloud, or your own AWS or GCP account where the UI is yours and the backend still calls us, or a licensed install on your own servers, which is an Enterprise arrangement we have not yet published a customer runbook for. Worth saying plainly — the tightest coupling in this comparison is ours. If walking away with a standalone codebase is the requirement, they win that outright.
The use-case split, stated plainly: marketing site, landing page, design-led consumer app, or a prototype for an investor next week — use them. That's the correct recommendation and we'd be the worse tool for it.
When the 30% starts to matter
It isn't a size threshold. It's a set of events:
- A second customer organisation logs in, and neither may see the other.
- Money moves, so a half-applied sequence becomes a reconciliation problem instead of a bug.
- A customer's security review asks how field permissions are enforced.
- Regulated data arrives — anything where a histogram is a leak.
- A second developer joins, and invariants stop being enforced by one person's memory.
Before the first of those, this post is a distraction. After it, retrofitting means auditing every query path and then installing a mechanism that makes the next one safe. That is expensive, and it is the one piece of work you cannot do file by file in isolation — but it is a refactor with a checklist, not a rewrite. Part 1's audit query is the Supabase-shaped version of that checklist.
FAQ
Can Lovable or Bolt build a multi-tenant SaaS?
They can generate the pieces, including Supabase RLS policies, and Lovable's publish scan will tell you which tables ended up with no rule at all. What's missing is anything that checks the rules you got express your tenancy model across the whole query surface.
Is Bolt or Lovable production ready?
For the frontend, and for single-tenant apps, frequently yes. The question to ask isn't about code quality — it's about the invariants: isolation, compensation, idempotency, audit.
What does an AI app builder not generate?
Anything globally correct rather than locally correct: authorization across the whole query surface, tenant isolation, compensating transactions, idempotency, and read audit.
Can I use an AI builder for the frontend and something else for the backend?
Yes, and it's a sensible split: generate the UI where generation is strongest, and put the invariants in a layer that enforces them for every query.
How do I test for the missing 30%?
The four tests above. Fail a step in the middle of a payment flow and see what state you're left in; retry a charge and check for a double; log in as a second tenant and try to read the first; run the read-audit query.
Part 3, which is already published
Part 3 goes after the leak itself: how field permissions escape through ORDER BY and aggregates — queries that are individually correct, return no hidden field, pass review, and still let a caller binary-search a value they were never allowed to see. It's up, and it includes the request that demonstrates the gap against our own storage: Your field-level permissions probably leak through ORDER BY. If you ran test 1 above and got clean 403s, that's the test that will surprise you.
What to do next
Audit your generated app against the table above. Six rows, an afternoon, and it's useful regardless of what you end up building on. Nothing in it requires talking to me.
Then, if you want to see what the other approach produces, applications generated by the platform are live and need no account — sixteen at the time of writing, with the current list at api.supero.dev/api/v1/public/showcase: concierge.supero.live · atelier.supero.live · ledgerline.supero.live. Two fair warnings. They scale to zero, so the first load is waking a cold container — give it ten seconds or so, and an idle one occasionally won't wake on the first try. And most are single-tenant public demos — they show you the generated application, the data model and the screens, not the governance this post argues about.
One exception is worth your time. sentinel.supero.live runs two insurers on one deployment — Northwind Mutual and Cascade Assurance — with four published logins across both tenants, so you can run test 1 against us directly. The two tenants use different claim-number prefixes, so a leak would be visible on sight. Same cold start as the others, so give the first page a few seconds. The login page hands you a working, shared, writable account, so treat anything you create there as public.
And if you'd rather read the output than click through it, the reference apps are on GitHub — MIT-licensed, readable Python and JS, nineteen of them: github.com/supero-platform/supero-apps. That's the honest test of the first half of this post: the generated layer is real code you can inspect in a browser tab. What you won't see in that repo is the part this whole piece is about — the invariants don't live in the generated files, which is rather the point.
Start free if the generated output is the standard you'd ship. One number the pricing page doesn't print, and you'll want it before you judge us on output: the free plan's budget is three app generations a month.
Lovable, Bolt and v0 capabilities checked 2 October 2026 against each vendor's own documentation — docs.lovable.dev/features/security, docs.lovable.dev/integrations/git-sync-overview, support.bolt.new and v0.app/docs (v0's pages are stamped 1 October 2026); CVE data from the NVD API the same day. This category changes monthly — if something here is out of date or wrong, tell me in the comments and I'll correct it. A comparison with an error in it is worse than no comparison.

Top comments (0)