AI website builders are everywhere right now. But almost all of them are platforms you rent: the users, the brand and the pricing belong to someone else.
I wanted the opposite: a small AI website builder that a freelancer or an agency could run themselves, under their own name, and charge for. That project became Talivo. In this post I'll walk through the architecture and the decisions behind it. The code snippets are simplified to show the ideas, but the patterns are the ones that matter.
Try the live demo (bring a free Gemini key; the first load takes ~30s while the free server wakes up).
The goal:
A user signs up, types something like "a warm website for a family bakery in Bologna", and gets back a complete, styled website. Every generation costs credits. The owner of the app decides who gets credits and how much to charge for them.
That gives us four building blocks:
Auth: who is making the request
Credits: whether they're allowed to make it
Generation: turning a prompt into a full HTML page
Admin: letting the owner manage users and balances
Why FastAPI + Gemini
FastAPI was an easy pick: async out of the box, very little boilerplate, automatic request validation with Pydantic, and it's pleasant for buyers to read and extend. Since the app is sold with full source code, readability was a real requirement.
Gemini had one killer feature for this use case: a free API key from Google AI Studio. Whoever installs the app can start generating sites without putting a credit card anywhere. For a product meant to be "download and run", removing that first barrier matters a lot. The Flash models are also fast enough that users don't stare at a spinner for too long.
Step 1: Generating a full website from a prompt
The core is a single call to Gemini with a carefully written system instruction. The trick is to be very explicit about the output format: one self-contained HTML file, no markdown fences, no explanations.
from google import genai
client = genai.Client(api_key=GEMINI_API_KEY)
SYSTEM_PROMPT = """You are an expert web designer.
Return ONE complete, self-contained HTML document.
- Inline all CSS in a <style> tag. No external files.
- Modern, responsive layout that works on mobile.
- Realistic copy based on the user's request (no lorem ipsum).
- Output ONLY the HTML. No markdown, no explanations."""
def generate_site(prompt: str) -> str:
response = client.models.generate_content(
model=GEMINI_MODEL, # e.g. a current Flash model from AI Studio
contents=prompt,
config={"system_instruction": SYSTEM_PROMPT},
)
return clean_html(response.text)
Even with a strict prompt, models sometimes wrap the output in a code fence anyway. So the HTML always goes through a small cleanup step:
import re
def clean_html(text: str) -> str:
text = text.strip()
# strip ```
{% endraw %}
html ...
{% raw %}
``` fences if the model added them
text = re.sub(r"^```
(?:html)?\s*|\s*
```$", "", text)
if "<html" not in text.lower():
raise ValueError("Model did not return an HTML document")
return text
Lesson learned: never trust the model's output format. Validate it, and treat a malformed response as a failed generation (more on why that matters below).
Step 2: Credits that can't go negative
This is the part that turns a demo into something you can charge for. The naive version looks like this:
# ā Don't do this
user = get_user(user_id)
if user.credits > 0:
user.credits -= 1
save(user)
Two quick requests from the same user can both read credits = 1, both pass the check, and you've given away a free generation. The fix is to make the check and the decrement one atomic database operation:
def spend_credit(db, user_id: int) -> bool:
cur = db.execute(
"UPDATE users SET credits = credits - 1 "
"WHERE id = ? AND credits > 0",
(user_id,),
)
db.commit()
return cur.rowcount == 1 # False = no credits left
If rowcount is 0, the user had no credits, and nothing changed. No race condition, no negative balances.
The second rule: charge first, refund on failure. Gemini calls can time out or return garbage, and users get (rightly) annoyed if they pay for nothing.
@app.post("/generate")
def generate(req: GenerateRequest, user=Depends(current_user)):
if not spend_credit(db, user.id):
raise HTTPException(402, "You're out of credits")
try:
html = generate_site(req.prompt)
except Exception:
refund_credit(db, user.id)
raise HTTPException(502, "Generation failed, your credit was refunded")
return {"html": html}
Step 3: Auth, kept boring on purpose
Auth is the place where I wanted zero creativity. Users register with email and password, passwords are stored as hashes (never in plain text), and every protected route goes through a FastAPI dependency like current_user above. Using Depends() means every endpoint that touches credits has to know who the user is. You can't forget the check.
Step 4: The admin panel
The admin panel is deliberately simple: a list of users, their balance, and a way to add credits.
A design choice I want to be transparent about: Talivo doesn't include a built-in card checkout. The owner tops up credits manually from the panel. That sounds like a limitation (and for some people it is), but it was intentional for v1:
The operator can charge however they want: a payment link, an invoice, a bank transfer, cash from a local client.
There's no payment provider lock-in and no platform cut.
Agencies often bundle credits into a bigger client deal anyway.
Since the full source is included, plugging in Stripe for automatic credit packs is a natural next step for anyone who needs it.
Step 5: Making it installable by non-developers
The target buyer isn't always a developer, so installation had to be as close to "double-click" as possible. On Windows, an install.bat script creates the virtual environment, installs the dependencies and starts the server. On macOS, Linux or any server it's the usual Python 3.10+ setup. The only thing the user needs to bring is their Gemini API key.
What I'd tell anyone building an AI wrapper product
The AI call is the easy part. Credits, auth, failure handling and installation are where most of the work (and most of the value) is.
Validate model output like user input. It will surprise you.
Make billing atomic from day one. Race conditions on credits are real money.
Refund on failure. It costs you almost nothing and builds a lot of trust.
Ship a demo people can actually use. Screenshots don't convince developers; a working URL does.
Try it
Live demo: demo-talivo.onrender.com (use your own free Gemini key; it isn't stored)
š¦ Get the full source code: Talivo on Gumroad, a one-time purchase at a launch price of ā¬49.99
I'd love feedback, especially from anyone who has built credit or usage-based billing before. What would you add next: Stripe credit packs, multiple AI providers, or one-click cloud deploy?
Top comments (2)
interesting that you chose fastapi for the ai layer, do you see any latency edge over a simple node wrapper? i tend to run uvicorn with multiple workers for better concurrency
Thanks! No real latency edge: the Gemini call takes seconds, so framework overhead is negligible either way. I chose FastAPI for readability, since buyers get the full source. Multiple uvicorn workers is a good call, and it's why credits are deducted with one atomic SQL UPDATE instead of in app memory. How many workers do you usually run?