The scary thing about an AI feature on a free plan is that the bill goes to you, not the user. One person with a loop script against your endpoint can burn through a month of budget while you sleep, and they're not even really a hacker, it's just cheaper than paying for the model themselves.
None of what follows is clever. It's the boring stuff I'd ship before the AI feature goes live, in the order I'd do it.
1. A hard spend cap at the provider, set lower than feels comfortable
First thing, before any code: go into the provider dashboard and set a monthly hard limit, plus a soft alert at about 50% of it. Most model providers let you do this per project.
My rule now is the cap should be a number that would annoy me, not ruin me. If the cap hits, the feature goes down and I get an email. That's a much better failure than a surprise invoice.
Also, make a separate project/key per environment. A staging key with no limit is the classic way to get surprised.
2. Per-user daily quota, counted in tokens not requests
The obvious first move is "50 requests per day". It's a trap. One request with a giant pasted PDF costs more than 50 tiny ones.
Have every call write a row: user id, input tokens, output tokens, model, timestamp. Before each new call, check the sum for that user in the last 24 hours. Free users get a small token budget, paid users get a bigger one, and nobody gets unlimited.
A plain table and one indexed query on (user_id, created_at) is fine at my size. I don't need Redis for this yet.
3. Cap the input before it reaches the model
Truncate or reject input above a size limit, and set max output tokens on every call. If you pass whatever the user sends, someone will paste an entire doc, and the model will happily write a novel back.
4. Rate limit by IP and by account
Quota is per day. Rate limiting is per minute. A script does several calls a second, which no human does in a normal UI. A limit of a handful of calls per minute per account (and per IP for signed-out stuff) stops a loop after a minute, not after a whole night.
Free accounts are also where abuse lives. I wrote about the small defense stack I use against fake signups before, and the same email verification step slows down this kind of abuse too.
5. A daily cost line in my morning check
Add one number to your morning glance: yesterday's model spend, and the top 3 users by tokens. Takes ten seconds. If one name is 10x everyone else, look.
6. Make the limit message human
When someone hits the quota, they see something like "You've used today's free summaries, they reset at midnight UTC. Paid plans get more." Not a 429 with a stack trace. That message is also a pretty honest upgrade prompt.
The short version
Treat every model call like it costs real money, because it does, and it's billed to you, not the user. Build the meter first, then the feature.
If you're keeping a side project cheap in general, this pairs well with the boring cost checklist in how to keep monthly bills under $20.
And if you've built an AI tool and want early users who will actually push it (in a nice way), you can list it on EarlyHunt when you're ready.
What limits do you put on free users of AI features? Curious if anyone does it per feature instead of per account.
Top comments (0)