You built a "chat with your documents" app. It works on your laptop. Then someone asks: "What will this cost me every month?"
Here is the short answer, using real prices I checked on October 6, 2026:
- A small test (1 PDF, 100 questions): under $1
- A small business (1,000 questions a month): about $8 to $15 a month
- A busy chatbot (10,000 questions a month): about $60 to $130 a month
With a small, cheap AI model, the answers are often the smaller part of the bill. Hosting and the database can cost more. In this post I will show where every dollar goes, how I did the math, and how to bring the bill down.
This is part 3 of my RAG series. Part 1 was How to Build a "Chat With Your PDFs" App. Part 2 was Why Your RAG Chatbot Gives Wrong Answers.
A note on prices: every price below comes from the official pricing page of each company, checked on October 6, 2026. Prices change often, so check the links at the end before you quote a client.
The 4 Parts of the Bill
- Embeddings: turning your text into numbers so it can be searched. You pay once when you upload a document, plus a tiny amount for each question.
- Vector database: the place where those numbers are stored.
- The AI model (LLM): it writes each answer. You pay for every question, so this part grows as your users grow.
- Hosting: the server that runs your app.
First, What Is a Token?
AI companies charge by the token. One token is about 4 letters, or about 0.75 of a word in English. Prices are given "per 1 million tokens".
You pay for two things: input tokens (what you send to the model) and output tokens (what the model writes back). Output costs more than input.
Current Prices (Checked October 6, 2026)
Embeddings
| Option | Price |
|---|---|
| OpenAI text-embedding-3-small | $0.02 per 1M tokens |
| OpenAI text-embedding-3-large | $0.13 per 1M tokens |
| A local model (sentence-transformers, like in part 1) | Free, but it runs on your own server and needs RAM |
Vector database
| Option | Price |
|---|---|
| Pinecone Starter | Free. Up to 2 GB storage and 1M read units a month |
| Pinecone Builder | $20 a month |
| Supabase Free (Postgres with pgvector) | $0. 500 MB database. Free projects pause after 1 week of no activity |
| Supabase Pro | From $25 a month. 8 GB disk. Does not pause |
AI model (price per 1M tokens)
| Model | Input | Output |
|---|---|---|
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 |
| Claude Haiku 4.5 | $1 | $5 |
| Claude Sonnet 5.5 | $2 | $10 |
OpenAI and other companies also have cheap models. Check their pricing pages and use the same math below.
One warning about free tiers: Google's pricing page says that on the Gemini free tier, your content is used to improve their products. On the paid tier it is not. Do not send client documents through a free tier.
Hosting (Render)
| Option | Price |
|---|---|
| Free web service | $0. 512 MB RAM. Has usage limits |
| Small paid service | $7 a month. Less than 1 CPU, 512 MB RAM |
| Medium paid service | $25 a month. 1 CPU, 2 GB RAM |
Render's Hobby workspace is $0 a month plus the compute you use. Its Pro workspace is $25 a month plus compute.
How Much Does One Question Cost?
Let us use simple numbers:
- Each chunk is about 400 words, which is about 530 tokens.
- You send
top_k = 4chunks, so about 2,100 tokens. - Your instructions and the question add about 170 tokens.
- So the input is about 2,300 tokens, and the answer is about 300 tokens.
The math for Claude Haiku 4.5:
- Input: 2,300 / 1,000,000 x $1 = $0.0023
- Output: 300 / 1,000,000 x $5 = $0.0015
- Total: $0.0038 per question. That is less than half a cent.
The same math for all three models:
| Model | Cost per question |
|---|---|
| Gemini 3.1 Flash-Lite | about $0.0010 |
| Claude Haiku 4.5 | about $0.0038 |
| Claude Sonnet 5.5 | about $0.0076 |
Note: models count tokens in different ways. Anthropic says its newer models, which include Sonnet 5.5, count about 30% more tokens for the same text. So Sonnet may cost a little more than the table shows. Treat all of these as good estimates, not exact bills.
What about embeddings? A 50-page PDF is about 25,000 words, or about 33,000 tokens. With text-embedding-3-small, that costs 33,000 / 1,000,000 x $0.02 = $0.00066. That is less than one-tenth of a cent. Embedding is almost never the problem.
Three Real Examples
| Small test | Small business | Busy chatbot | |
|---|---|---|---|
| Questions per month | 100 | 1,000 | 10,000 |
| Documents | 1 PDF | 20 PDFs | 20 PDFs |
| Setup | Your laptop and free tiers | Pinecone Starter (free) + Render $7 | Supabase Pro $25 + Render $25 |
| Fixed cost per month | $0 | $7 | $50 |
| AI cost, Gemini Flash-Lite | $0.10 | $1.03 | $10.25 |
| AI cost, Haiku 4.5 | $0.38 | $3.80 | $38.00 |
| AI cost, Sonnet 5.5 | $0.76 | $7.60 | $76.00 |
| Total with Flash-Lite | $0.10 | $8.03 | $60.25 |
| Total with Haiku 4.5 | $0.38 | $10.80 | $88.00 |
| Total with Sonnet 5.5 | $0.76 | $14.60 | $126.00 |
A few things to know about this table:
- Embedding 20 PDFs adds only about one cent in total.
- The small business setup uses an embedding API, not a local model, because a 512 MB server is too small for a local model.
- Pinecone's own examples show its free plan handling thousands of searches a day for a small knowledge base, so 1,000 questions a month is far below its limit.
- The busy chatbot uses Supabase Pro because it never pauses and it keeps daily backups.
- These numbers do not include your domain name, taxes, bank fees for paying in dollars, or your own time.
Free Stack or Paid Stack?
The free stack ($0 a month) uses a local embedding model, FAISS from part 1 (or a free database tier), your laptop or Render's free plan, and a free AI tier. It is great for learning and for demos.
It is not good for real clients, for three reasons: the Gemini free tier may use your content to improve Google's products, a free Supabase project pauses after a week without activity, and free hosting has limits.
The cheap paid stack (about $8 to $15 a month) is the small business column above. It is always on, it keeps client data private, and it is still very cheap. For most small clients, this is the right place to start.
7 Ways to Cut Your Bill
1. Send fewer chunks. Going from top_k = 4 to top_k = 2 drops the input from about 2,300 tokens to about 1,230 tokens. That is about 46% less input cost. Test it with your own questions, because too few chunks can make answers worse (see part 2).
2. Limit the length of answers. Output tokens cost more than input tokens. On Haiku 4.5 they cost 5 times more. Most APIs let you set a maximum answer length:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model='claude-haiku-4-5-20251001',
max_tokens=300, # stops very long answers
messages=[{'role': 'user', 'content': prompt}],
)
print(response.content[0].text)
3. Use the cheapest model that answers well. Test 20 real questions on two or three models. In the example above, Haiku costs half as much as Sonnet, and Flash-Lite costs about 70% less than Haiku. If the cheap model answers your questions well, use it.
4. Use prompt caching for the fixed part of your prompt. Anthropic charges only 10% of the normal input price when it reads text from its cache. This only helps for text that stays the same on every call, like a long set of instructions. The retrieved chunks change with every question, so they will not be cached.
5. Save answers to repeated questions. A business FAQ bot gets the same questions again and again. Save the answer, and the second time you pay nothing:
answer_cache = {}
def ask_with_cache(question):
key = question.strip().lower()
if key in answer_cache:
return answer_cache[key]
answer = ask_question(question, chunks, index)
answer_cache[key] = answer
return answer
This cache is lost when your app restarts. For a real app, keep it in your database. Clear it when the documents change.
6. Do not embed the same file twice. Save a fingerprint (hash) of every file. When a client re-uploads the same file, skip the embedding step:
import hashlib
def file_hash(path):
with open(path, 'rb') as f:
return hashlib.sha256(f.read()).hexdigest()
# Save the hash when you embed a file.
# Next time, if the hash is the same, skip it.
7. Protect a public chatbot. If you put the chatbot on a public website, anyone can send it thousands of questions, and you pay for all of them. Add a daily question limit for each visitor, and set a monthly budget or spending alert in your AI account.
How to Price This for a Client
If you build RAG chatbots for clients, do not guess a price. Use this:
- Add up the monthly cost: hosting + database + AI cost at the client's expected number of questions. Use the table above.
- Add your time: updating documents, fixing problems, answering messages.
- Add some room for busy months.
- Charge a one-time setup fee plus a monthly fee, and put a limit on questions per month (for example 1,000). Charge extra above the limit.
- Decide who pays the AI bill. One simple way is for the client to create their own AI account and give you the key. Then their usage is their own cost, and you carry no risk.
Do not charge only your costs. In the small business example, the costs are about $8 to $15 a month. That will not pay for your time.
Wrap-Up
A RAG chatbot is much cheaper to run than most people think. For a small business, expect around $8 to $15 a month. The cost grows with questions, so choose a cheap model, send fewer chunks, limit answers, and protect public bots.
If you would like a working codebase to start from, I made a RAG chatbot source-code kit: Source code: .
What does your RAG chatbot cost you today? Tell me in the comments, and I will try to cover your biggest question in the next part.
Where These Prices Come From
All checked on October 6, 2026:
Top comments (0)