DEV Community

Ashish Sharma
Ashish Sharma

Posted on AI-assisted

How I keep Cloudflare under $5 (and what quietly runs up your bill)

I run a few sites on Cloudflare's $5 Workers plan. The base plan covers a lot more than people expect. Most of the surprise bills I've read about come from a few products doing far more work than their owners realized. Here's where the money goes.

Turn on billing notifications and usage alerts first. A runaway loop should show up as an email, not on your card statement.

Workers: the base is generous

The Paid plan is $5/month and includes 10M requests and 30M CPU milliseconds. Beyond that, it's $0.30 per million requests and $0.02 per million CPU ms. Bandwidth isn't charged.

• Serve static files as static assets. Static assets are free and unlimited. Routing them through Worker code uses requests and CPU.
• Cache responses that don't change per user. On Workers Caching, a cache hit still counts as a request, but it skips CPU time, because the Worker doesn't run. (I haven't verified billing for the Cache API, so check that separately.)
• Block bots before they reach your code. WAF rate-limiting rules stop traffic before a Worker runs.

Queues: three operations per message

Each message costs three operations: one write, one read, one delete. The Paid plan includes 1M operations a month, then it's $0.40 per million.

One million messages a month is 3M operations, with 2M billable, which is about $0.80. Cheap, but it grows with volume, so don't queue work you can do inline.

• Batching doesn't lower the count. Operations are counted per message, so batching improves efficiency but not the bill.
• Retries cost extra. Each retry adds a read.

KV: writes are the expensive part

The Paid plan includes 10M reads and 1M writes a month. After that, reads are $0.50 per million and writes are $5.00 per million, so a write costs about ten times a read. Deletes and list operations are also billed at $5 per million beyond 1M included.

If every request updates a counter in KV, 5M writes a month costs about $20 in writes alone, even though reads look cheap.

Keep KV writes out of hot paths. Batch updates, or put counters in D1 or a Durable Object. Avoid listing keys in request paths.

R2: cheap storage, metered operations

The free tier covers 10 GB of storage, 1M Class A operations and 10M Class B operations a month. Beyond that, storage is $0.015 per GB-month, so 100 GB is about $1.50.

Class A (writes, listing, multipart uploads) is $4.50 per million. Class B (reads, HEAD) is $0.36 per million. Deletes are free. Egress is free for data served through Workers, the S3 API and r2.dev.

The expensive habit is calling HeadObject or ListObjects on every request. Store the metadata in D1 or KV instead.

D1: you pay for rows scanned

The Paid plan includes 25B rows read and 50M rows written a month. Beyond that, rows read cost $0.001 per million and rows written cost $1.00 per million.

You're billed for rows the query scans, not the rows it returns. A filter on an unindexed column can scan the whole table even if it returns three rows. Index the columns you filter on often. An indexed write adds one extra row written, which the docs say is typically offset by the reduction in rows read.

Durable Objects: idle is free, active isn't

The Paid plan includes 1M requests and 400,000 GB-seconds a month. Beyond that, requests are $0.15 per million and duration is $12.50 per million GB-seconds.

Duration is billed at 128 MB per object, whether or not it uses that memory. One always-on object uses about 336,000 GB-seconds a month, which fits inside the included amount. Two always-on objects go over by about 273,000, roughly $3.40.

• WebSockets without hibernation. Calling accept() bills for the whole connection. The WebSocket Hibernation API lets the object sleep between messages.
• Pending I/O. It can keep an object in memory for up to 15 minutes, and that time is billed.
• Alarms count as requests.
• Idle hibernatable objects aren't billed. The goal is to let them sleep.

Workers AI: cache the answers

Both Free and Paid include 10,000 Neurons a day. Beyond that, the Paid plan is $0.011 per 1,000 Neurons. Free users can't go past the daily allowance.

For @cf/meta/llama-3.2-1b-instruct, output tokens cost about 18,252 Neurons per million, roughly $0.20 per million output tokens. Input tokens cost about 2,457 Neurons per million. The free daily allowance covers around 550k output tokens a day on that model. Rates differ by model, so check the model's page.

The biggest waste is calling the model on the same input every time. Cache the results, batch embedding jobs, and pick the smallest model that gives acceptable output.

Vectorize: the query formula scales with index size

Vectorize bills queried dimensions and stored dimensions. The docs give this formula for queries:

((queried vectors + stored vectors) × dimensions × $0.01 / 1,000,000)

So each query's cost depends on how many vectors are in the index, not just how many results it returns. Stored dimensions cost about $0.05 per 100M. Queries are free if you don't run any.

Before you launch, work the formula out with your own vector count and dimensions, and test it. Keep indexes small, and filter with metadata where you can.

Browser Rendering: browsers are expensive

The Paid plan includes 10 browser hours a month, then $0.09 per additional hour. For Browser Sessions (Puppeteer, Playwright, CDP), each concurrent browser beyond 10 costs $2, averaged monthly.

Billing rounds the monthly total to the nearest hour, with 30 minutes or more rounding up. So 100 hours is 90 billable hours, about $8.10.

Use fetch when you can. Close sessions when you're done. Headless browsers are slow to run and costly to keep open, so only use one when the page needs JavaScript.

Stream: storage bills whether anyone watches or not

Stored minutes are $5 per 1,000, prepaid. Delivered minutes are $1 per 1,000. Ingress and encoding are free, and there's no separate bandwidth charge.

A video you never play still costs you storage. Delete originals you don't need.

Images: unique transformations

Images bills per unique transformation each month. The Paid plan includes 5,000, then $0.50 per 1,000. Repeat requests for the same transformation in the same month aren't billed again.

Don't generate sizes nobody uses. Calls to the binding's .info() method aren't billed.

A checklist

  1. Turn on billing alerts first.
  2. Serve static files as static assets.
  3. Cache what doesn't change per user.
  4. Rate-limit bots at the WAF.
  5. Count operations per request for Queues, KV and R2.
  6. Index the D1 columns you filter on.
  7. Keep KV writes out of hot paths.
  8. Let Durable Objects hibernate.
  9. Cache AI results and use the smallest model that works.
  10. Work out the Vectorize formula with your numbers before launch.
  11. Close browser sessions, and don't use a browser when fetch will do.
  12. Delete unused Stream videos, and avoid generating Images sizes nobody uses.

Most of these products have free allowances big enough for a small site. The bills get big when a single product does work you didn't count.

Top comments (0)