We run a small content site on Cloudflare Workers, D1, and R2 — no origin server, no VM. It's been cheap and mostly boring, which is what you want from infrastructure. But "mostly" is doing work in that sentence. Here are six things that went wrong quietly enough that a green deploy log told us everything was fine while it wasn't.
None of this is a takedown of the platform. It's the list of assumptions we made that turned out to be wrong, in the order we tripped over them.
1. D1's statement limit doesn't fail loudly if your chunking is wrong
D1 caps a single SQL statement at 100,000 bytes. That's documented, and for short rows you'll never notice. We noticed the first time we tried to write a long article body in one UPDATE.
The obvious fix is to split the content into chunks and send them as separate statements that concatenate onto the column:
for (const chunk of chunks) {
await db.prepare(
`UPDATE posts SET content_html = content_html || ? WHERE id = ?`
).bind(chunk, id).run();
}
That works — until one chunk in the middle of a batch silently fails to apply for an unrelated reason (a flaky connection, a retried request that landed twice, whatever) and every other statement still returns success. The row is now missing a chunk. The write API told you 200 the whole way through. Nothing about the response shape tells you the final string is short.
The only way we've found to catch this is to re-measure after the fact:
const { results } = await db.prepare(
`SELECT length(content_html) AS len FROM posts WHERE id = ?`
).bind(id).all();
if (results[0].len !== expectedLength) {
throw new Error(`content truncated: got ${results[0].len}, expected ${expectedLength}`);
}
length() in SQLite counts characters, not bytes, so compare it against your source string's character count, not its byte size, or a batch of multi-byte text will fail the check for the wrong reason.
2. Something smaller than D1's limit can kill you first
Before we ever got near 100KB, a completely different ceiling showed up: on Windows, passing SQL inline to wrangler d1 execute --command breaks at something close to 8KB, with a plain "command line is too long" error that has nothing to do with D1.
# breaks around ~8KB, nowhere near D1's own limit
wrangler d1 execute my-db --remote --command "UPDATE posts SET content_html = '...' WHERE id = 42"
It's a shell-level limit, not a database one, and it's easy to misdiagnose as a D1 problem because the symptom (a write that refuses to go through) looks identical. The fix is to stop passing SQL on the command line at all:
printf '%s' "$SQL" > patch.sql
wrangler d1 execute my-db --remote --file=patch.sql
The lesson generalizes past Windows: whatever transport you use to get SQL into wrangler, check which layer's limit you're actually hitting before you start shrinking your data to fit a number that isn't the real constraint.
3. Writing that SQL file can quietly rewrite your data
Once we moved to --file, a second problem showed up that's easy to miss because it doesn't produce an error at all: opening a file in plain Windows text mode converts \n to \r\n on write — including inside a quoted string literal that's supposed to be HTML content, not a source-code line you'd want reformatted.
We measured it directly on one patch: a 28,326-character string came back as 28,637 characters after a round trip through a naive file write — one extra byte for every line break the content already had.
# quietly turns every \n inside the string into \r\n
with open("patch.sql", "w") as f:
f.write(sql)
# doesn't
import io
with io.open("patch.sql", "w", encoding="utf-8", newline="") as f:
f.write(sql)
The insert still "succeeds." The row still looks plausible if you eyeball it. The only way we catch this now is the same length re-check from gotcha #1, run against the file on disk before it ever reaches wrangler.
4. A token that verifies fine can still be the wrong token
wrangler d1 execute started failing with an authentication error on a token we knew was valid — it had a real expiration date years out and passed a plain token-verify check. The error was real, just misleading: the token existed and was active, it just had never been granted D1 permissions in the first place. It had been issued for a completely different job (DNS and Access rules).
Cloudflare's own permission model explains why this happens: API tokens are scoped per service — a token can hold, say, "Workers Scripts Read" without holding "D1 Write," and there's no single "is this token good" check that covers both. "Valid" and "authorized for this call" are different questions, and only one of them is answered by a token-verify endpoint.
The fix was mechanical once we understood it: keep a separate, narrowly-scoped token specifically for D1, under its own environment variable, instead of reusing one general-purpose token everywhere.
# an API token without D1 permission looks fine right up until this call
export CLOUDFLARE_API_TOKEN="$GENERAL_PURPOSE_TOKEN"
wrangler d1 execute my-db --remote --command "select 1"
# Authentication error [code: 10000]
5. A wall of 404s can mean you're checking the wrong host, not a broken upload
We serve images through R2 behind a Worker, and after a batch upload our verification script reported every single image as a 404 — including ones that had been live and working for months. Before assuming the upload pipeline was broken, we re-checked one already-known-good, already-live image with the exact same request. It also came back 404.
That's the tell: if a control you know is fine fails the same way as the thing you're actually testing, the bug is in the check, not the subject. In our case it was two things stacked on top of each other. First, we were hitting the public-facing custom domain's media path instead of the actual serving Worker's own subdomain, and the routing between the two isn't 1:1. Second, once we fixed the host, we still got failures — because the default user-agent on a plain HTTP client gets blocked by the Worker's own bot protection, independent of whether the file exists.
import urllib.request
req = urllib.request.Request(image_url, headers={
"User-Agent": "Mozilla/5.0 (verification script)",
"Referer": "https://my-blog.org/",
})
urllib.request.urlopen(req)
Run the control check first. It's one extra request and it tells you whether you're about to chase a bug that doesn't exist. The tool pages behind this particular pipeline live at my-blog.org/tools, if you want to see what didn't 404.
6. Cross-origin requests throw away your referrer path, not just strip it a little
We wanted to know which article page a newsletter signup came from, and for a while every single row in the database said the same thing: the site's homepage. That's not what was happening — nobody was subscribing exclusively from the homepage.
The frontend Worker and the API Worker sit on different origins, and a same-origin-looking site can still be split across two Worker subdomains under the hood. The browser's default referrer policy (strict-origin-when-cross-origin) only sends the full path when the request stays on the same origin; cross-origin, it sends the origin and nothing else. Every request looked like it came from /, because that's literally what the browser was willing to disclose, not because that's where the click happened.
The fix isn't a header trick — you can't get the path back out of Referer once the browser has decided to withhold it. You have to send the path yourself, explicitly, in the request body:
fetch(apiUrl, {
method: "POST",
body: JSON.stringify({
email,
source: location.pathname, // don't rely on Referer for this
}),
});
If your pages Worker and your API Worker are different origins — which is common the moment you put one behind a custom domain and hit the other directly — budget for this before you build any attribution on top of Referer.
None of these are exotic. Every one of them looked, at the moment it happened, like something else was broken — the database, the network, an expired credential, a missing file. What they had in common is that the actual failure was one layer away from where the symptom showed up, and a plain "it returned 200" or "it deployed" told us nothing about that layer.
If you're running D1 and R2 behind Workers, the site these lessons came from is my-blog.org/tools — a small pile of calculator pages that gave us most of this list one incident at a time.
Top comments (0)