DEV Community

Cover image for Our MCP server was fine. Cloudflare was returning HTML.
Alexander Lukashov
Alexander Lukashov

Posted on Edited on Originally published at foxnose.net

Our MCP server was fine. Cloudflare was returning HTML.

This is what our MCP endpoint returned to a client speaking JSON-RPC:

HTTP/2 403
content-type: text/html; charset=UTF-8
server: cloudflare

Attention Required! | Cloudflare
Please enable cookies.
Sorry, you have been blocked
Enter fullscreen mode Exit fullscreen mode

A program was being asked to enable cookies.

We did not see that for three days, because that is not what the client reported. The client reported a parse failure, and a parse failure points at your own serialization, not at a machine four thousand kilometres away.

What it looked like from inside

The connector could not talk to our server. The tools never listed. Everything we could check on our side looked correct: valid JSON-RPC, right content type, right status codes, protocol revision we support.

So we did the reasonable thing and started fixing the things that were wrong but adjacent.

GET /{prefix}/_mcp was returning JSON. Per the Streamable HTTP spec it should return 405 if the server does not offer an SSE stream on GET. We changed it. It is a real fix and we kept it.

It changed nothing.

That is the value of shipping one variable at a time. If we had bundled that change with anything else, the next result would have been unreadable.

The thing that actually settled it

We wrote a minimal MCP server that imitated our own response shape exactly: application/json on POST, 405 on GET. A hundred lines, no auth, no database, nothing of ours in it. Then we exposed it through an ngrok tunnel and pointed the same connector at it.

It worked immediately. Tools listed, first try.

That one result eliminated most of the search space. Our response shape was fine, because a server with the same shape worked. The protocol was fine. The connector was fine. What was left was everything between the connector and our application, which is exactly the part we had not been looking at, because it is not in the repository.

If you take one thing from this: when you cannot find the bug in your code, build something that cannot possibly contain it and see whether the problem follows. The mimic server took an hour. It saved the rest of the week.

It was the user agent

Here is the same request, same body, same endpoint, varying only User-Agent:

user-agent                          code  content-type              bytes
Claude-User/1.0; +claude.ai          200  application/json            156
Claude-SearchBot/1.0; +claude.ai     200  application/json            156
anthropic-ai                         200  application/json            156
curl/8.7.1                           200  application/json            156
ClaudeBot/1.0; +claude.ai/bot        403  text/html; charset=UTF-8   4543
GPTBot/1.2; +openai.com/gptbot       403  text/html; charset=UTF-8   4543
Enter fullscreen mode Exit fullscreen mode

A Cloudflare managed rule, on by default, blocking AI crawlers. Nobody on our side had turned it on. Nobody had turned it off either, which is the point.

Look at the fourth row. curl passes. Every tool you reach for when something is broken is on the allowlist, which is why this survives so long. You test the endpoint, it answers, you conclude the endpoint is fine, and you go back to reading your own code.

One limitation worth stating: every row above was sent by curl, so the TLS fingerprint was held constant while only the header changed. That proves the user agent alone is enough to trigger the block. It does not prove the user agent is the only signal, and bot detection also reads the ClientHello. If you allow an agent by name and it still gets refused, that is the next variable to vary.

The taxonomy is where it goes wrong

Anthropic runs three crawlers on purpose, so that a site owner can make three separate decisions.

ClaudeBot collects content that may contribute to training. Claude-User fetches a page because a person just asked Claude a question. Claude-SearchBot indexes for search results. Three names, three robots.txt entries, three different trade-offs. The split exists precisely so you can refuse training and stay reachable.

Cloudflare has a category for exactly that middle case. It is called AI Assistant, and here is who is in it, as the crawler list shows it today:

ChatGPT-User            OpenAI        AI Assistant
MistralAI-User          Mistral       AI Assistant
Perplexity-User         Perplexity    AI Assistant
DuckAssistBot           DuckDuckGo    AI Assistant
Meta-ExternalFetcher    Meta          AI Assistant
Manus Bot               Manus         AI Assistant

Claude-User             Anthropic     AI Crawler
Enter fullscreen mode Exit fullscreen mode

Look at the naming. ChatGPT-User, MistralAI-User, Perplexity-User, Claude-User. One convention, one job, four vendors. Three of them are filed as assistants and one is filed as a crawler.

Meta gets it right twice over: Meta-ExternalAgent is an AI Crawler and Meta-ExternalFetcher is an AI Assistant, which is the same split Anthropic makes and the same split OpenAI makes.

So this is not a taxonomy that has failed to catch up with agents. The category exists, it is populated, and six vendors are in it correctly. One row is filed wrong.

The consequence lands on one vendor's users. Block the AI Crawler category, which is what the default does, and ChatGPT-User, Perplexity-User and the rest keep working. Claude-User does not. You made one policy decision and got a different outcome depending on which assistant your customer happens to use, without being told that is what you were choosing.

Fixing it on your side means allowing that agent by name, and first you have to work out that you need to. Ours now refuses the training crawlers and passes the user-initiated traffic, which took a deliberate change rather than anything the default did for us.

Fixing it properly is one row in Cloudflare's own table.

Nothing in the response says you were blocked

This is what makes it expensive rather than annoying.

There is no JSON error. No error code. No header explaining the refusal. cf-mitigated is absent. All you get that a machine can read is server: cloudflare and a cf-ray id, and neither of those means anything to a JSON-RPC client that expected an object and got a document.

So the client raises a parse error, and a parse error is a lie about where the problem is. It points inward, at your serializer, your framework, your content type. Every hypothesis it suggests is about code you own.

The part almost nobody uses: this is configurable. The same screen has a Configure Response control that sets the status code and message returned to blocked crawlers. If the thing behind your CDN is an API, a JSON body with an explicit reason costs nothing and turns three days of debugging into one line in a log. Blocking somebody is fine. Blocking them in a format they cannot parse is a choice you probably did not mean to make.

The raw evidence

Both sides of the same request, so you can see what a client has to work with.

Allowed, Claude-User:

HTTP/2 200
content-type: application/json

{"jsonrpc":"2.0","id":1,"error":{"code":-32002,
 "message":"Server not initialized",
 "data":{"hint":"Call initialize first and reuse Mcp-Session-Id header."}}}
Enter fullscreen mode Exit fullscreen mode

That is our server refusing the call, correctly, in the protocol, with a hint saying what to do next. 156 bytes.

Blocked, GPTBot:

HTTP/2 403
content-type: text/html; charset=UTF-8
server: cloudflare
cf-ray: a38529b8ec8bc239-BEG

Attention Required! | Cloudflare
Please enable cookies.
Sorry, you have been blocked
You are unable to access fxns.io
Enter fullscreen mode Exit fullscreen mode

4543 bytes of HTML. No cf-mitigated header, no JSON, nothing naming a rule. A client written against the MCP spec has no branch for this.

The same thing decides whether agents can discover you

Agent cards, OAuth protected-resource metadata, MCP discovery documents: all of it is converging on /.well-known/. That only works if the thing fetching it is allowed to fetch.

Ours is reachable, and here is how you tell:

/.well-known/oauth-protected-resource     404  application/json
/definitely-not-a-route                   404  application/json

body: {"message":"Route not found","error_code":"route_not_found",...}
Enter fullscreen mode Exit fullscreen mode

The 404 is ours. Same JSON envelope as any unknown route, which means the request reached the application. A block page instead of your own error format means it did not. Check whose 404 it is, not whether you got one.

The five minute version

Take your own endpoint and run the request you care about six times, changing only the user agent:

for ua in "curl/8.7.1" "Claude-User/1.0" "ClaudeBot/1.0" "GPTBot/1.2"; do
  curl -s -o /dev/null -w "$ua %{http_code} %{content_type}\n" -A "$ua" \
    -X POST https://your.endpoint/_mcp -H 'Content-Type: application/json' \
    -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
done
Enter fullscreen mode Exit fullscreen mode

Run it again in a month. The category a crawler sits in is Cloudflare's data, not yours, and it can be refiled without anything changing on your side, including the test you ran to prove your rule was safe. Check per user agent, not per category, and check more than once.

If the codes differ, or the content types do, the layer in front of you is making decisions you did not make. Watch the content type especially: a proxy can answer 200 with an HTML body, and no status check will ever catch that one. In Cloudflare they live under AI Crawl Control, in the Security section, one row per crawler. Then decide which of those decisions you actually want, because refusing training crawlers and refusing your customers' agents are not the same choice, and by default they are the same switch.


Updated after publishing: added the AI Assistant category table after a reader asked how the vendors line up, noted that the user agent matrix held the TLS fingerprint constant, and extended the check to content type after a reader pointed out that a proxy can answer 200 with an HTML body.

Top comments (5)

Collapse
 
raknaos profile image
Raknaos •

The user-agent matrix is the part most people miss. When curl works but an agent gets a block page, I now check the identity the client presents first: same request, different UA and TLS fingerprint, 200 vs 403. Proving the response is HTML instead of the JSON-RPC you expected is a good way to reframe it from 'my server is broken' to 'my infra is answering the wrong protocol'. Solid writeup, thanks for the debugging path.

Collapse
 
alexander_lukashov profile image
Alexander Lukashov •

TLS fingerprint is the one I didn't test, and it's a real hole in that matrix.

Every row was sent by curl. So I varied the UA string while holding the fingerprint constant at "curl". That proves the UA alone is enough to trigger the block. It does not prove it's the only signal, and a real ClaudeBot presents a different UA and a different ClientHello, which I never separated.

Matters for anyone following the recipe: if you allow Claude-User by name and traffic still gets refused, the next variable is the one I didn't vary.

Your reframing is better than the one in the post, btw. "My infra is answering the wrong protocol" says in six words what I spent two paragraphs getting to 😅

Collapse
 
moogh profile image
MOOGH •

This class of failure deserves more write-ups. The nasty part is that the client usually reports it as a protocol or validation error, so you go read your server code while the actual response body is an HTML interstitial. One thing worth adding to the post: the tell is almost always the content-type and the first byte of the body, not the status code. Logging the raw first ~200 bytes of every failed tool response has saved me more debugging time than any structured error handling — and it survives the case where the proxy returns 200 with an HTML body, which no status-code check will ever catch.

Collapse
 
alexander_lukashov profile image
Alexander Lukashov •

The 200-with-HTML case is the one I did not cover, and you are right that it is worse. My check script prints content type but the sentence under it says "if the codes differ", which is exactly the advice that misses it. Fixed, and credited.

The logging habit is better than what our own harness does, which stung a bit to notice. On a transport failure we record that it happened and throw the body away, so the trace shows mcp_transport_error and nothing else. If the block page had hit us during a scenario run instead of during manual poking, we would have had the error class and no idea it was HTML. Keeping the first couple of hundred bytes costs nothing and is the difference between a category and a cause.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.