DEV Community

Cover image for Cloudflare Markdown for Agents now streams converted responses
Dave Kurian
Dave Kurian

Posted on Originally published at otf-kit.dev

Cloudflare Markdown for Agents now streams converted responses

A page fetch can look successful while your agent pipeline quietly depends on details that are no longer in the response. Cloudflare’s 9 October change to Markdown for Agents replaces buffered conversion with an in-process streaming path. Alongside that change, converted responses lose three headers that clients sometimes use for accounting and response handling.

If your application asks Cloudflare for Markdown by sending Accept: text/markdown, update the consumer before relying on the new behavior. Read the body as a stream, stop treating token-count headers as authoritative, and make size handling explicit. Cloudflare also raised the conversion ceiling from 2 MiB to 6 MiB of decompressed HTML. That is an input limit at conversion time, not a promise that the returned Markdown body is 6 MiB or smaller.

What changed in the conversion path

Markdown for Agents is Cloudflare’s content-negotiation feature for enabled zones. A client requests a page with an Accept header that includes text/markdown; Cloudflare retrieves the origin HTML and converts it when possible. The 9 October changelog says conversion now runs in-process at the edge and streams content as it arrives instead of buffering the HTML response and sending it to a separate conversion service.

For an agent application, the practical change is about the contract at the response boundary. Code that calls await response.text() still obtains a complete string, but it waits for the body to finish and accumulates it in memory. That may be acceptable for a small document processed as one unit. If your service handles many responses concurrently, forwards output as it arrives, or wants a bounded memory profile, consume the response body incrementally.

Streaming the body does not mean every downstream step can also stream. A model request may need the whole document, a retrieval index may need chunk boundaries, and a summarizer may need a clean section map. Choose where you buffer deliberately: for example, stream from the network into a size-limited temporary store, then parse or chunk from that store. Avoid accidentally buffering once in the HTTP client, again in a string conversion, and a third time in a queue payload.

Byte, Dex, Luna, and Nova track converted Markdown chunks as they move through an agent pipeline

Remove assumptions about response headers

Converted responses no longer include x-markdown-tokens or x-original-tokens. If your cost dashboard or routing logic reads either value, those calculations need another source. Count tokens with the tokenizer used by the model or service that will process the text. Keep the estimate next to the exact text version you submit, because conversion or cleanup can change the token count.

The converted response also omits Content-Length; Cloudflare says it is removed rather than recalculated because the Markdown body is streamed. Do not use a missing length as evidence that the response is empty, and do not wait for that header to size a progress bar, reserve exact storage, or decide whether a request has completed. Completion comes from the stream ending. If your UI needs progress, show an indeterminate state or measure bytes received locally.

The docs say that Vary includes Accept, preserving any dimensions already declared by the origin. That matters if a cache sits between your consumer and the page: HTML and Markdown are distinct representations. Let the cache respect the response’s variation metadata rather than building a cache key that ignores the requested content type.

Read the response incrementally

A minimal client should request Markdown, inspect the response status, and consume the body in chunks. The docs show converted responses with a Markdown content type, but do not define it as a contract in their description; if your application checks that header, treat it as a defensive validation rather than the signal that streaming has completed.

In a JavaScript runtime with Web Streams, the core loop can look like this:

const response = await fetch(url, {
  headers: { Accept: "text/markdown" },
});

if (!response.ok) {
  throw new Error(`Markdown request failed: ${response.status}`);
}

if (!response.body) {
  throw new Error("Response body is unavailable");
}

const reader = response.body.getReader();
const decoder = new TextDecoder();
let receivedBytes = 0;
const maxMarkdownBytes = 4 * 1024 * 1024;
const parts = [];

try {
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    receivedBytes += value.byteLength;
    if (receivedBytes > maxMarkdownBytes) {
      await reader.cancel("Local Markdown output limit exceeded");
      throw new Error("Markdown output exceeded the local byte budget");
    }

    parts.push(decoder.decode(value, { stream: true }));
  }

  parts.push(decoder.decode());
} finally {
  reader.releaseLock();
}

const markdown = parts.join("");
Enter fullscreen mode Exit fullscreen mode

The example’s 4 MiB output budget is an application choice, not a Cloudflare limit. Set it based on the memory and model-input budgets of your own service. For large documents, do not collect every string fragment into an array: send chunks to a bounded file, object store, or parser that can apply backpressure. The important properties are that you count bytes yourself, handle stream errors and cancellation, and finish decoding any partial UTF-8 sequence after the final chunk.

The response can also contain useful metadata in YAML frontmatter when the source page has supported meta tags, and JSON-LD may be preserved at the end. Keep those structures in mind when parsing; avoid assuming that the Markdown body begins at byte zero or that every page has the same sections.

Treat the 6 MiB cap as an input boundary

The documented conversion ceiling is 6 MiB (6,291,456 bytes) of decompressed HTML, up from 2 MiB (2,097,152 bytes). The limit applies after decompression, not to compressed transfer size. A compressed origin response can expand substantially, so checking only compressed bytes does not establish that it is within the conversion limit. The changelog does not specify the exact response behavior for a page beyond that ceiling. Do not build retry logic around an assumed status code, truncated body, or automatic fallback.

If you control the origin-fetch path, enforce a decompressed-input budget before conversion or parsing and define what your application does when it is exceeded. If Cloudflare is the component doing conversion and you cannot inspect the original HTML, record that distinction: your local Markdown-output limit is a separate safeguard, not a measurement of Cloudflare’s decompressed-HTML input. For a source that exceeds the documented conversion size, choose a fallback you can test, such as fetching a smaller page, using a site-provided text endpoint, or returning a clear unsupported-document result.

Test that boundary with a representative large page before routing production traffic through it. Verify that your consumer handles whatever response Cloudflare returns for that case, since the public documentation states the limit but does not promise a particular over-limit response. Keep the test separate from ordinary network timeouts and invalid content types so failures remain diagnosable.

Dex, Luna, Byte, and Nova separate an oversized HTML page stack from documents ready for conversion

Choose a conversion path for each source

Markdown for Agents is designed for HTML pages from zones where the feature is enabled. Cloudflare’s documentation lists other options for cases where Markdown for Agents is unavailable or where an application needs different conversion behavior: Workers AI AI.toMarkdown() supports multiple document types and summarization, while the Browser Run /markdown endpoint can render a dynamic page in a browser before converting it. Check those options against the source type and rendering needs in your own pipeline.

Before enabling the request path broadly, check three pieces of your pipeline: whether the requested zone has the feature enabled, whether your consumer handles streamed bodies and missing headers, and whether your own input/output budgets have a defined fallback. Then canary a small set of known pages and compare local token estimates with the token counts recorded for your actual model requests. Cloudflare’s changelog reports reduced conversion overhead and memory use, but it gives no measured end-to-end application result; measure your own traffic before attributing any change in latency or cost to the feature.

If your concern is protecting an AI endpoint when security inspection itself fails, see our separate guide on choosing a response for failed WAF detections. That is a policy decision at the request boundary; this post is about consuming converted page content.

What to change before relying on it

Search your client and middleware for x-markdown-tokens, x-original-tokens, and Content-Length reads tied to Markdown responses. Replace token headers with a local estimate, remove exact-length assumptions, and make the stream reader own completion and cancellation. Add a separate output cap that matches your service budget.

Finally, identify where the original decompressed HTML can be measured. If your code never sees it, do not claim that your Markdown-body counter enforces Cloudflare’s HTML conversion ceiling. Instead, document the boundary as upstream behavior and test the fallback your system can control. This keeps the new streaming response useful without turning an undocumented over-limit outcome into a hidden production dependency.

Sources

Top comments (0)