DEV Community

Mikhail Savchenko
Mikhail Savchenko

Posted on Originally published at inite.ai

Cloudflare's Clef-omni adds audio, video decisions

Cloudflare expanded its open-weight Clef family of decision models — lightweight models built for classification and schema-constrained scoring rather than text generation — with three changes: a new multimodal model, a price cut, and a speed upgrade.

Clef-omni is the first model in the family to natively process audio (wav/mp3) and video (mp4/webm) alongside text and images in a single pipeline, built on a Qwen3-Omni-30B-A3B-Instruct foundation with the text-to-speech components stripped out. Cloudflare reports median response times of about 130ms for text-only input, 150ms for images, and roughly 1.5 seconds for a full 21-second video clip with sound. It launches at $0.15 per million input tokens, with weights published on Hugging Face.

Clef-flash pricing drops from $0.09 to $0.038 per million input tokens, which Cloudflare says makes it cheaper than TypeSafe's Jev model. The trade-off is a smaller hosted context window — cut from 64k to 24k tokens — though the underlying weights still support 256k tokens for anyone self-hosting. Cloudflare says only 0.24% of its observed requests exceeded 24k input tokens, which is why it made the cut rather than leave the window unchanged.

Clef itself keeps its $0.24 per million token price and 64k context window but gets serving-infrastructure optimizations — including a move to SGLang — that cut median latency by up to 2.0x depending on input size, with no changes to the model weights.

Cloudflare lists internal use cases including closing spam issues on its GitHub docs repo, moderating plugin libraries for phishing in its EmDash content management system, scanning for personally identifiable information, and detecting malicious domains. Clef remains API-compatible with Jev and is available through Cloudflare's AI Gateway.

Top comments (0)