Redson Dev brief · PRIMARY SOURCE
Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Cloudflare Blog · October 9, 2026

The expansion of the Clef decision model family presents a significant opportunity to streamline how organizations process and derive insights from diverse data streams. This latest announcement from Cloudflare introduces Clef-omni, a new model capable of natively handling audio, video, images, and text within a unified processing pipeline. Simultaneously, the core Clef model has seen inference speeds improve by up to 2.0x, and the more economical Clef-flash model now benefits from reduced pricing, making advanced decision-making tools more accessible and efficient. For a logistics startup based in Atlanta, Georgia, optimizing delivery routes and warehouse operations, this could mean feeding real-time traffic camera footage, voice commands from dispatchers, scanned package labels, and written customer service queries into a single system. The Clef-omni model could then identify potential delays, prioritize urgent deliveries based on multimodal inputs, and automatically reroute drivers, saving both time and fuel. An independent e-commerce entrepreneur operating out of Los Angeles, California, selling custom apparel might use this to analyze product images for quality control, customer review text for sentiment, and even video submissions for custom design approvals, all while reducing the complexity of managing disparate AI services. Similarly, an internal IT team at a mid-sized healthcare provider in Boston, Massachusetts, could leverage these advancements to enhance patient intake by processing scanned insurance cards, dictated physician notes, and patient portal messages through one system to flag critical information or potential issues faster, improving administrative efficiency and patient care coordination. To begin exploring these capabilities, consider a specific, recurring data processing challenge within your current operations that involves at least two different data types—for instance, image and text, or audio and text. Spend a few hours this week sketching out a simple workflow that would feed these two data types into a theoretical unified pipeline. Don't worry about integration details yet, but rather identify what decisions or insights you'd want to extract and how a single, multimodal model could simplify the process compared to your existing or imagined multi-tool approach. This exercise will clarify the potential for efficiency gains and highlight specific use cases where such a capability could unlock new value.
Source / further reading
Learn more at Cloudflare Blog →