Cloudflare has improved how it handles and reports client-cancelled HTTP/3 requests across Free, Pro, Business, and Enterprise plans. Customers now get a clearer view of client behavior in Cloudflare analytics and, where available, logs.
Previously, Cloudflare did not always stop an HTTP/3 request when the client cancelled its request stream. Some cancellations were already recorded as 499, while others continued to the origin and showed the eventual upstream status.
Cloudflare now stops affected requests sooner, reducing unnecessary origin work, and records them as 499. Customers may notice more 499 status codes for HTTP/3 traffic. This reflects more consistent reporting of existing cancellations, not an increase in failed requests.
Customers who use 499 status codes in availability calculations should consider excluding them from server-side error rates because they represent requests cancelled by clients.
Markdown for Agents now converts HTML with an in-process streaming engine at the edge. It processes content as it arrives instead of buffering the HTML response and sending it to a separate conversion service. This reduces conversion overhead and memory use.
This release also changes the conversion limit and response headers:
Conversion supports up to 6 MiB (6,291,456 bytes) of decompressed HTML, increased from 2 MiB (2,097,152 bytes). The limit applies after decompression, not to the compressed response size.
Converted responses no longer generate the x-markdown-tokens or x-original-tokens headers. Clients that use these values need to calculate token counts themselves.
Content-Length is removed from converted responses rather than recalculated, because the Markdown body is streamed.
You can now view R2 bandwidth by the Cloudflare location that served each request in the Cloudflare UI. This helps you see which locations consume the most bandwidth with options to select a specific bucket and download (read) vs upload (write) bandwidth.
By default, the chart shows the top five locations by total bandwidth consumed during the selected time range. Use the Top 5 locations dropdown to select other locations, up to six at a time.
You can now use cf.appsec.request.failed_detections to control how your rules handle requests when a security detection reports a failure.
The field is an Array<String> of detection IDs that reports failures from content scanning, WAF attack score, attack signature detection, leaked credentials detection, and AI prompt detections for personally identifiable information (PII), prompt injection, custom topics, and unsafe topics.
The field does not alter the existing behavior of detections. Use it in rules to choose how to handle requests with reported failures.
When no failures are reported, the field returns []. You can use it on all plans, but your plan must still include the detections and rule features you want to use.
Supported rules:
Custom rules at the zone and account levels
Rate limiting rules at the zone and account levels
Request Header Transform Rules at the zone level
Match any reported failure:
len(cf.appsec.request.failed_detections) gt 0
Match a reported leaked credentials detection failure:
@cf/cloudflare/clef-omni is now available on Workers AI. Clef-omni is a decision model that takes audio (WAV or MP3) and video (MP4 or WebM) input alongside text and images. It joins Clef and Clef-flash in the Clef family of open-weight decision models. We also cut the price of Clef-flash, so it now costs less than Jev, and made Clef faster.
Clef-omni: one decision model for every modality
Previously, making a decision about a voice recording or a video meant chaining models together: transcribe the speech, split the audio and visual tracks, then pass the results to a text decision model. Clef-omni reads every modality directly in one request. A video's soundtrack is aligned with its frames, so the model can reason over what is seen and heard at the same time.
Clef-omni is built on a 30B-parameter mixture-of-experts (MoE) backbone with 3B active parameters. Like the rest of the Clef family, it does not generate text. It scores every allowed answer in a single pass, so decisions return quickly:
Text requests: about 20 ms
Image or audio inputs: under 100 ms
A 21-second video clip with sound: about 300 ms
Pass media as base64 data URLs in the images, audio, and videos fields:
const response = await env.AI.run("@cf/cloudflare/clef-omni", { model: "clef-omni", state: "Review the installation: a photo of the unit, an audio recording of it running, and a video of the fan.", images: ["data:image/png;base64,<base64-png>"], audio: ["data:audio/mpeg;base64,<base64-mp3>"], videos: ["data:video/mp4;base64,<base64-mp4>"], questions: { label_visible: { type: "noul", instructions: "Is the model and serial number label visible in the photo?", }, sounds_normal: { type: "noul", instructions: "Does the unit sound like it is running smoothly, without rattling or grinding?", }, fan_running: { type: "noul", instructions: "Is the fan running in the video?", }, },});
const response = await env.AI.run("@cf/cloudflare/clef-omni", { model: "clef-omni", state: "Review the installation: a photo of the unit, an audio recording of it running, and a video of the fan.", images: ["data:image/png;base64,<base64-png>"], audio: ["data:audio/mpeg;base64,<base64-mp3>"], videos: ["data:video/mp4;base64,<base64-mp4>"], questions: { label_visible: { type: "noul", instructions: "Is the model and serial number label visible in the photo?", }, sounds_normal: { type: "noul", instructions: "Does the unit sound like it is running smoothly, without rattling or grinding?", }, fan_running: { type: "noul", instructions: "Is the fan running in the video?", }, },});
Clef-omni scores highest of the Clef family on BANKING77, CLINC150+OOS, and Amazon ESCI:
Benchmark
Clef-omni
Clef
Clef-flash
Jev
BFCL (case exact)
98.2
98.47
98.76
95.75
BANKING77 (macro-F1)
94.8
94.20
90.93
79.74
CLINC150+OOS (macro-F1)
97.7
97.43
66.77
89.27
Amazon ESCI (macro-F1)
57.8
57.48
57.39
55.21
PhishNChips (accuracy)
73.2
79.60
75.05
62.55
Clef-flash is now cheaper
Clef-flash now costs $0.038 per million input tokens, down from $0.090, which makes it cheaper than Jev. To offer this price, the hosted Clef-flash context window is now 24K tokens, down from 64K. Based on usage data, only 0.24% of requests exceed 24K input tokens. If you need a larger context window, use Clef, which keeps its 64K context window.
The Clef-flash weights on Hugging Face are unchanged and support up to a 256K context window if you self-host.
All Clef models convert image inputs to input tokens, and Clef-omni does the same for audio and video. For details on how each input type is tokenized, refer to the Clef, Clef-flash, and Clef-omni model pages.
Clef is now faster
We optimized how Clef is served on Workers AI, so it now returns decisions up to 2x faster. The model weights are unchanged.
Input size
Before: median / p95 (ms)
Now: median / p95 (ms)
Median speedup
~800 tokens
262 / 438
152 / 351
1.7x
~3,400 tokens
616 / 777
305 / 531
2.0x
~16,000 tokens
2,721 / 3,250
1,635 / 1,805
1.7x
Part of this speedup comes from moving Clef to SGLang ↗︎. Clef support is coming to SGLang in version 0.5.22 (PR #42721 ↗︎). If you self-host Clef, launch commands are available in the Clef collection on Hugging Face ↗︎.
Get started
Clef-omni follows the same System One API as Clef and Clef-flash, and works with AI Gateway. To try it, change the model ID to @cf/cloudflare/clef-omni and set the model selector to clef-omni.
createBatch() now accepts an options object that creates up to 100 Workflow instances in one call. The result lists the created instances and explains why any others were not created. To use this form in local development and get its types from wrangler types, use Wrangler 4.148.0 or later.
To create instances that share the same options, pass count. Each instance receives a generated ID:
created contains the created instances. errors contains each entry that was not created, identified by its position in the input. IDs that already exist and IDs repeated within the batch are reported as errors instead of being skipped silently.
Passing an array to createBatch() is deprecated. Existing code that uses the array form continues to work.
Traffic to split tunnel excluded resources is no longer briefly blocked while the client is connecting or reconnecting. The client now keeps its learned split tunnel configuration across reconnects.
Support for routing non-RFC 1918 local IPv4 networks through the tunnel when unrestricted LAN inclusion is enabled by policy or MDM.
Faster connects and lower memory use. The hosts file is now read once and shared across the client’s DNS resolvers instead of being reloaded by each one.
Additional changes and improvements
Improved reauthentication reliability and fixed an issue where a reauthentication could force a new registration.
Improved client reaction to the current network lowering its MTU.
Improved DNS reliability on networks with lower MTUs by clamping the TCP maximum segment size (MSS) for DNS-over-HTTPS connections sent through the tunnel.
Individual DNS-over-HTTPS queries now time out instead of hanging when the upstream server stops responding.
Improved API reliability by retrying requests dropped when reusing pooled connections.
Added an MDM setting to prefer IPv4 when resolving hostnames in proxy mode. The setting is off by default.
Fixed the client reconnecting while Emergency Disconnect was active after switching organizations or re-registering.
Fixed the client being unable to connect after an upgrade when its stored registration credentials no longer matched its configuration.
Fixed the client service restarting unexpectedly when it was slow to respond, such as after waking from sleep.
Fixed Extra Logging failing to capture packets across all interfaces.
Fixed an issue that could prevent remote diagnostics from completing.
Fixed DNS connectivity checks failing on IPv6-only networks.
Fixed the client service exiting when its route-monitoring socket was closed after sleep or wake.
Fixed DNS enforcement checks making the client service unresponsive on systems with large routing tables.
Fixed slow captive portal checks causing the client service to become unresponsive or restart while connecting.
Fixed a race when switching tunnel protocols during key rotation that could prevent WireGuard from connecting.
Fixed the client continuing to report “No network” after a successful manual disconnect.
Fixed a client UI crash that could occur when the daemon connection was reset during an IPC request.
Fixed a startup crash when date formatting data for the system locale had not yet loaded.
Traffic to split tunnel excluded resources is no longer briefly blocked while the client is connecting or reconnecting. The client now keeps its learned split tunnel configuration across tunnel reconnections.
Support for routing non-RFC 1918 local IPv4 networks through the tunnel when unrestricted LAN inclusion is enabled by policy or MDM.
Improved connection reliability on devices with very large hosts files. The client now detects a large hosts file, allows more time for its initial DNS check, and shows a banner letting the user know that connecting may take longer.
Faster tunnel reconnections and lower memory use. The hosts file is now read once and shared across the client’s DNS resolvers instead of being reloaded by each one.
A service recovery mechanism, backed by a Windows scheduled task, now starts the client service on system unlock if it is not already running. This is enabled by default.
Additional changes and improvements
Improved reauthentication reliability and fixed an issue where a reauthentication could force a new registration.
Improved client reaction to the current network lowering its MTU.
Improved DNS reliability on networks with lower MTUs by clamping the TCP maximum segment size (MSS) for DNS-over-HTTPS connections sent through the tunnel.
Individual DNS-over-HTTPS queries now time out instead of hanging when the upstream server stops responding.
Improved API reliability by retrying requests dropped when reusing pooled connections.
Added an MDM setting to prefer IPv4 when resolving hostnames in proxy mode. The setting is off by default.
The client no longer requires the Windows WLAN AutoConfig service to be running.
Fixed the client reconnecting while Emergency Disconnect was active after switching organizations or re-registering.
Fixed the client being unable to connect after an upgrade when its stored registration credentials no longer matched its configuration.
Fixed the client service restarting unexpectedly when it was slow to respond, such as after waking from sleep.
Fixed the client service failing to restart after an unexpected termination.
Fixed the client UI getting stuck in a connecting state after sleep and wake even though the tunnel was connected.
Fixed slow captive portal checks causing the client service to become unresponsive or restart while connecting.
Fixed a race when switching tunnel protocols during key rotation that could prevent WireGuard from connecting.
Fixed the client continuing to report “No network” after a successful manual disconnect.
Fixed Digital Experience Monitoring (DEX) HTTP tests failing TLS validation.
Fixed latency spikes and traffic interruptions during TPM-backed API authentication when hardware-backed registration is enabled.
Fixed trailing whitespace in BIOS serial numbers causing serial-number and client-certificate device posture checks to fail.
Fixed the client UI crashing at startup when it could not write to the Windows registry.
Fixed a client UI crash that could occur when the daemon connection was reset during an IPC request.
Fixed a startup crash when date formatting data for the system locale had not yet loaded.
Known issues
A Windows DNS client regression may cause connectivity check failures on systems containing large hosts files. While this release includes a fix to mitigate this issue, users may still experience reduced DNS performance and connectivity check failures.
Traffic to split tunnel excluded resources is no longer briefly blocked while the client is connecting or reconnecting. The client now keeps its learned split tunnel configuration across reconnects.
Support for routing non-RFC 1918 local IPv4 networks through the tunnel when unrestricted LAN inclusion is enabled by policy or MDM.
Faster tunnel reconnections and lower memory use. The hosts file is now read once and shared across the client’s DNS resolvers instead of being reloaded by each one.
Added an MDM setting to prefer IPv4 when resolving hostnames in proxy mode. The setting is off by default.
Additional changes and improvements
Improved reauthentication reliability and fixed an issue where a reauthentication could force a new registration.
Improved client reaction to the current network lowering its MTU.
Improved DNS reliability on networks with lower MTUs by clamping the TCP maximum segment size (MSS) for DNS-over-HTTPS connections sent through the tunnel.
Individual DNS-over-HTTPS queries now time out instead of hanging when the upstream server stops responding.
Improved API reliability by retrying requests dropped when reusing pooled connections.
Fixed the client reconnecting while Emergency Disconnect was active after switching organizations or re-registering.
Fixed the client being unable to connect after an upgrade when its stored registration credentials no longer matched its configuration.
Fixed the client service restarting unexpectedly when it was slow to respond, such as after waking from sleep.
Fixed the client window not appearing on first launch after a fresh install on RHEL 10.
Fixed duplicate WARP routing policy rules accumulating on reconnect.
Fixed slow captive portal checks causing the client service to become unresponsive or restart while connecting.
Fixed a race when switching tunnel protocols during key rotation that could prevent WireGuard from connecting.
Fixed the client continuing to report “No network” after a successful manual disconnect.
Fixed a client UI crash that could occur when the daemon connection was reset during an IPC request.
Fixed a startup crash when date formatting data for the system locale had not yet loaded.
Known issues
When in DNS Only mode, the client may send DNS queries for names that are configured for Local Domain Fallback to the encrypted DNS server instead of falling back to the system configuration. Local Domain Fallback works as expected in other client modes.
Log Explorer datasets are now queried from the Logs page under Observability in the Cloudflare dashboard. The Logs page brings Log Explorer and Workers Observability datasets together with a shared filter builder, SQL editor, and visualizations.
As part of this change, the Log Explorer menu is no longer shown in the dashboard navigation.
Your enabled datasets, saved queries, and SQL queries continue to work on the Logs page.
To manage datasets, open the dataset selector on the Logs page and select Configure next to the Log Explorer datasets.
The previous Log Search page remains available at its direct URL.
Cloudflare Organizations is now generally available for Enterprise customers and MSSP/Distributor partners.
Organizations provides a top-level container for centrally managing accounts, members, analytics, and shared policies. Organization Super Administrators receive implicit access to every account in their Organization without requiring separate account memberships.
Enterprise customers can manage accounts in a single-tier Organization. MSSP/Distributor partners can use nested sub-organizations to manage customer accounts.
Organization Roles remains in beta, and current product limitations still apply.
AI Security for Apps now supports an updated set of categories for detecting unsafe topics in incoming prompts.
The values available in cf.llm.prompt.unsafe_topic_categories have changed. Existing WAF custom rules remain valid, but rules that reference a removed or renamed category will no longer match that category. Review any rules that use this field and update their expressions to use the currently supported values.
For category descriptions and configuration guidance, refer to Unsafe topics.
The DNS records page in the Cloudflare dashboard now shows a warning once you have used 85% of your DNS records quota.
The warning reflects the quota that applies to you. If your zone has its own quota, the warning shows that zone's usage. If your account uses an account-level DNS records quota, the warning shows your usage across all zones in the account.
This release introduces a new detection to mitigate a heap-based buffer overflow vulnerability in F5 BIG-IP, and enhances existing command injection protections by incorporating tested beta logic into the baseline rule.
Key Findings
CVE-2026-94127: A heap-based buffer overflow vulnerability in F5 BIG-IP. Attackers can exploit this flaw to execute arbitrary code on the affected system.
Ruleset
Rule ID
Legacy Rule ID
Description
Previous Action
New Action
Comments
Cloudflare Managed Ruleset
N/A
Command Injection - Generic 8 - uri - Beta
Log
Block
This rule is merged into the original rule "Command Injection - Generic 8 - uri" (ID: ).
Cloudflare CASB now supports custom finding types, giving security teams full control over the security conditions CASB detects across their SaaS and cloud integrations.
In addition to CASB's library of standard finding types, you can now write your own detection logic using Rego ↗︎, the open-source policy language from Open Policy Agent (OPA). Use custom finding types to match your organization's own thresholds and exceptions, such as flagging admin accounts without two-factor authentication, and get higher-confidence findings to act on.
Key capabilities
Write your own detection logic — Define exactly what CASB flags using Rego expressions evaluated against asset data from your connected integrations.
Target any supported provider and asset class — Scope a custom finding type to a provider (such as Google Workspace or Microsoft 365) and asset class (such as users, files, or groups), and apply it to all integrations for that provider or a selected subset.
Built-in validation — Select Validate to check your expression for syntax errors and schema issues before you create the finding type.
Inspect and duplicate standard finding types — Open any standard finding type to view its detection logic, then duplicate it as the starting point for a custom finding type.
Works with CASB policies — Use custom finding types in CASB remediation policies to send matching findings to Slack, ServiceNow, or any other webhook destination.
Get started
In Cloudflare One ↗︎, go to Cloud & SaaS findings > Findings library.
Select Create finding.
Enter a name, description, and severity.
Select a provider and asset class, then set the integration scope.
Write your Rego expression and select Validate.
Select Create finding.
CASB evaluates the custom finding type against assets as they are created or updated within the selected scope. Matching assets appear as posture finding instances under Posture Findings.
BGP peering over IPsec and GRE tunnels is generally available for Cloudflare WAN and Magic Transit. You can use it for production workloads.
BGP peering exchanges routes dynamically between your devices and your Cloudflare virtual network routing table. You no longer need to update static routes manually as your network changes.
BGP over IPsec and GRE tunnels is available to all accounts that use Unified Routing. No enablement is required. BGP over CNI remains in closed beta.
Web Search API is now available in beta. Web Search API lets your AI agents and applications search the Internet and ground their responses in live information, instead of guessing URLs or relying on a model's training cutoff.
At launch, you can choose between three search providers: Ceramic.ai, Exa, and Linkup. All three support Zero Data Retention for requests made through Cloudflare, and all have committed to Cloudflare's verified bot crawling standards.
Web Search API runs through AI Gateway, so search requests appear in your gateway logs and are billed to your AI Gateway credits at each provider's list API price, with no additional markup. You can also bring your own provider API key.
Call Web Search API with the REST API:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \ --request POST \ --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "query": "What are some fun things to do in Salt Lake City as fall approaches?", "provider": "ceramic", "limit": 5, "options": { "gateway": { "id": "default" } } }'
Or from a Worker with the AI binding:
const response = await env.AI.websearch({ gatewayId: "default", query: "What are some fun things to do in Salt Lake City as fall approaches?", provider: "exa", limit: 5,});const results = await response.json();
The strict service token authentication setting applies consistent behavior to requests made with service tokens. When the setting is on for a Zero Trust organization, Access handles requests with service token headers as follows:
If authentication or authorization fails, Access always returns 401 or 403 instead of redirecting the client to the login page with 302.
Only Service Auth policies can authorize the request. Access ignores Allow policies and any CF_Authorization cookie sent with the request.
Access does not return a CF_Authorization cookie to the client after successful authentication. Subsequent requests should continue to use service token headers.
Zero Trust organizations created on or after October 5, 2026 have strict service token authentication turned on by default and cannot turn it off. Cloudflare recommends that existing organizations turn it on as well.
Organizations created before October 5, 2026 can configure the setting in the dashboard or through the API.
The Agents SDK now provides first-class support for building agents using the Pi harness. You can build long-running agents using the combination of Pi 1.0 ↗︎, Pi Durable ↗︎, and the new PiHarness class that the Cloudflare Agents SDK provides, ensuring your agent's work is durably persisted, even if interrupted mid-turn.
Built with Earendil ↗︎, this integration is our first step toward first-class support for third-party agent harnesses on Cloudflare.
PiHarness is a new "Lifecycle capability" provided by the Cloudflare Agents SDK. Pi Durable provides the agent harness and the Lifecycle is responsible for keeping the agent running in the Durable Object. The Lifecycle is a core concept in the Agents SDK ensuring that long-running work can run in a Durable Object, surviving restarts, crashes, and network issues. We will share more on Lifecycle capabilities in the near future.
Install
npm i agents@latest @earendil-works/pi-durable @earendil-works/pi-ai
The agents/models/pi-ai entry point supports AI Gateway and Workers AI models, so you can get started with Cloudflare models right away or use your existing Pi AI provider.
Add tools with extensions
Both tools and system prompt sections are provided to the Pi Harness via extensions.
import { Type } from "@earendil-works/pi-ai";import { skills } from "agents/harness/pi";const WordCount = Type.Object({ text: Type.String() });const wordCount = { name: "word_count", description: "Count the words in a text.", parameters: WordCount, replay: "safe", async execute({ text }) { const words = text.split(/\s+/).filter(Boolean).length; return { content: [{ type: "text", text: String(words) }] }; },};// In the harness factory, before Harness.open():registry.install({ name: "editor", sections: [ { key: "preamble", render: () => "You are an editor.", tag: false }, ], tools: [wordCount],});registry.install(await skills(sources));
import { Type } from "@earendil-works/pi-ai";import type { ToolRegistration } from "@earendil-works/pi-durable";import { skills } from "agents/harness/pi";const WordCount = Type.Object({ text: Type.String() });const wordCount: ToolRegistration<typeof WordCount> = { name: "word_count", description: "Count the words in a text.", parameters: WordCount, replay: "safe", async execute({ text }) { const words = text.split(/\s+/).filter(Boolean).length; return { content: [{ type: "text", text: String(words) }] }; },};// In the harness factory, before Harness.open():registry.install({ name: "editor", sections: [ { key: "preamble", render: () => "You are an editor.", tag: false }, ], tools: [wordCount],});registry.install(await skills(sources));
For more information on creating and configuring extensions, refer to Extensions.
Every plan now gets at least 30 days of analytics data. Adaptive analytics datasets, such as HTTP requests, security events, and DNS analytics, retain at least 31 days of data for Free and Pro domains, and you can query up to 30 days in a single request. Previously, Free and Pro domains could see between 24 hours and 8 days of history depending on the dataset.
A full month of history lets you investigate an issue after it happens, compare today with the same day in previous weeks, and tell a one-time spike from a longer trend. The change applies in the Cloudflare dashboard, in Custom Dashboards, and through the GraphQL Analytics API.
Domain analytics also now live in one place. In the Cloudflare dashboard, select a domain and go to Analytics to see Traffic, Performance, Security, Cache, Origin, DNS, and Visitors as tabs that share one time range and one set of filters. Account-level analytics are under Observability > Analytics.
This change does not alter which datasets or fields your plan can access. Aggregated datasets, such as httpRequests1hGroups, keep their existing per-plan limits. To check the exact retention and query window for a zone or account, query the settings for each dataset.
You can now build Custom Dashboards charts from Workers Observability data. Two new datasets, Workers Observability — Logs and Workers Observability — Traces (OTel), let you chart Worker invocations, log levels, errors, CPU and wall time, span counts, and durations next to HTTP traffic, security events, and other analytics datasets.
This gives you one dashboard for an application that spans Cloudflare's network and your Workers. For example, you can put request volume, WAF blocks, and Worker error rates on the same view, filter all three by time range, and spot whether a spike in errors lines up with a change in traffic.
The datasets are available for every Worker in your account that has Workers Logs or Workers Traces turned on. Custom Dashboards also now allow up to 100 dashboards for every account.