Cloudflare tested its WAF with AI attackers, and a trailing dot slipped past
Cloudflare has published how it tested its own web application firewall with frontier AI models acting as attackers. The motivation, the company writes, is that what language models are good at is iterating on attack payloads faster than a human: changing the encoding, moving the payload to another part of the request, or switching to the next vulnerability class based on each response.
The tester, written in Python, starts from exploits the WAF already blocks and asks a model to propose the next variation; a second model call reviews the response. Neither call sees the WAF's rules, rule IDs or scores. Models never send requests themselves: code checks the target against an allowlist, disables redirects, logs every attempt and enforces a limit, and response text is treated as untrusted input. A request that was not blocked became a lead for human review, not a confirmed exploit.
Against an authorized customer staging environment, the run covered 45 scenarios, 44 of them across six categories: cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal and Log4j. It recorded 1,107 attempts. After removing malformed, benign, duplicate and out-of-scope results, Cloudflare says the vast majority were blocked, and the requests that got through were turned into new detections for all customers.
The post walks through one SSRF session. The cloud metadata address 169.254.169.254 was blocked as a dotted address, as the decimal integer 2852039166 and in octal. At attempt 18 the model kept the request identical to a blocked one and only added a trailing dot to the host; the client got a redirect instead of a block page. There was no evidence the application fetched metadata, but the pair of requests gave engineers a precise question to fix.

Why it matters
The method matters more than the score: using a model as a tireless mutator, with code keeping it on a leash, is something any team can run against its own staging stack. Cloudflare's own conclusion is the unglamorous one: a WAF bypass still needs a vulnerable application, so patching remains the stronger defence.