Cloudflare Used AI to Attack Its Own WAF and Found 49 Gaps, Mostly SSRF
Cloudflare just published the results of an experiment that every API developer should read, even if you’ll never run a firewall. The company put frontier AI models inside a test harness and told them to get past its own Web Application Firewall.
They found gaps. Not many, but the pattern of what got through is the real lesson.
What Cloudflare did
The test started with attack payloads the WAF already blocked. Instead of replaying a fixed list, the models looked at how each attempt was handled and proposed a tweak for the next one: different encoding, different placement in the request, different delivery. They had no access to the WAF rules or source code. Pure black box.
The models never sent anything themselves. A Python harness built and replayed the HTTP requests, kept state, enforced limits and collected responses. One model call proposed the next mutation, a second reviewed the result.
Across 45 scenarios the system made 1,107 attempts. After human review, 558 were blocked and 49 were flagged as real findings. Forty-eight of those 49 were command injection or server-side request forgery (SSRF). The work led to two new SSRF detections in Cloudflare’s managed ruleset and an improved third one.
The SSRF trick worth understanding
One example from the write-up shows how this works. The model kept trying to reach a cloud metadata address, the internal IP that hands out credentials on most cloud VMs. It tried decimal form. Octal. Other encodings. The decimal version got blocked.
So it kept the same request shape and switched to a trailing-dot version of the address. That one didn’t hit the WAF block at all. It got a redirect.
Here’s why that matters to you. These are all the same address:
169.254.169.254
2852039166 (decimal)
0251.0376.0251.0376 (octal)
169.254.169.254. (trailing dot)
A filter that checks the string “169.254.169.254” catches one of these. An HTTP client resolves all four to the same place.
Lesson: don’t filter strings, check destinations
If your API ever fetches a URL a user gives you (webhooks, link previews, image imports, “import from URL” features) you have an SSRF surface. Most people defend it wrong, with a blocklist of hostnames or IP strings. Cloudflare’s result shows exactly how that fails.
Do it this way instead:
- Parse the URL with a real parser, not a regex
- Allow only
httpandhttpsschemes - Resolve the hostname yourself
- Check every resolved IP against private, loopback and link-local ranges
- Connect to the IP you checked, not the hostname, so DNS can’t change between check and use
- Turn off automatic redirects, or re-run the check on every hop
A rough Python version of the core check:
import ipaddress
import socket
from urllib.parse import urlparse
def is_safe_target(url: str) -> bool:
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or not parsed.hostname:
return False
try:
infos = socket.getaddrinfo(parsed.hostname, None)
except socket.gaierror:
return False
for info in infos:
ip = ipaddress.ip_address(info[4][0])
if ip.is_private or ip.is_loopback or ip.is_link_local or ip.is_reserved:
return False
return True
Step 6 is the one people skip. Notice that the trailing-dot request in Cloudflare’s test produced a redirect. If your client follows redirects blindly, an attacker can pass your check with a harmless URL and then bounce you to the metadata endpoint.
Lesson: a WAF is a backstop
Cloudflare runs one of the most heavily tuned WAFs on the internet, and a model with no inside knowledge still found 49 things worth fixing. That’s not a knock on Cloudflare. It’s what happens when you defend at the edge using pattern matching.
Your API has to be safe on its own. Validate inside the handler, at the point where you actually make the outbound call. Treat the WAF as the second lock.
Lesson: keep the model away from the trigger
Look at how the harness was built. The models suggested and judged. Plain code did the sending, the rate limits and the bookkeeping. Humans made the final call on every finding.
That split is a good pattern for any AI feature that touches your API. Let the model decide what to try. Let deterministic code decide what actually goes out.
Try it yourself
Find one place in a project where your code fetches a user-supplied URL. Feed it the four forms of the metadata address above, plus a URL on your own server that 302-redirects to http://127.0.0.1. If any of them get through, you know what to fix this week.