Dev News Daily ENDE

Cloudflare's security agents gather evidence in plain code first and must cite it

Cloudflare published on 7 October how its Managed Defense service uses AI agents to investigate security alerts — and, more usefully for anyone building agents, why its first design failed.

Why one agent failed. The first prototype gave a single general-purpose agent the whole investigation. It produced useful analysis but also claims the evidence did not support. Cloudflare names three recurring problems: context became authority (a detection is a hypothesis, not proof that an attack succeeded), scope drifted (an agent can query the wrong account or time range — "you can't rely on a language model prompt to be a boundary"), and failure disappeared (a timed-out lookup may not distinguish "not checked" from "checked and not found").

The fix: recon first, inference second. Before any model is called, deterministic code runs fixed, versioned reconnaissance — the customer's detection history, traffic baseline, enforcement outcome and network observations — and stores each item with its source, version and timestamp. The same snapshot can be replayed, so differences between agents come from interpretation, not retrieval.

Cloudflare's security agents gather evidence in plain code first and must cite it
Cloudflare's security agents gather evidence in plain code first and must cite it — Dev News Daily

Then the agents.

  • A lightweight model, Clef — Cloudflare's open-source decision model on Workers AI — scores how likely an alert is a false positive; known high-volume noise is classified as passive and stays out of the active queue.
  • For alerts that need more, a coordinator runs four specialist agents in parallel: traffic analysis, customer context, global telemetry (aggregates only, never another customer's records) and threat intelligence.
  • A synthesis agent combines their findings; it cannot fetch new evidence or choose a classification outside an approved list.
  • Specialists must cite items in a versioned evidence package, and application code checks that every citation exists and supports the claim.

The advisory explicitly separates "not checked" from "checked, with no matching result", and a human analyst confirms the scope. For deeper analysis Cloudflare says it uses approved models from OpenAI and Anthropic.