Skip to content
Agents tracked: 258 Downloads (7d): 219M up 6.1% GitHub stars: 5.5M VS Code installs: 148M Releases (7d): 293 Agent status: 2 with issues Updated Oct 7, 2026
Case study Cloudflare

How Cloudflare built a multi-agent security operations harness for Managed Defense

Cloudflare describes an internal harness for its Managed Defense analysts that gathers evidence with deterministic code, filters noisy alerts with a triage model, and runs specialist AI agents in parallel to produce evidence-backed advisories. Analysts keep final responsibility for decisions, and the feature is in early beta.

The problem

Security alerts often arrive in bursts, and analysts must collect data, decide which alerts are related and account for missing sources while new ones keep coming. A first prototype that gave one general-purpose agent the whole investigation produced claims the evidence did not support, queried out-of-scope data, and hid lookup failures.

How they did it

  1. Run a fixed set of versioned reconnaissance workflows in application code before any model call, storing each datum with its source, version and timestamp.
  2. Score each alert against that recon data with Clef on Workers AI, and classify known high-volume noise as passive so it skips the active queue.
  3. Have a coordinator run four specialist agents in parallel: traffic analysis, customer context, global telemetry (aggregates only) and threat intelligence.
  4. Build a versioned evidence package, require specialists to cite items from it, and validate citations in application code.
  5. Let a synthesis agent combine the typed findings, with Clef picking from a reduced list of classifications, then generate an advisory report for the analyst.
  6. Orchestrate the stages with Workflows so a failed stage reuses already validated work, and record gaps as not checked, checked with no match, or checked with evidence of absence.

Results

  • As reported by Cloudflare, the harness cuts the time analysts spend assembling and analyzing alerts, though the article gives no figures.
  • Known high-volume noise is deterministically classified as passive so analysts are not paged repeatedly for it.
  • The advisory shows related alerts, admitted evidence, visible gaps and recommended next steps, such as rate limiting or WAF rules.
  • The early beta is available for eligible application-security alerts and cases in Cloudflare Managed Defense.

As reported by the source (Cloudflare Blog engineering post); AgentGid did not measure these figures.

Takeaway. Move evidence collection and scope enforcement into deterministic code, then give narrow specialist agents only validated evidence they must cite, rather than relying on one general agent's prompt.
AgentGid's take

This suits security teams with in-house engineers who can write versioned recon workflows and citation validators, and who already run on Cloudflare's stack (Workers AI, Workflows, Durable Objects, D1). Watch the evidence: the time savings are Cloudflare's own claim with no figures, and the feature is an early beta limited to eligible Managed Defense alerts. AgentGid has no listed alternatives for this case.

Similar use cases