Sid Techno
HomePricingPortfolioContactSign in+1 (725) 465-8325WhatsAppGet Your Website
Sid Techno

Web & mobile app development agency. Websites, e-commerce stores, and custom software — built by the team behind DeployBase, TradeLeap, ClearAgent, and DeskLeap.

30 N Gould St, Ste NSheridan, WY 82801United States+1 (725) 465-8325

Stay Updated

Get product updates, new features, and hosting tips.

Products

  • DeployBase
  • TradeLeap
  • ClearAgent
  • DeskLeap

Services

  • CMS Development
  • Shopify Websites
  • React Native Apps
  • Laravel Development
  • Graphic Design
  • Wix Websites

Company

  • About Us
  • Pricing
  • Blog
  • Portfolio
  • Testimonials
  • Custom Solutions
  • Careers
  • Request a Quote
  • Contact
  • Client Portal
  • Sign In

Legal

  • Terms of Service
  • Privacy Policy
  • Refund Policy

© 2026 Sid Techno LLC. All rights reserved.

TermsPrivacyRefund
Home/Blog/Don't Trust the Model With the Evidence: How Cloudflare Fences In Its Security AI
October 8, 2026

Don't Trust the Model With the Evidence: How Cloudflare Fences In Its Security AI

Cloudflare's first AI security agent hallucinated. Its fix: gather the evidence in code before the model runs, stop the summary adding facts, and check every citation.

By Sid Techno

Written on 9 October 2026, about a Cloudflare post published on 7 October 2026.

In August we wrote about six "critical" SQLite vulnerabilities that turned out to be fabricated — a fix that didn't exist, a proof-of-concept that didn't even parse — and still reached the National Vulnerability Database, because nothing in the pipeline checked them. The lesson was not "AI is bad at security". It was that a convincing claim with nobody checking it travels a long way.

This week Cloudflare published the other half of that story: an engineering write-up of how it built AI agents for security operations so that the model is never the thing deciding what the facts are. It is worth reading even if you never touch a security product, because the pattern applies to almost anything you might build with AI.

First, the scope, in Cloudflare's own words: "The early beta is available in Cloudflare Managed Defense for eligible application-security alerts and cases." It is an early beta, inside one managed service, for eligible alerts only. The post gives no pricing. Everything below is about the design, not a product you can switch on.

They started where everyone starts — and it failed

The most useful thing in the post is that Cloudflare describes its own failed first attempt. One general-purpose AI agent was given the whole investigation. In their words: "It produced useful analysis, but it also hallucinated claims the evidence did not support."

They name three problems, and each one will be familiar to anyone who has put a language model near real data:

  • "Context became authority." An alert is a guess, not proof. As the post puts it, "A detection is a hypothesis, not proof that an exploit succeeded or an attack occurred." A single agent blurred the two.
  • "Scope drifted." An agent can look up the wrong account or the wrong time range. The post's verdict: "You can't rely on a language model prompt to be a boundary."
  • "Failure disappeared." This is the subtle one: "If a lookup times out, the result may not distinguish 'not checked' from 'checked and not found.'"

Their fix, in one sentence: "we moved evidence collection and scope enforcement into application code, before model analysis begins."

The design: the model never fetches its own facts

Here is how the system is laid out, step by step.

1. Ordinary code collects the evidence first. "Before we even call inference, deterministic code runs a fixed set of reconnaissance workflows with versioned API calls." Every piece is "stored with its source, version, and timestamp." Because the inputs are a fixed snapshot, a run can be replayed — so, in Cloudflare's words, differences between the agents' findings "come from interpretation rather than retrieval."

2. A small model filters out the noise. Most alerts are not incidents. A lightweight triage step compares each alert with its history, and alerts very likely to be false positives skip the deeper analysis.

3. Narrow specialists, in parallel. For alerts that need a closer look, "a coordinator AI agent runs four specialist AI agents in parallel" — one each for traffic, customer history, internet-wide signals and threat intelligence. Each sees only its own slice.

4. The summariser cannot go looking for new facts. "A synthesis AI agent combines their typed findings into one advisory; it can't fetch new evidence or choose a classification outside the approved vocabulary." The step that writes the conclusion is exactly the step that is not allowed to add evidence.

5. Code checks every citation. The specialists must cite items from a versioned evidence package, and "Application code checks that every citation exists, belongs to the investigation, and supports the attached claim. Invalid findings are corrected or recorded as limitations."

And around all of it, the boundary is enforced by code, not by asking nicely: "The model never receives authority to cross tenant boundaries or act for the Managed Defense Analysts."

The idea worth stealing: three answers, not two

If you take one thing from the post, take this. When a lookup fails, most systems quietly report "nothing found". Cloudflare's advisory instead distinguishes three states:

  • Not checked
  • Checked, with no matching result
  • Checked, with evidence supporting absence

Those look similar and mean completely different things. "We didn't look" and "we looked and it's clean" lead to opposite decisions, and a system that merges them will one day tell you everything is fine when it simply never checked. The post is explicit about what happens when the evidence runs out: "When the evidence is insufficient, it makes no classification or disposition recommendation."

What this does not claim

It would be easy to read this as "Cloudflare solved AI hallucination". It doesn't say that, and the design doesn't do that. The models can still be wrong. What the design does is narrow where a wrong answer can enter and make it checkable when it does: the facts are gathered before the model runs, the summary cannot add new ones, and every citation is verified by ordinary code. And the final call stays with a person — "the Managed Defense Analyst remains responsible for the decision and any mitigation."

That is a smaller claim than "AI you can trust". It is also a far more believable one.

What it means if you are building with AI

You don't need a security operations centre to use these ideas. If you are adding AI to a business workflow — summarising enquiries, drafting replies, classifying documents — the same questions apply:

  • Where do the facts come from? If the model looks things up for itself, you can't replay or audit what it saw. Gather the inputs in ordinary code first.
  • Can the final step add new claims? The step that writes the answer should be the one least able to invent.
  • Is every claim checked by something other than the model? If the AI says "according to the invoice", something should confirm that invoice exists.
  • Does "nothing found" mean you looked? Record "not checked" separately, or a silent failure will look like a clean result.
  • Who makes the decision? Keep a person on anything that matters.

If you are planning to put AI into a process and want a second opinion on where it should — and should not — be allowed to act, get in touch.

Sources

  • Cloudflare, "Building an evidence-grounded agentic security operations harness on Cloudflare" — blog.cloudflare.com/agentic-security-operations/, published 7 October 2026 (read 9 October 2026)
  • Sid Techno, "AI Found 1,072 Real Bugs. Then It Invented 6 Fake Ones." — published 3 August 2026
#sidtechno#cybersecurity
All Articles