Back to News
RSS feedwww.divert.cloud

Divert Reports Blocking 14 AI Models Before Customer-Edge Exploits

Summary

Divert’s field report describes a blind security test in which eight OpenAI models and six Anthropic models operated autonomously against the same simulated customer edge. The environment combined a real customer website, two intentionally vulnerable public web applications, and five Divert deception mazes covering web, SSH, FTP, API, DNS, and infrastructure services. The two applications contained OS command injection and server-side template injection flaws, each capable of root code execution and an internal pivot. Divert says every model interacted with the edge, received a unique threat identity, and triggered a blocking signal. Eight models later reached at least one customer-application RCE when enforcement was deliberately withheld for observation, while six exploited both vulnerable paths. Across those eight successful runs, the blocking signal preceded the first verified exploit by an average of 9 minutes 55 seconds; the average time to first RCE was 12 minutes 47 seconds. The models also solved 46 deception milestones, progressed through 16 model-maze pairings, and completed 11 mazes. Two runs ended in refusals after interaction. The test used identical hosts and clues, one blind run per model, no human steering, separate time limits for OpenAI and Anthropic, and server-side beacons with per-run canaries as RCE evidence. Divert emphasizes that enforcement was held back after detection, so the observed exploits represent activity allowed by the lab rather than attacks that bypassed an active block.