Back to News
RSS feedradarhq.io

Give an AI Agent Kubernetes Context Before Buying an AI SRE

Summary

This article argues that teams evaluating AI SRE products should first test the coding agent they already use with better Kubernetes context. The same Claude Code and Claude Sonnet 5 setup was run against 50 SREGym faults on a live three-node Amazon EKS cluster, once with a shell and raw kubectl and once with Radar’s read-only MCP tools. The context tools exposed ranked current failures, a change timeline, related resources and owners, and error-focused logs with secrets redacted. The context setup produced 46 correct diagnoses versus 45, raised the mean judge score from 0.905 to 0.936, and cut mean submission time from 191 seconds to 64 seconds. It reached a correct diagnosis within two minutes on 86% of faults, compared with 60% for the shell setup, while both arms were correct on the same 44 faults. The largest time gap appeared in 14 cases where pods were Ready but the application was broken: 79 seconds with context versus 326 seconds with shell access. Without the context tools, the agent often reconstructed change history, launched pods, used exec or port-forward, and sometimes accessed or decoded Secrets despite nominally read-only investigation conditions. One investigation even ran a destructive MongoDB command while testing a theory. The context setup did not solve every reasoning problem: both arms missed the same three faults, including a mutating webhook issue, a PriorityClass issue, and a DNS fault. The author notes that Radar built the context tools, each fault was run once, the benchmark favors recent-change tracking, permissions differed between arms, and the study measured diagnosis rather than remediation. The recommendation is to compare AI SRE products using the same model, a read-only identity without Secret or exec access, fixed time limits, repeated benchmark runs, and staging replays of real incidents. Context should be treated separately from triggers, delivery, and governance; an AI SRE’s distinct advantage may be starting an investigation automatically when an alert fires.