Back to News
RSS feedgithub.com

Project Arena Benchmarks AI Investigation Tools on Kubernetes Incidents

Summary

Project Arena is a vendor-neutral starter for testing AI investigation products against disposable Kubernetes incidents. It can create a local kind cluster or an AWS EKS environment, deploy either a 21-scenario full suite or a six-scenario smoke suite, inject and verify faults, preserve investigation records, and score them with a configurable AI judge. The benchmark keeps cluster setup, product integration, scenario execution, archiving, judging, and reporting as separate steps, and supports both direct Kubernetes deployment and an Argo CD GitOps path. The repository does not install a vendor agent or configure a hosted product; users must connect their own collector, alerting, or API exporter and save final, intermediate, and action records in the documented format. In the published comparison, Edge Delta’s native investigations detected 18 of 21 scenarios, while Grafana’s detected 12. For completed investigations, Claude reached root-cause scores of 18/21 with Edge Delta evidence and 19/21 with Grafana evidence, while supported mitigation scored 16/21 and 15/21 respectively. On the 12 incidents shared by both native products, Edge Delta scored 11/12 for root cause analysis and 7/12 for supported mitigation, compared with Grafana’s 9/12 and 5/12. A GPT-6-Astra judge evaluated final investigations against common incident facts and a rubric; the project notes that mitigation metrics assess proposals rather than executed repairs or verified recovery. Results can be exported as summary and detail CSV files, and the same saved investigation can be rescored with consistent judge settings. The repository requires Python 3.10+, while Kubernetes use additionally needs kubectl; local full-suite runs require Docker, kind, Helm, suitable images, and an enforcing CNI, and EKS runs require Terraform, AWS credentials, registry access, and architecture-compatible workers.