Back to News
RSS feedgithub.com

Humanbound Open-Sources an Adversarial Testing Engine for AI Agents

Summary

Humanbound is an open-source adversarial testing engine, SDK, and CLI designed to test AI agents rather than only individual prompts. It drives multi-turn conversations against a real agent endpoint, probes tool use and scope boundaries, and scores behavior against a stated security policy. A scope file can describe the agent’s business purpose, permitted actions, and restricted actions, allowing generic jailbreak tests to become targeted tool-abuse tests. Failed findings can be exported as guardrail or firewall rules, and the project also supports training a Tier 2 classifier with its firewall component. The tool can run locally with providers including OpenAI, Anthropic, Gemini, or Ollama, while Ollama enables a full air-gapped workflow with no external API calls. Users can invoke the same implementation through the CLI or Python SDK, inspect posture scores and conversation logs, and generate HTML reports. An Arena command runs intentionally vulnerable agents in isolated Docker environments for practice, although the project describes that isolation as best effort. The repository is licensed under Apache-2.0, includes an MCP server path through its runner architecture, and sends anonymous CLI telemetry that users can disable.