Best of Agent Harnesses Catalogs 167 Frameworks, Tools, and Techniques
Summary
Best of Agent Harnesses is a curated, searchable catalog of 167 AI agent harnesses, orchestration frameworks, configurations, evaluation systems, infrastructure components, and related techniques. It defines a harness as the runtime around the model's tool-use loop: the layer that controls available tools, approvals, context, memory, execution environment, error recovery, and guardrails. The repository ranks projects by relevance to those concerns and by GitHub stars, with weekly refreshes; its tables also classify adoption surface, autonomy, recovery, and license status. The authors argue that better models make harness design more important because broader capabilities create more failure modes, while retry logic, validation, fallbacks, permissions, and crash recovery determine whether agents work in production. The page cites measurements claiming that changing only the harness moved GLM-5.2 from 23% to 52% pass@1 and Gemma 4 26B from 15% to 36% on SWE-bench Pro, while harness rankings had only a -0.05 rank correlation across models. It also reports a comparison on ARC-AGI-3 in which a bare frontier model scored about 30% and Prime Agent with Opus 5 scored 95.5%, while noting that the accompanying pages discuss what these claims do and do not establish. The repository includes AGENTS.md and Claude Code settings templates, a roughly 180-line Python harness template, setup playbooks, use-case guides, and category comparisons. Agents can consume the list through harnesses.json, llms.txt, or an MCP server exposing recommendation, comparison, search, filtering, and template operations. The project is published to PyPI and the official MCP registry, and also provides three agent skeletons for harness selection, codebase auditing, and weekly movement reporting.