Anthropic Proposes Metrics to Track the Pace of Frontier AI Development
Summary
Anthropic proposes a public measurement framework for making the pace of frontier AI development more visible to governments, third parties, and the public. The framework covers three areas: how much AI contributes to building future AI systems, whether oversight keeps pace with increasingly autonomous agents, and how developers allocate compute among capability and safety work. In Anthropic’s snapshot from August 2026, Claude led 26% of the company’s AI R&D work, while more than 90% of the work was performed at or above the “AI collaborates” level; no measured category was fully autonomous. Anthropic also reported about 30,000 research and engineering agents active on its most-used internal platform. Online monitors covered all reported agent actions before execution and blocked about 0.002% of more than one billion analyzed decisions, while offline monitoring flagged roughly one to two transcripts per thousand for further review. During a week in July 2026, about 6% of AI R&D compute was allocated to safety work, rising to about 12% for compute used in AI-driven AI R&D. The company says these figures are snapshots rather than trends, and that classifications rely partly on best-effort labels and model-based judgments. It also notes that safety work is difficult to distinguish from capability work and that compute share is an imperfect proxy for research effort. Anthropic argues that common methodologies, independent verification, and recurring disclosure could make comparisons across labs more credible. It says it plans to give independent third-party evaluators access to internal processes, systems, and data to verify safety practices, report incidents, and monitor these metrics.