Praxa: An Evidence-Bound Harness for Governed AI Agent Execution
Summary
Praxa is presented as an agent harness for separating and testing claims that are often conflated in AI-agent systems: what an agent proposes, what it is authorized to do, what is dispatched, what external effect is verified, and what is promoted into service. Its architecture uses deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion to make those transitions explicit. The paper reports four evidence lanes. In an author-run repository-local audit at a pinned revision, 1,027 of 1,027 unit tests and 89 of 89 Workerd tests passed; all 363 expected source files were instrumented and four coverage floors were met, but raw per-test transcripts and independent reproduction are unavailable. In a provider-backed Terminal-Bench Core 0.1.1 pilot over 12 curated tasks, both the baseline and reliability-layer arms passed 17 of 36 strict trials. The reliability layer used 37.49% more input tokens and 50.73% more output tokens, so the pilot does not establish superiority. In a post-debug, two-order coordination-proxy comparison, the baseline and a source-authored candidate each completed 180 of 180 trials with equal measured accuracy, full hermetic crash recovery, and zero protected violations. The candidate used 37.11% fewer tokens, reduced estimated endpoint cost by 33.84%, and required 11.63% fewer steps, but these results do not establish better quality, latency, or production behavior. Deployed source and configuration evidence also shows bounded reflection, recall accounting, memory compilation, and tool-health paths, without a reported production outcome lift. The supported contribution is therefore an evidence-bound architecture for making authority-to-effect transitions testable. The evidence does not establish adversarial security, production safety, broad specialist superiority, autonomous recursive optimization, or user benefit.