runtape Uses Counterfactual Replay to Debug AI Agent Decisions
Summary
runtape is a local Python tool for investigating failed AI agent runs and keeping verified fixes from regressing. It records model calls, tool arguments, results, errors, latency, and the surrounding context in JSONL traces, then reruns only the selected model call while leaving the agent and its tools untouched. Its `why` command measures the original decision on the unchanged context, removes or replaces system prompts, messages, tool results, and other context pieces, and uses repeated trials with a corrected one-sided Fisher exact test to identify text whose removal changes the decision. The tool narrows confirmed causes to JSON items, paragraphs, or sentences and distinguishes headline causes from context that is merely required for the task. The project’s examples include an email assistant forwarding an invoice because of an HTML comment, a refund agent using an amount from another order, and an operations agent following an old runbook. In one reported llama3.2 3B case, removing the earlier order lookup changed the refund decision from 9 of 40 reruns to zero, with p = 0.001. A benchmark across support, email, operations, disk cleanup, and access-control scenarios planted harmful instructions in realistic documents; among 14 counted cases where the model’s harmful behavior remained sufficiently stable, the planted sentence was the headline cause in all cases and was isolated exactly in 12. `fix` tests an untrusted-content rule, an action guard, and removal of the source text on the exact failed context, while `--write-test` generates a pytest regression test. The tool supports OpenAI and Anthropic SDKs, LangChain, LangGraph, OpenAI-compatible local servers, and custom loops, but it is intended for individual failures rather than continuous monitoring. Its findings apply only to the tested model and context: routed APIs, rare decisions, replacement text, and interactions among three or more unrelated pieces can limit attribution.