The paper introduces Proxy Confidence, a training-free method for auditing tool calls made by black-box LLM agents whose token probabilities are hidden. A low-cost open-weight surrogate reads the same context, tool schema, and proposed action, then scores the call using several log-probability readouts: teacher forcing and request-PMI assess individual argument values, a discriminative verdict evaluates the call holistically, and tool-choice competition compares the selected function with alternatives. The authors argue that generative likelihood is useful for locating incorrect arguments, while a verdict is better at detecting calls that are wrong as a whole; when the error type is unknown, an ensemble is the low-regret choice. The system requires no access to the acting model's internals and adds one prefill pass alongside the tool call. On difficult coding tasks, it reaches an AUROC of 0.825, compared with 0.598 for the actor's stated confidence, and generative readouts outperform that baseline by 0.07 to 0.28 across three additional actors. It also beats self-consistency by 0.14 to 0.19 for near-deterministic actors at 1/K of the sampling cost. In a real-time gate, escalating the least-trustworthy calls for review improves accepted-action accuracy by 0.05 to 0.30 at 50% coverage. Returning the score with the tool result also improves task success on live-execution benchmarks by 0.119 and 0.137, with p <= 1e-4, and beats a random-value control by 0.078 when step errors are silent, with p = 0.003.
AI News
The latest AI releases, research, products, and industry updates.
Loading...