Auditing Event-Selection Skill in Financial LLM Agents
Summary
Financial LLM agents are often judged by comparing their end-to-end returns with a baseline, but that test can mistake passive exposure for event-selection skill. An agent that changes from flat to long frequently may earn positive paired returns simply because the underlying event pool trends upward, even if it selects events at random. The proposed Agent Policy-Value Audit keeps the observed number of each ordered action-change type fixed and randomly reassigns those changes across eligible events. The average payoff from these reassigned actions is the composition benchmark; the gap between the agent's observed deployment value and that benchmark is selection value. In semi-synthetic tests using real earnings-event returns, a conventional zero-centered paired test falsely identified selection skill in 11.6% of no-skill replications, while the transition-matched audit reduced that rate to 5.3%. Applied to 723 earnings events at 44 U.S. consumer-facing firms, the audit decomposed gross deployment value of +15.2 basis points per event into a +25.8-basis-point composition benchmark and -10.6 basis points of selection value. The agent did not detectably outperform matched random assignments. The authors argue that financial-agent evaluations should report deployment value and event-selection value separately.