Back to News
RSS feedarxiv.org

JusticeAxis Benchmarks Legal Judgment Between Rigid Rules and Ungrounded Discretion

Summary

JusticeAxis frames legal judgment as a reference-anchored task: a decision must remain grounded in both the applicable statute and the circumstances of the case. The benchmark contains 256 real-world criminal cases from 18 countries, each paired with audio, image, and text evidence. For every case, lawyers wrote three judgments: the recorded judgment and two counterfactual judgments representing the benchmark’s two failure modes. The authors also introduce JusticeAgent, a harness in which element agents establish facts and a judge agent applies the law using experience-oriented skills. These skills are distilled from execution trajectories and admitted only when they meet Bayesian credible bounds. Experiments show that failure modes shift with model scale: open-weight backbones tend to introduce unsupported grounds, while frontier models tend to fall back on the statutory default. The authors report that JusticeAgent, used as a plugin, raises a frozen open-weight backbone to commercial-level performance. Project resources are available on GitHub.