Enterprise structured-data agents must reason across schemas, relationships, policies, and recurring business roles, but repeatedly reconstructing that structure for each query is costly. The paper introduces latent equivalence learning to divide this work between semantic agents and learned systems that predict recurring structure. Its framework separates persistent, task-relevant identities from the ways those identities appear in a particular dataset. Supporting and opposing evidence are used to form support-realized Gaussian prototypes, while soft-membership profiles preserve distinctions that hard assignments would lose. A separate learned query-prototype system represents recurring evidential requirements and maps them into the same persistent identity space through a learned compatibility function. Query-conditioned routing then materializes the relevant dataset-specific evidence, allowing the downstream agent to reason over an organized evidential state instead of rebuilding cross-schema relationships on every request. On the Data Agent Benchmark, which contains 54 queries across 12 heterogeneous datasets, the full implementation achieved 94.67% dataset-macro stratified Pass@1 across five complete trials and succeeded on 258 of 270 raw query attempts. The reported score was 55.51% for the benchmark’s Claude Opus 4.6 reference agent. The system ranked first among 40 leaderboard entries at submission. These results support the paper’s claim that learned compression and routing can improve the evidence organization available to enterprise data agents, although the abstract reports results for this benchmark and implementation rather than broader deployment conditions.
AI News
The latest AI releases, research, products, and industry updates.
Loading...