Cost-Aware AI Scoring and Human Escalation for Valid Hypothesis Testing
Summary
Large language models are increasingly used to judge outputs and label data, but their noisy or biased judgments cannot simply be treated as ground truth in formal statistical inference. This paper studies how AI evaluations and selective human verification can be combined to test hypotheses while controlling type-I and type-II errors at minimum cost. The setting uses a fixed pool of items with hidden binary labels: an item may receive an AI score, go directly to a human, be escalated to a human after the AI report, or be left unqueried as evidence accumulates. The authors derive an information-theoretic lower bound on the cost of achieving specified testing errors and use a report-dependent information frontier to characterize the value of AI information and human verification. They then introduce SCALE, a sequential cost-aware policy that selects AI-scored items and adaptively escalates selected cases to people. SCALE remains valid at finite sample sizes and approaches the lower bound to first order as the target error probabilities go to zero. The framework is also extended to an unknown AI-output model estimated from paired AI-human pilot data. Numerical results show that SCALE approaches human-only or AI-only testing when one source is clearly better, while producing its largest savings when inexpensive AI judgments and selective human checks are both useful.