Back to News
RSS feedarxiv.org

KnowBench Proposes a Deployment-Grounded Effort-Reduction Benchmark for Clinical AI

Summary

KnowBench is a proposed benchmark for evaluating clinical AI in deployment rather than judging how closely an output resembles a reference artifact or satisfies an expert rubric. Its central metric, Effort Reduction (ER), is the proportion of system-generated clinical work product accepted by the responsible clinician after expert and safety review. The clinician's review and attestation event serves as ground truth: accepted units represent work completed by the system, while corrections represent residual effort returned to the clinician. The paper applies the same construction to several administrative and clinical tasks, including visit notes, diagnosis and billing codes, orders, EHR chart summaries, after-visit summaries, and clinical decision support. It also defines degenerate cases and a reporting protocol intended to make ER claims auditable and comparable across systems. As an initial documentation measurement, the authors report more than one million signed encounters over a production window longer than six months across thirteen medical specialties. Knowtex's proprietary fine-tuned clinical foundation models, operating within a closed feedback architecture, achieved an aggregate ER of 97.99%, with specialty-level aggregates ranging from 96.8% to 98.9%. The authors present this number as an initial headline measurement rather than a complete validation of the benchmark. They state that the release only partially provides the protocol checklist and withholds some companion statistics. KnowBench is offered as a common standard against which this result and future clinical AI measurements can be assessed.