MedCode Improves LLM Medical Calculations Through Embedded Code
Summary
Large language models often perform well on medical exams and question-answering benchmarks but remain unreliable when a task requires an exact numerical result. This is important for calculations used in medication dosing, organ-function assessment, and prognostic scoring, where small errors may have serious clinical consequences. The paper introduces MedCode, a framework that trains an LLM to identify the relevant clinical calculator, extract its input variables, and generate executable code for the calculation. A deterministic interpreter performs the arithmetic and returns the value with an explanation and unit, reducing reliance on the model’s own numerical manipulation. The authors build supervised fine-tuning and preference datasets from the MedCalc benchmark and add an ICU-focused calculation dataset. They also propose weighted Direct Preference Optimization, which gives greater emphasis to preference pairs that the model finds difficult to distinguish. Experiments with LLaMA3-8B, Qwen2.5-7B, and Mistral-7B report absolute accuracy improvements of 20–30 percentage points, supporting embedded code generation as a way to improve medical calculation performance.