Back to News
RSS feedarxiv.org

GAD-RL Improves OCR Faithfulness with Adaptive On-Policy Distillation

Summary

Vision-language models can replace anomalous text in images with linguistically plausible wording, reducing OCR transcription faithfulness. This paper introduces GAD-RL, an adaptive on-policy distillation method for joint post-training that regulates a frozen teacher's guidance according to the student's task performance and local output distributions. Distillation is disabled for response groups containing an output with a task reward of at least 0.95, while its strength is continuously reduced as the group-mean reward rises. The method also weights the forward Kullback-Leibler update by the student's probability of the teacher's Top-1 token, limiting auxiliary updates when the student gives little support to that candidate. The design is motivated by offline analysis showing that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across checkpoints and across response groups with different rewards. On Qwen3.5-2B, GAD-RL reaches 59.92% Micro Recall on CHAOS-Bench, exceeding GRPO by 8.45 percentage points and fixed-weight GRPO+OPD by 4.43 points. It also reports an Overall score of 91.18 on OmniDocBench v1.6.