Facebook Research presents Reinforcement Learning from eXpert-Aligned Rubrics (RL-XAR), a method designed to improve language-model writing on tasks where quality is difficult to verify automatically. The approach starts with high-quality human texts, trains a rubric generator to score those texts above model continuations, and then uses reinforcement learning against the learned rubrics. The rubric-generation and model-training process is repeated until the optimizer can no longer find a clear human-model gap. The work first finds that ordinary LLM judges can prefer model writing to expert human writing: pairwise judgments favored models in 63.5% to 84.6% of academic-section comparisons, while a standard generated rubric selected model writing 100% of the time. On paper-section writing, RL-XAR trained Qwen3.5-27B and achieved a worst-rubric score of 9.60 on a human-normalized scale where human writing scored 10, while blind expert comparisons preferred the trained model 16 to 2. On story continuation, the score rose from 2.8 to 8.2 and human readers preferred the RL-XAR model 19 to 1. Gains were smaller on Wikipedia section writing, increasing the score from 2.7 to 4.0 and leaving the model behind frontier systems. The authors attribute this partly to weaknesses in the judge used during training. Additional experiments indicate that judge capability, the quality of human reference writing, rubric iteration, and meta-optimizer choice all affect results. The authors caution that the learned rubrics may still favor their model and that neither the trained writer nor the graders are perfect. They plan to release a technical report on arXiv.
