CARAT Tests Whether Materials LLMs Reason or Recite
Summary
CARAT examines whether materials language models reason from crystal structures or simply repeat answers printed in their inputs. The benchmark holds the question and gold answer constant across eight matched views, represents structural relations separately in GraphSpace, and uses fine-tuning, answer masking, evidence injection, paired inference, and claim-withholding rules to test evidence use. On the hardest families, a grounded view improves accuracy by 17.3 points over formula inputs. GraphSpace exceeds a plain periodic graph by 19.3 points, but the authors show that this mostly reflects missing information in the baseline: the gap is 1.96 points when the baseline contains the needed fields and 46.7 points when it does not. The authors also attack their own benchmark. A shortcut that reads a list directly solves four of seven hardened families, so they rebuild the benchmark until eleven such shortcuts perform near chance. The frozen model repeats a linked relation while ignoring the link in 95.6% of paired cases. After matched supervision, it reaches 99.8%; deleting the link reduces performance to 23.4%, below the best shortcut's 27.0%. The results suggest that the original model learned to recite relations, while evidence use can be trained, but benchmark design must carefully separate grounding from information availability and shortcuts.