ArgGYM Introduces a Verified Benchmark for Defeasible Reasoning
Summary
ArgGYM is a procedural benchmark and reinforcement-learning-with-verifiable-rewards environment for structured defeasible reasoning, where conclusions can be supported, defeated by counter-evidence, reinstated, or revised as new arguments appear. It divides this capability into twelve tasks and scores outputs with a symbolic argumentation engine that computes formal reasoning states. The frozen benchmark contains 1,440 verified instances arranged across fifteen curriculum configurations, with two argument preference orderings, weakest-link and last-link, and two set orderings, elitist and democratic. The same generators and verifiers can create fresh evaluation instances, reducing dependence on static test sets, and can provide verifiable rewards for training. On the frozen benchmark, frontier and open-weight models display sharply different reasoning profiles: they often recover parts of structured answers without completing the full task. Performance also falls in later curriculum configurations, where dependencies are longer and structures interact more heavily. The authors release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.