The study examines PRM-Pruned Fragment Grafting (PPFG), an inference-time method that extracts a high-scoring prefix from a chain pruned by a process reward model and inserts it as an in-context demonstration into a sibling chain that is still decoding. The authors test the mechanism at an operating point where earlier fragment-grafting work reported gains only with additional compensating components. On Qwen2.5-7B-Instruct with Math-Shepherd over all 500 MATH500 problems and three random seeds, both stagnation-targeting and random-targeting PPFG are statistically indistinguishable from an independent parallel chain-of-thought baseline on every measured axis. Among 322 stagnation-rule injection events, only 14% targeted a genuinely struggling chain; the remainder affected chains that had already succeeded, were near completion, or were on a flat process-reward plateau. Refining the compound gate did not simultaneously produce accurate targeting and enough firing density, while a random control achieved the same parity at 2.4 times the firing rate, indicating that the result is not specific to the targeting heuristic. The finding replicated across three base language models, six benchmarks, a second process reward model, and a compatibility-gate sweep. Two-one-sided-tests analysis supported positive equivalence in all twelve Qwen/LLaMA cells. Injected chains were pruned at 2.75 times the matched-step rate in a per-event check, but a surviving-sibling counterfactual found no population-level compensation. A hindsight oracle placed the maximum per-problem gain from choosing PPFG over independent decoding at 0.13 percentage points. The authors present an equivalence-testing template for establishing null inference-time mechanisms and limit every claim to the tested operating point.
AI News
The latest AI releases, research, products, and industry updates.
Loading...