Back to News
RSS feedarxiv.org

HarvestBench Tests Whether LLM Agents Pay to Avoid Killing Animals

Summary

HarvestBench is a reproducible benchmark for measuring whether LLM agents will pay to avoid causing harm to living creatures. In a cooperative corn-harvesting simulation, memoryless sub-agents control two tractors and choose whether to drive over an animal at no fuel cost or steer around it for a posted price. The benchmark also tests responses to rocks, which damage tractors, and harmless hay bales, as well as whether agents take crops from a neighbor’s field. Across nine models and 7,201 priced decisions, 3,951 involved animals, and kill rates ranged from 0.4% to 98.8%; Terra and Sol were the most merciful while GPT-4o-mini was the most cruel, with behavior not ordered by overall capability. Four of six models responded significantly to price, with estimated elasticities from 0.09 to 1.69. Every model killed wild animals more often than farmed animals on the default map, a pattern that held across tested map geometries. A morality briefing reduced kill rates below 6% in five of six reasoning models, while removing it pushed rates above 84% in all six. Because scoring uses game-log events rather than an LLM grader, the authors say the benchmark is fully reproducible and measures willingness to pay rather than stated beliefs.