HARDEN Uses Constrained Evolutionary Search to Create Harder LLM Evaluations
Summary
Language models are often evaluated on curated benchmarks that may not capture the complexity of enterprise deployments. The paper introduces HARDEN, a constrained evolutionary search method that transforms existing evaluation cases into harder variants while keeping their expected outputs fixed. It searches across generated, domain-specific complexity axes and enforces constraints on task semantics, realism, and execution validity. The authors evaluate it on FinQA, PubMedQA, and ContractNLI using three Qwen3.5 scales: 35B-A3B, 122B-A10B, and 397B-A17B. Across these settings, HARDEN lowers task-model accuracy by an average of 22.7%. Compared with single-pass baselines using the same feasibility checks, it produces reductions of up to 49.9% relative to baseline accuracy. The results indicate that evolutionary search can create substantially harder evaluation cases that remain valid and answer-preserving.