Back to News
RSS feedarxiv.org

PFArena Benchmarks AI Models for Protein Modification

Summary

PFArena is a benchmark for evaluating AI systems that assist protein modification, where the sequence space is large and laboratory validation is expensive and slow. It defines four controlled task interfaces covering single-mutant generation and multi-mutant ranking, with different amounts of mutation-fitness data to represent research settings with different levels of prior experimental evidence. The study evaluates six protein language models, six large language models, and five LLM-based agents using complementary measures of peak performance and overall performance. Protein language models are strongest at open-ended single-mutant generation, where protein-specific prior knowledge is useful. Large language models and agents perform well at ranking multiple mutants, especially when target-specific fitness data are available. Across all model families, performance deteriorates as the search space expands and mutation depth increases. The authors release the benchmark suite and code to support reproducible research on model-assisted protein modification.