Back to News
RSS feedgamowlabs.com

LabBench Evaluates Whether AI Agents Can Choose the Next Wet-Lab Experiment

Summary

Gamow Labs introduces LabBench, a benchmark for testing whether AI agents can decide which biological experiment should be run next. It contains 20 held-out tasks built from real wet-lab records in drug discovery and genomics; each task gives an agent the records available at a historical decision point while withholding the laboratory’s interpretation and decision. Agents are scored against 20–22 binary criteria tied to the actual decision, with biologists reviewing the tasks and additional hardening applied to prevent trivial solutions. Five frontier agents were each run once per task in their vendors’ harnesses. GPT-6 Astra achieved the highest mean score at 40.8%, narrowly ahead of Claude Opus 5.5 at 40.5%; Astra had a median task time of five minutes, compared with 34.6 minutes for Opus. Grok 4.7, Muse Spark 1.3, and Gemini 3.8 Flash scored 29.6%, 28.2%, and 18.4%. Across 406 criteria, 182 were passed by no agent and 31 by all five, suggesting substantial shared failure modes. Agents were better at identifying and explaining prior results than at choosing, committing to, or ranking experiments: they passed 47% of identification criteria but 21% of choice criteria, and none passed any of the 13 criteria about which experiment should come first. The best agent passed only 9% of experiment-design criteria. In a small follow-up, adding one sentence that directed attention to evidence caused GPT-6 Astra to pass all five tested core decisions, raising individual task scores in four examples and slightly improving a fifth. The authors interpret this as evidence that the models possess relevant biological knowledge but do not reliably retrieve and apply it when making autonomous choices. They argue that dependable experimental design remains a key bottleneck for AI-assisted and eventually autonomous wet-lab research.