Phylo Scales Biology AI with Frontier Open Models
Summary
Phylo, the applied research lab behind Biomni Lab, uses long-running AI agents to help biologists review literature, generate hypotheses, design experiments, and analyze data. These workflows can run for hours or days, make hundreds of tool calls, span datasets of tens to hundreds of gigabytes, and connect to HPC clusters and specialized protein models. Because inference spending determined user quotas, Phylo evaluated proprietary and open-weight models through its BiomniBench tests for quality, latency, and cost. It moved most traffic to frontier open-weight models served through Fireworks, while retaining proprietary models for the hardest tasks and letting users choose models. Phylo reports 2x month-on-month user growth at 60% lower cost, with lower cost per task improving quota utilization and retention. A fast serverless endpoint also roughly halved time to first token, an effect that compounds across hundreds of turns in an agent run. The company says it can evaluate and move new open models into production within about 24 hours. Phylo does not self-host the serving stack, relying instead on Fireworks for serverless inference and engineering support. It plans to use millions of monthly user traces as training data for fine-tuning and reinforcement learning aimed at difficult biology problems.