Profession-Specific Agent Prompts Raise Costs Without Improving Scientific Accuracy
Summary
A study evaluates 503 profession-specific AGENTS.md profiles with Gemini 3.8 Flash in the Pi agent harness. Across nine text-based science benchmarks covering 4,531 sampled questions and 100 matched profiles, 4,488 items completed all five prompt conditions after API-error retries. Full profession profiles were 0.6 percentage points less accurate than a minimal baseline, with a 95% bootstrap interval from -1.5 to +0.2, and no benchmark showed a statistically clear improvement. They generated 1.5 to 2.3 times as many output tokens and cost 2.2 to 4.5 times more per successful call. On 60 tool-using BioMysteryBench problems, profiles achieved a 46.7% mean solve rate versus 56.7% for the baseline, largely because profile runs hit token and time limits more often. Longer prompts did improve first-pass reliability on SuperGPQA when provider API drops were frequent: the profile reached 71.6% versus 54.0% for the short baseline, but generic and mismatched prompts performed similarly. The authors conclude that full profession profiles do not improve accuracy for the tested model and tasks, while selective retrieval and open-ended scientific tasks remain untested.