Back to News
RSS feedana15.substack.com

A Practical Guide to Automating White-Collar Work with AI Agents

Summary

This practical guide explains how to turn repeatable white-collar work into environments that AI agents can navigate. It uses a management-consulting forecasting task as an example: an agent must inspect spreadsheets, calculate a 2030 revenue range, and send the result through workplace tools. The proposed setup recreates a workplace with files, mail, chat, calendars, and application state, then exposes structured operations through MCP servers for files, spreadsheets, documents, and other tools. Docker and the Archipelago runtime provide a reproducible environment, while a small agent loop passes the task, tool schemas, and previous results to an LLM and executes its structured tool calls. The guide adds a final-answer mechanism, context summarization, and error recovery so the agent can continue after incorrect calls or context limits. In the worked example, the author manually identifies LATAM customer counts, calculates a 4.91% annual growth rate, and derives the correct 2030 revenue range. A Qwen3-8B run fails after choosing the wrong workbook and relying on unsupported assumptions introduced during context compaction, demonstrating why trajectory quality matters beyond successful execution. The article then describes LLM-based judging, exact verification where possible, and a reinforcement-learning loop in which rollouts receive rewards and update the model. It cites reported results showing improvement from 16.11% to 27.29% on the discussed white-collar benchmark, while noting that the best cited result was still about 30%. The conclusion stresses that expert demonstrations, human judgment, creativity, and reliable checking remain important because current systems are not fully autonomous across measured AI research work.