What Happens If an AI Agent Writes One Million Tokens per Second?
Summary
This field note presents a thought experiment: imagine one capable AI agent generating one million output tokens per second. It explicitly makes no claim that any current model can reach that rate, and distinguishes output speed from context capacity, input processing and aggregate service throughput. At that hypothetical speed, 1,000 story endings would take about one second to write, 40 app drafts about 0.8 seconds, 10,000 critiqued candidates three seconds, and 10,000 conversation rehearsals ten seconds. Those figures count generated output only; they exclude reading, ranking, tool calls, permissions, tests, deployment, human review and physical experiments. The article argues that generation would stop being the slow part, while choosing what to trust would remain difficult. In one illustrative software example, 800,000 output tokens fall from 2 hours 13 minutes 20 seconds at 100 tokens per second to 0.8 seconds at the imagined rate, but a fixed 60-second check leaves an end-to-end speedup of about 132.57x rather than 10,000x. Longer fixed checks reduce the gain further, illustrating the logic of Amdahl’s law. Four scenarios show the remaining constraints: taste in selecting a story ending, clear specifications and tests for software, physical evidence for scientific ideas, and fidelity when rehearsing a real conversation. The note also warns that more drafts from one model may share the same blind spots, faster does not mean cheaper because hardware consumes energy, and token volume is not understanding. Its practical advice is to define “done,” set evaluation criteria before reviewing options, identify the slow check, and request disagreement built on different assumptions.