AI-Generated Tests Can Add Noise Without Protecting Code
Summary
Niklas Gruhn argues that coding agents often add tests that look responsible but provide little protection against regressions. He describes a test that scanned source files for the phrase “wall-clock,” causing a failure after the phrase appeared in a comment, plus tests that merely asserted a list of statuses, reimplemented production logic elsewhere, or duplicated static configuration values. He also criticizes tests that check whether exact headings remain in prompts, arguing that prompt effects are nondeterministic and should be monitored with evaluations rather than unit tests. In the codebase he examined, 404 of 6,429 tests fit these patterns. The tests did not materially slow the suite, had no production impact, and were easy to remove, but they produced red failures after ordinary changes and consumed time and tokens. Gruhn says the larger cost is lost confidence: developers can no longer tell which tests protect meaningful requirements or trust AI-generated tests to provide coverage. He suggests that better architecture, such as interceptable shared logging, may be preferable when behavior genuinely needs testing. The post presents these claims as the author’s engineering judgment, not as a controlled study.