Back to News
RSS feedblog.glyph.im

What Would a Serious AI Product Look Like?

Summary

Glyph’s essay argues that current AI products do not take their own reliability and safety limitations seriously. Because language models can produce incorrect information, the author proposes making claim verification a first-class workflow: each claim should have space for human checking, notes, and an explicit completion mark. Research tools should present large, clearly identified citations with publication dates, authors, and unmodified quotations, while AI summaries should remain secondary until verified. The essay also calls for less first-person language and fewer apologies, task-specific interfaces instead of vague natural-language commands, and clear indicators showing whether data came from an authoritative source, a tool, a user, or the model. For reproducibility, users should be able to inspect and replay workflows, understand stochastic variation, freeze deterministic steps, and see context usage, compaction, and harness-generated prompts. For coding agents, the author argues that containers and approval prompts are inadequate: products should enforce filesystem boundaries independently of the model, snapshot repositories after operations, eliminate dangerous automatic modes, group actions into reviewable plans, and support mock services for testing API behavior. The essay extends these concerns to organizations, recommending protected periods without AI output, deliberate practice to prevent skill loss, and resources for mental-health risks. Finally, the author says that available evidence has not demonstrated meaningful productivity gains under the author’s preferred measurement approach and argues that vendors should expose task-level effectiveness metrics rather than relying on benchmarks. These conclusions are presented as the author’s critique and hypothesis, not as a reported controlled study.