A Polyphonic Conception of AI Understanding
Summary
The paper asks what it means to say that an AI model understands something when people must decide whether to trust its output. It argues that purely mathematical or statistical descriptions cannot cleanly distinguish trustworthy from untrustworthy outputs without implicitly returning to the question of understanding. The authors criticize a “monophonic” conception, which assumes that one underlying mechanism supports all the capacities associated with understanding. Drawing on mechanistic evidence, they characterize large language models as “polyphonic”: their outputs arise from coalitions of parallel mechanisms with uneven reliability. These mechanisms may complement or duplicate one another, or suppress competing mechanisms, and multiple coalitions may be sufficient for a task without any single coalition being indispensable. This structure makes straightforward inferences from a model’s behavior to its understanding hazardous. In response, the paper proposes a conception suited to polyphonic systems. An attribution of understanding should concern internal organization, especially whether sound circuitry is reliably and correctly recruited and remains in control of the output. The authors present this as a tractable basis for guiding judgments about when an AI output deserves trust.