CHOIR Reveals Distinct Model Voices Beneath Surface Agreement
Summary
Open-ended language-model answers can create false plurality when different systems repeat the same familiar defaults. The paper introduces CHOIR, or Collective Hierarchically-Ordered Inquiry Responses, a framework that adapts free-list elicitation from cognitive anthropology to language-model ensembles. It repeatedly asks models for ranked lists, clusters responses into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. On the external Infinity-Chat 100 prompt bank, CHOIR found high surface agreement: 93 of 100 prompts showed agreement above chance. It nevertheless distinguished narrow prompts from broader prompts in which additional answer depth could be recovered. Across Infinity-Chat 100 and a separate 27-question diagnostic bank designed to isolate mechanisms, base-model identity was the strongest recoverable signature. Persona prompts changed which concepts were surfaced, but those concepts remained within the signatures associated with the underlying base models. The framework also includes a source-blind ranking module that prioritizes rare but stable candidates for later inspection. The authors present CHOIR as a way to turn apparent open-ended homogeneity into a diagnostic question: where models agree, why they agree, and what alternatives remain reachable through structured probing.