Study Finds Frontier AI Models Are Getting Funnier at Original Jokes
Summary
A LessWrong study examines whether progress in measurable AI capabilities also appears in the harder-to-verify domain of humor. The author tested six closed models from OpenAI and Anthropic: Haiku 4.5, GPT-5 mini, GPT-5.2, Opus 4.5, GPT-6 Astra, and Fable 5.1. To reduce memorized or copied jokes, each model had to write an English joke of no more than 40 words using two randomly selected dictionary words, while avoiding country-specific references. Sixty-two human respondents, blinded to model identity, rated individual jokes and compared pairs generated for the same word prompts. Newer and larger models performed better in preference share, with GPT-6 Astra, Fable 5.1, and Opus 4.5 forming the stronger group. Opus 4.5 produced a chuckle-or-laughter response in up to 86.4% of its rated jokes, while GPT-6 Astra reached 80.0% in one reported measure. Across the models, chuckle-or-better ratings had a Spearman correlation of 0.714 with release recency. Performance on humor also appeared positively correlated with creative-writing and several math, science, and coding benchmarks, although the study had only four to six comparisons for each external benchmark and wide error bars. Human raters agreed substantially on whether jokes were funny at all but disagreed more on which specific joke was better. An LLM judge showed substantial agreement with human model rankings, but the author says careful ablations are still needed to distinguish effects from reinforcement learning, safety training, and model scale.