What AI Can and Cannot Tell Us About Pain in a Cat’s Face
Summary
Research on the Feline Grimace Scale suggests that AI can help with a narrow, validated task, but general chatbots are not reliable pain scorers. The scale, developed at the Université de Montréal, rates five facial action units—ears, eyes, muzzle, whiskers and head position—from 0 to 2. A total of 4 or more out of 10 separated cats needing pain relief from other cats in the original study, with 90.7% sensitivity and 86.6% specificity. In 2023, researchers trained a two-stage neural-network system on 3,447 expert-scored cat-face photos: one model located 37 facial landmarks and another converted them into a grimace score. The best smartphone-sized combination distinguished painful from pain-free cats with 95.5% accuracy. A separate 2025 Scientific Reports study tested ChatGPT with GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro and Perplexity on 50 photos, repeating the test after two months and comparing results with a veterinary anaesthesiologist. All four systems scored pain lower than the expert; Claude was closest on average in the second test but individual scores varied, while the others were both low and inconsistent. The authors describe the result as a snapshot because chatbot behavior changes with updates. Owners can apply the scale to a calm, awake cat after observing it for 30 seconds, but the tool is intended for acute rather than chronic pain, is less accurate soon after sedation, and cannot rule out illness. A score is a record to discuss with a veterinarian, not a diagnosis.