Multimodal language models perform near the ceiling on conventional food-recognition benchmarks, but this success may reflect visual matching rather than cultural understanding. The authors introduce CulturalMenuBench, a benchmark with 4,870 items in 10 languages spanning 18 regions. Its 10 tasks combine final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, covering recognition through process-grounded cultural attribution. Tests of 12 models found that systems scoring above 94% on standard multiple-choice tasks achieved no more than 56% when attributing dishes to Chinese regional cuisines, even with the same four-choice format. Error patterns were consistent with random guessing, and performance tracked visual distinctiveness more closely than cultural structure. Models classified cuisines 7 to 18 percentage points more accurately from dish names than from images, suggesting that relevant knowledge exists but is not reliably activated by visual input. Removing sequential cooking images selectively reduced performance on process-grounded tasks while leaving other tasks stable, confirming the importance of procedural evidence. The authors argue that future training should connect perception, cooking procedure, and cultural context. Code and data are publicly available.
AI News
The latest AI releases, research, products, and industry updates.
Loading...