Eight AI Models Identify Animal Traces with 37.15% Peak Accuracy
Summary
Labqoat's Wildlife CSI benchmark tests whether AI models can identify the animal that left a photographed trace. It uses 2,000 images from the AnimalClue datasets: 400 each of droppings, footprints, feathers, eggs, and bones. Each of eight models received one image and its country, and had to return a single species name; exact species, genus, and family matches were scored, while failed or unusable calls counted as misses. Claude Opus 5.5 led with 743 exact species identifications, or 37.15%, followed by Muse Spark 1.3 at 33.35% and GPT-6 Astra at 31.70%. The results varied sharply by trace type: Opus identified 205 of 400 egg images but only 82 of 400 footprint images, and footprints were the hardest category overall. Across the models, 842 images received no correct species identification, while even the leading model identified the correct family on only 51.75% of images. The comparison also reports estimated or recorded costs, ranging from about $0.74 to $102.43, though some figures are provisional and some model failures were caused by image-safety filters. The author cautions that the benchmark is not enough to establish what a wildlife expert should score, and that species labels in source observations may be more precise than a single trace photo warrants.