Back to News
RSS feedhieraticbench.vercel.app

HieraticBench Tests AI’s Ability to Read Ancient Egyptian Hieratic

Summary

HieraticBench is a benchmark designed to test whether generative AI models can identify, read, and translate hieratic, the cursive form of ancient Egyptian hieroglyphs used in everyday writing. Its first version contains 268 items, including real documents, signs, other Egyptian scripts used as controls, and two renditions of an unpublished sentence. The creator says current foundation models often fail to identify the sentence, sometimes labeling it as Tibetan, Urdu, Korean, or “Reformed Egyptian.” On real documents, the best model identified hieratic correctly 95% of the time, but the best result for reading individual signs was only about 13%, even when the models were explicitly told that the script was hieratic. No tested model could reliably translate the sealed sentence, which has never been published and has no answer key. The evaluation is incomplete because not every model was run on every task; only Claude models had completed the sign-reading task, and the creator planned to test Astra and Gemini 3.1 Pro on real documents. The creator also says they have no prior benchmark or evaluation experience and are not fully certain that the scoring is correct. The project is seeking additional model runs, sealed sentences, labeled signs, and input from people with knowledge of hieratic or Egyptology.