Jevman is an open-source benchmark for testing AI decision models through Pac-Man, chosen as a simple and fast environment for evaluating real-time decisions. The project compares jev 1.13, kev, clef, clef flash, GPT-6 Luna, and Laya while each model plays against bot-controlled ghosts. The authors say the models' low latency enables real-time play. Each model was run through 100 games, and the results were published on a leaderboard. The repository is open source so other people can run their own models and submit them to the ranking. The project also provides a playable version in which people control Pac-Man while models act as the ghosts, either as a mixed group or as a selected model group. The games are run through the authors' startup, Opper, and cost about two cents per game; free credits were added for users who want to try it. The article presents the benchmark as both a model comparison and an interactive demonstration of low-latency decision making.
AI News
The latest AI releases, research, products, and industry updates.
Loading...