hn.today

Show HN: Jevman – AI decision models play Pac-Man

opper.ai20 points2 comments
Screenshot of Show HN: Jevman – AI decision models play Pac-Man

Jevman is a public, open-source benchmark that pits AI decision models against Pac-Man’s classic arcade ghosts to measure real-time decision quality. Six models each played 100 full games; results are shown on a leaderboard with high scores and mean scores reported with 95% margins of error. Users can watch recorded games, play against the models, and inspect per-game replays for verification. Any model reachable via an HTTP endpoint can join: the repo includes a 34-line example endpoint, and hosted, fine‑tuned or locally run models are all supported. The project is licensed AGPL-3.0 and CI replays submitted runs to validate reported scores before adding them to the public ranking.

The benchmark enforces strict, repeatable rules: every junction the game sends the current maze state and asks for a direction, and models return probabilities per direction; Pac-Man then selects an action. Responses longer than 2 seconds are replaced by a simple backup rule and counted as backup moves. Each game runs until three lives are lost or a 5‑minute cap (no run approached the limit; the longest was 2:24). Rankings use the mean score with ±2 standard errors to declare ties, and every game is recorded to allow exact replay and independent checking.

Read on opper.ai2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.