Skip to content

LLM evaluation · bring your own key

Can a language model beat our 2022 agent?

The model plays short games against the original minimax agent. Each turn it gets the rules, the board, the move history and the full list of legal moves, and must answer with a JSON move. The referee checks every answer: an illegal move, or a reply with no usable move (malformed, cut off or refused), is counted and the model is asked again (up to three times) before it forfeits. Both providers are scored by the same rule.

The same seeds and colours are replayed with the random and greedy agents and a scripted first-legal-cell player in the model's seat, so the comparison is like for like (if a run stops early, on the games the model finished). This page is an evaluation harness, not a claim: a few games give wide intervals. What is sent to the provider.

Evaluation setup

Colours alternate, so half are played as Red.

Drives the minimax opponent's move shuffle and tie-breaks.

Side by side against the original agent

Same 4 seeds and colours for every row, on 4 × 4; the opponent is always the original minimax agent. Baselines are computed in your browser. Win rates: Wilson 95% intervals, wide with a handful of games, which is the honest answer. The paired column compares each baseline with the model game by game: only games exactly one of them won count, with an exact McNemar test (with 4 games it can only detect a very large difference). Turns within a game are not independent, so the rejected-answer interval is a Wilson interval on the effective number of turns, after a design effect estimated from how much the rate varies between games (never below 1). The scripted baseline plays the first empty cell in (r, q) order; as Blue that line wins unless Red blocks it, so a model that beats the agent should also beat this row.
Player in the seatWin rate intervalWin rate (95% CI)First answer rejectedper turn (95% CI, clustered by game)Paired with the LLMsame games, exact McNemarLatency · tokens
LLM (your key)not run yet––––
Random baseline0 draws0/4 [0.0, 49.0]0 by constructionafter a runinstant
Greedy one-ply baseline0 draws1/4 [4.6, 69.9]0 by constructionafter a runinstant
First legal cell (scripted)0 draws2/4 [15.0, 85.0]0 by constructionafter a runinstant