LLM evaluation · bring your own key
Can a language model beat our 2022 agent?
The model plays short games against the original minimax agent. Each turn it gets the rules, the board, the move history and the full list of legal moves, and must answer with a JSON move. The referee checks every answer: an illegal move, or a reply with no usable move (malformed, cut off or refused), is counted and the model is asked again (up to three times) before it forfeits. Both providers are scored by the same rule.
The same seeds and colours are replayed with the random and greedy agents and a scripted first-legal-cell player in the model's seat, so the comparison is like for like (if a run stops early, on the games the model finished). This page is an evaluation harness, not a claim: a few games give wide intervals. What is sent to the provider.
Evaluation setup
Colours alternate, so half are played as Red.
Drives the minimax opponent's move shuffle and tie-breaks.
Side by side against the original agent
| Player in the seat | Win rate interval | Win rate (95% CI) | First answer rejectedper turn (95% CI, clustered by game) | Paired with the LLMsame games, exact McNemar | Latency · tokens |
|---|---|---|---|---|---|
| LLM (your key)not run yet | – | – | – | – | |
| Random baseline0 draws | 0/4 [0.0, 49.0] | 0 by construction | after a run | instant | |
| Greedy one-ply baseline0 draws | 1/4 [4.6, 69.9] | 0 by construction | after a run | instant | |
| First legal cell (scripted)0 draws | 2/4 [15.0, 85.0] | 0 by construction | after a run | instant |