Evaluation
How strong is the agent, really?
A seeded round robin between the original _4399 agent and variants that each change exactly one knob, with colour-swapped pairs to cancel the first-move advantage. Every number comes with its sample size and a 95% interval.
Jump to: original benchmark · reference round robin · Python cross-check · run your own · alpha-beta efficiency
Restated, not changed
The original benchmark, with its uncertainty
The headline result from the original Python (scripts/benchmark_agent.py, Python 3.12.7): the agent beat the subject's random agent in 144 of 160 games on boards 4 to 7, 20 seeded games per size and colour. The count is unchanged; the interval is new.
144 of 160 games won against the random agent
Wilson 95% CI 84.4% to 93.8%
| Board | Interval | Win rate (95% CI) |
|---|---|---|
| 4 × 4 | 32/40 · 80.0% [65.2, 89.5] | |
| 5 × 5 | 37/40 · 92.5% [80.1, 97.4] | |
| 6 × 6 | 36/40 · 90.0% [76.9, 96.0] | |
| 7 × 7 | 39/40 · 97.5% [87.1, 99.6] |
Reference round robin
1,200 games, five agents, three board sizes
Every pairing played 40 games per board size (20 seeds, each played twice with colours swapped) on 4 × 4, 5 × 5, 6 × 6. Tournament seed 2022. Played by the TypeScript port, which reproduces the original Python move for move in deterministic mode; the original is too slow for depth-3 search at this scale, so a feasible subset was re-run in Python below.
- Depth matters; dynamic depth barely runs. Fixed depth 3 beat the original in 84 of 120 games (70.0%, 95% CI 61.3% to 77.5%). The original searched deeper than one ply on only 34 of 4,602 searched moves (0.7%): its thresholds only deepen once fewer than 15% of cells are empty, and most games end before that.
- No detectable difference from greedy one-ply. Head to head the original won 61 and lost 59 (50.8%, 95% CI 42.0% to 59.6%), and its strength minus greedy's is +15 Elo (95% CI −24 to +56). The interval still allows the original to win up to about 10 points more or 8 points less than half its games against greedy, so only larger differences are ruled out; the opening book and the late deepening add no detectable strength.
- Against random it reproduces the original benchmark. 106/120 wins (88.3%, 95% CI 81.4% to 92.9%) on boards 4 to 6, consistent with the 90% the Python run recorded on boards 4 to 7.
- Strength costs time. Depth 3 is 151 Elo above the original (95% CI 106 to 195) but spends about 279× as long per move, partly because of the pruning bug described below. No first-move advantage was detectable: Red won 50.3% (95% CI 47.5% to 53.2%).
- Games
- 1,200
- board sizes 4, 5, 6
- First-move (Red) win rate
- 50.3%
- 95% CI 47.5% to 53.2%; 1 draw
- Mean game length
- 22.2 turns
- Bradley-Terry bootstrap: 1000 resamples, seed 2022
| Agent | Strength interval | Elo vs random (95% CI) | Win rate (95% CI) | W–D–L | Illegal | ms / move (95% CI) |
|---|---|---|---|---|---|---|
| Minimax, fixed depth 3depth fixed at 3 | 490 [432, 559] | 76.0% [72.0, 79.6] | 365–1–114 | 0 | 27.50 [25.05, 29.88] | |
| Minimax, fixed depth 2depth fixed at 2 | 369 [321, 427] | 58.3% [53.9, 62.7] | 280–1–199 | 0 | 1.39 [1.29, 1.49] | |
| Minimax, dynamic depth (original)none (as submitted) | 339 [290, 396] | 53.8% [49.3, 58.2] | 258–0–222 | 0 | 0.10 [0.09, 0.10] | |
| Greedy one-plydepth fixed at 1, no opening book | 324 [274, 386] | 51.5% [47.0, 55.9] | 247–0–233 | 0 | 0.10 [0.10, 0.11] | |
| Random baselinenot a minimax agent | reference (0) | 0 | 10.2% [7.8, 13.2] | 49–0–431 | 0 | 0.01 [0.01, 0.01] |
| Pairing | W–L (D) | Win rate interval | Win rate (95% CI) | As Red / as Blue |
|---|---|---|---|---|
| Original vs Depth 2 | 55–65 | 45.8% [37.2, 54.7] | 23/60 · 32/60 | |
| Original vs Depth 3 | 36–84 | 30.0% [22.5, 38.7] | 12/60 · 24/60 | |
| Original vs Greedy | 61–59 | 50.8% [42.0, 59.6] | 34/60 · 27/60 | |
| Original vs Random | 106–14 | 88.3% [81.4, 92.9] | 53/60 · 53/60 | |
| Depth 2 vs Depth 3 | 40–79 (1) | 33.3% [25.5, 42.2] | 23/60 · 17/60 | |
| Depth 2 vs Greedy | 62–58 | 51.7% [42.8, 60.4] | 29/60 · 33/60 | |
| Depth 2 vs Random | 113–7 | 94.2% [88.4, 97.1] | 56/60 · 57/60 | |
| Depth 3 vs Greedy | 90–30 | 75.0% [66.6, 81.9] | 50/60 · 40/60 | |
| Depth 3 vs Random | 112–8 | 93.3% [87.4, 96.6] | 58/60 · 54/60 | |
| Greedy vs Random | 100–20 | 83.3% [75.7, 88.9] | 52/60 · 48/60 |
| Difference | Difference interval | Elo difference (95% CI) |
|---|---|---|
| Original − Depth 2 | −30 [−70, +11] | |
| Original − Depth 3 | −151 [−195, −106] | |
| Original − Greedy | +15 [−24, +56] | |
| Original − Random | +339 [+290, +396] | |
| Depth 2 − Depth 3 | −121 [−168, −76] | |
| Depth 2 − Greedy | +45 [+3, +86] | |
| Depth 2 − Random | +369 [+321, +427] | |
| Depth 3 − Greedy | +166 [+119, +214] | |
| Depth 3 − Random | +490 [+432, +559] | |
| Greedy − Random | +324 [+274, +386] |
Validation
Cross-check against the original Python
scripts/crosscheck_tournament.py rebuilt the same variants from the unchanged Python (patching only dynamic_depth_allocation and the opening-book flag) and replayed the pairings that are feasible in Python on boards 4 and 5, with the same colour-swapped protocol. Random streams differ between the languages, so games differ; the question is whether win rates agree. With 80 games per pairing this can only rule out large discrepancies.
| Pairing | Original Python | TypeScript port | Difference interval | Difference (pp) |
|---|---|---|---|---|
| Original vs Depth 2 | 30/80 [27.7, 48.5] | 37/80 [35.7, 57.1] | −8.8 [−23.4, +6.4] | |
| Original vs Greedy | 41/80 [40.5, 61.9] | 39/80 [38.1, 59.5] | +2.5 [−12.7, +17.6] | |
| Original vs Random | 66/80 [72.7, 89.3] | 68/80 [75.6, 91.2] | −2.5 [−14.1, +9.1] | |
| Depth 2 vs Greedy | 43/80 [42.9, 64.3] | 43/80 [42.9, 64.3] | 0.0 [−15.1, +15.1] | |
| Depth 2 vs Random | 74/80 [84.6, 96.5] | 75/80 [86.2, 97.3] | −1.2 [−9.9, +7.3] | |
| Greedy vs Random | 66/80 [72.7, 89.3] | 63/80 [68.6, 86.3] | +3.7 [−8.6, +16.0] |
Run your own
A round robin in your browser
Same harness, same statistics, run on your machine in Web Workers. Pick a seed to make the run reproducible, then export the games and the summary as CSV.
Set up a round robin
120 games, played in parallel Web Workers.
Search efficiency
Alpha-beta pruning saves almost nothing, and here is why
120 mid-game positions (40 per board size, from seeded greedy-vs-random games, seed 2022) were searched three ways at the same depth and move order: plain minimax (no pruning), the original alpha-beta as submitted, and a textbook alpha-beta. The share of the full tree each one visits is the pruning ratio; lower is better.
In the agent's own move order the original prunes very little: averaged per board size and depth it still visits at least 93.4% of the tree, and 222 of 240 searches visited every node. The cause is one line in minimax.py: the minimising branch updates beta with if beta <= min_score: beta = min_score, which can only raise beta, so beta stays at +infinity and a cut-off can only happen once the search finds a line it scores as a forced win for Red. The 18 searches where that happened visited 12.5% to 86.2% of the tree. The answer is still right (all 480 searches returned the same root value as plain minimax), just slow. With beta = min(beta, min_score) the same search on 6 × 6 at depth 3 would visit 18.1% of the tree instead of 94.9%. The agent on this site is left as submitted.
| Board · depth | Full tree (nodes) | Share of tree visited | Original: share visited | Textbook: share visited | Same value |
|---|---|---|---|---|---|
| 4×4 · d240 positions | 124 | 95.9% [90.4, 100.0]122 nodes | 44.5% [39.3, 49.8]49 nodes | 40/40 · 40/40 | |
| 4×4 · d340 positions | 1,275 | 93.4% [86.8, 98.3]1,249 nodes | 29.6% [25.0, 34.0]281 nodes | 40/40 · 40/40 | |
| 5×5 · d240 positions | 322 | 97.8% [93.4, 100.0]322 nodes | 35.1% [30.9, 39.3]103 nodes | 40/40 · 40/40 | |
| 5×5 · d340 positions | 5,599 | 95.9% [90.8, 100.0]5,527 nodes | 21.2% [17.8, 25.4]886 nodes | 40/40 · 40/40 | |
| 6×6 · d240 positions | 566 | 98.2% [95.0, 100.0]562 nodes | 32.4% [28.8, 35.7]165 nodes | 40/40 · 40/40 | |
| 6×6 · d340 positions | 13,614 | 94.9% [89.3, 99.3]13,299 nodes | 18.1% [15.9, 20.2]2,007 nodes | 40/40 · 40/40 |
Canonical (r, q) move order
| Board · depth | Full tree (nodes) | Share of tree visited | Original: share visited | Textbook: share visited | Same value |
|---|---|---|---|---|---|
| 4×4 · d240 positions | 124 | 95.4% [89.3, 100.0]121 nodes | 39.4% [32.9, 46.3]43 nodes | 40/40 · 40/40 | |
| 4×4 · d340 positions | 1,275 | 92.3% [85.1, 98.1]1,245 nodes | 26.4% [21.0, 32.1]254 nodes | 40/40 · 40/40 | |
| 5×5 · d240 positions | 322 | 100.0% [100.0, 100.0]322 nodes | 38.2% [30.7, 45.4]103 nodes | 40/40 · 40/40 | |
| 5×5 · d340 positions | 5,599 | 97.3% [94.0, 100.0]5,559 nodes | 25.5% [20.5, 31.0]1,116 nodes | 40/40 · 40/40 | |
| 6×6 · d240 positions | 566 | 95.6% [88.5, 100.0]552 nodes | 30.1% [25.1, 35.6]160 nodes | 40/40 · 40/40 | |
| 6×6 · d340 positions | 13,614 | 92.7% [84.8, 99.3]13,210 nodes | 16.8% [13.5, 20.3]2,247 nodes | 40/40 · 40/40 |