Skip to content

Evaluation

How strong is the agent, really?

A seeded round robin between the original _4399 agent and variants that each change exactly one knob, with colour-swapped pairs to cancel the first-move advantage. Every number comes with its sample size and a 95% interval.

Jump to: original benchmark · reference round robin · Python cross-check · run your own · alpha-beta efficiency

Restated, not changed

The original benchmark, with its uncertainty

The headline result from the original Python (scripts/benchmark_agent.py, Python 3.12.7): the agent beat the subject's random agent in 144 of 160 games on boards 4 to 7, 20 seeded games per size and colour. The count is unchanged; the interval is new.

90.0%

144 of 160 games won against the random agent

Wilson 95% CI 84.4% to 93.8%

By board size (both colours pooled, 40 games each). Small samples give wide intervals.
BoardIntervalWin rate (95% CI)
4 × 432/40 · 80.0% [65.2, 89.5]
5 × 537/40 · 92.5% [80.1, 97.4]
6 × 636/40 · 90.0% [76.9, 96.0]
7 × 739/40 · 97.5% [87.1, 99.6]

Reference round robin

1,200 games, five agents, three board sizes

Every pairing played 40 games per board size (20 seeds, each played twice with colours swapped) on 4 × 4, 5 × 5, 6 × 6. Tournament seed 2022. Played by the TypeScript port, which reproduces the original Python move for move in deterministic mode; the original is too slow for depth-3 search at this scale, so a feasible subset was re-run in Python below.

  • Depth matters; dynamic depth barely runs. Fixed depth 3 beat the original in 84 of 120 games (70.0%, 95% CI 61.3% to 77.5%). The original searched deeper than one ply on only 34 of 4,602 searched moves (0.7%): its thresholds only deepen once fewer than 15% of cells are empty, and most games end before that.
  • No detectable difference from greedy one-ply. Head to head the original won 61 and lost 59 (50.8%, 95% CI 42.0% to 59.6%), and its strength minus greedy's is +15 Elo (95% CI −24 to +56). The interval still allows the original to win up to about 10 points more or 8 points less than half its games against greedy, so only larger differences are ruled out; the opening book and the late deepening add no detectable strength.
  • Against random it reproduces the original benchmark. 106/120 wins (88.3%, 95% CI 81.4% to 92.9%) on boards 4 to 6, consistent with the 90% the Python run recorded on boards 4 to 7.
  • Strength costs time. Depth 3 is 151 Elo above the original (95% CI 106 to 195) but spends about 279× as long per move, partly because of the pruning bug described below. No first-move advantage was detectable: Red won 50.3% (95% CI 47.5% to 53.2%).
Games
1,200
board sizes 4, 5, 6
First-move (Red) win rate
50.3%
95% CI 47.5% to 53.2%; 1 draw
Mean game length
22.2 turns
Bradley-Terry bootstrap: 1000 resamples, seed 2022
Bradley-Terry strengths on the Elo scale, fitted to every game (draws count half), with random fixed at 0 and percentile intervals from a bootstrap that resamples games within each pairing and board size. Win rates use Wilson intervals. Move times come from one run of the TypeScript port (Apple M4, Node v26.10.0) and are only comparable within this table.
AgentStrength intervalElo vs random (95% CI)Win rate (95% CI)W–D–LIllegalms / move (95% CI)
Minimax, fixed depth 3depth fixed at 3490 [432, 559]76.0% [72.0, 79.6]365–1–114027.50 [25.05, 29.88]
Minimax, fixed depth 2depth fixed at 2369 [321, 427]58.3% [53.9, 62.7]280–1–19901.39 [1.29, 1.49]
Minimax, dynamic depth (original)none (as submitted)339 [290, 396]53.8% [49.3, 58.2]258–0–22200.10 [0.09, 0.10]
Greedy one-plydepth fixed at 1, no opening book324 [274, 386]51.5% [47.0, 55.9]247–0–23300.10 [0.10, 0.11]
Random baselinenot a minimax agentreference (0)010.2% [7.8, 13.2]49–0–43100.01 [0.01, 0.01]
Head to head. Win rate of the first-named agent; the vertical line marks 50%. Each pairing played equal games as Red and as Blue.
PairingW–L (D)Win rate intervalWin rate (95% CI)As Red / as Blue
Original vs Depth 255–6545.8% [37.2, 54.7]23/60 · 32/60
Original vs Depth 336–8430.0% [22.5, 38.7]12/60 · 24/60
Original vs Greedy61–5950.8% [42.0, 59.6]34/60 · 27/60
Original vs Random106–1488.3% [81.4, 92.9]53/60 · 53/60
Depth 2 vs Depth 340–79 (1)33.3% [25.5, 42.2]23/60 · 17/60
Depth 2 vs Greedy62–5851.7% [42.8, 60.4]29/60 · 33/60
Depth 2 vs Random113–794.2% [88.4, 97.1]56/60 · 57/60
Depth 3 vs Greedy90–3075.0% [66.6, 81.9]50/60 · 40/60
Depth 3 vs Random112–893.3% [87.4, 96.6]58/60 · 54/60
Greedy vs Random100–2083.3% [75.7, 88.9]52/60 · 48/60
Strength differences, first-named minus second, in Elo points. Each interval comes from the same 1000 bootstrap refits as the leaderboard, so it is the right way to compare two agents (their separate intervals are correlated, so overlap is not a test). An interval that excludes 0 is a detectable difference.
DifferenceDifference intervalElo difference (95% CI)
Original − Depth 2−30 [−70, +11]
Original − Depth 3−151 [−195, −106]
Original − Greedy+15 [−24, +56]
Original − Random+339 [+290, +396]
Depth 2 − Depth 3−121 [−168, −76]
Depth 2 − Greedy+45 [+3, +86]
Depth 2 − Random+369 [+321, +427]
Depth 3 − Greedy+166 [+119, +214]
Depth 3 − Random+490 [+432, +559]
Greedy − Random+324 [+274, +386]

Validation

Cross-check against the original Python

scripts/crosscheck_tournament.py rebuilt the same variants from the unchanged Python (patching only dynamic_depth_allocation and the opening-book flag) and replayed the pairings that are feasible in Python on boards 4 and 5, with the same colour-swapped protocol. Random streams differ between the languages, so games differ; the question is whether win rates agree. With 80 games per pairing this can only rule out large discrepancies.

Win rate of the first-named agent. Difference = Python − TypeScript, with Newcombe's 95% interval; an interval that contains 0 is consistent with no difference.
PairingOriginal PythonTypeScript portDifference intervalDifference (pp)
Original vs Depth 230/80 [27.7, 48.5]37/80 [35.7, 57.1]−8.8 [−23.4, +6.4]
Original vs Greedy41/80 [40.5, 61.9]39/80 [38.1, 59.5]+2.5 [−12.7, +17.6]
Original vs Random66/80 [72.7, 89.3]68/80 [75.6, 91.2]−2.5 [−14.1, +9.1]
Depth 2 vs Greedy43/80 [42.9, 64.3]43/80 [42.9, 64.3]0.0 [−15.1, +15.1]
Depth 2 vs Random74/80 [84.6, 96.5]75/80 [86.2, 97.3]−1.2 [−9.9, +7.3]
Greedy vs Random66/80 [72.7, 89.3]63/80 [68.6, 86.3]+3.7 [−8.6, +16.0]

Run your own

A round robin in your browser

Same harness, same statistics, run on your machine in Web Workers. Pick a seed to make the run reproducible, then export the games and the summary as CSV.

Set up a round robin

Agents
Board sizes
Colour-swapped pairs per pairing and size

120 games, played in parallel Web Workers.

Choose agents and board sizes, then run. Every game is reproducible from the seed; each pairing plays the same seed twice with colours swapped.

Search efficiency

Alpha-beta pruning saves almost nothing, and here is why

120 mid-game positions (40 per board size, from seeded greedy-vs-random games, seed 2022) were searched three ways at the same depth and move order: plain minimax (no pruning), the original alpha-beta as submitted, and a textbook alpha-beta. The share of the full tree each one visits is the pruning ratio; lower is better.

In the agent's own move order the original prunes very little: averaged per board size and depth it still visits at least 93.4% of the tree, and 222 of 240 searches visited every node. The cause is one line in minimax.py: the minimising branch updates beta with if beta <= min_score: beta = min_score, which can only raise beta, so beta stays at +infinity and a cut-off can only happen once the search finds a line it scores as a forced win for Red. The 18 searches where that happened visited 12.5% to 86.2% of the tree. The answer is still right (all 480 searches returned the same root value as plain minimax), just slow. With beta = min(beta, min_score) the same search on 6 × 6 at depth 3 would visit 18.1% of the tree instead of 94.9%. The agent on this site is left as submitted.

The agent's own shuffled move order. Upper mark: original; lower, fainter mark: textbook alpha-beta. Intervals: percentile bootstrap over positions (2,000 resamples). Last column: positions where each pruned search returned the same root value as plain minimax.
Board · depthFull tree (nodes)Share of tree visitedOriginal: share visitedTextbook: share visitedSame value
4×4 · d240 positions12495.9% [90.4, 100.0]122 nodes44.5% [39.3, 49.8]49 nodes40/40 · 40/40
4×4 · d340 positions1,27593.4% [86.8, 98.3]1,249 nodes29.6% [25.0, 34.0]281 nodes40/40 · 40/40
5×5 · d240 positions32297.8% [93.4, 100.0]322 nodes35.1% [30.9, 39.3]103 nodes40/40 · 40/40
5×5 · d340 positions5,59995.9% [90.8, 100.0]5,527 nodes21.2% [17.8, 25.4]886 nodes40/40 · 40/40
6×6 · d240 positions56698.2% [95.0, 100.0]562 nodes32.4% [28.8, 35.7]165 nodes40/40 · 40/40
6×6 · d340 positions13,61494.9% [89.3, 99.3]13,299 nodes18.1% [15.9, 20.2]2,007 nodes40/40 · 40/40
Canonical (r, q) move order
Same positions, moves searched in sorted (r, q) order instead of shuffled.
Board · depthFull tree (nodes)Share of tree visitedOriginal: share visitedTextbook: share visitedSame value
4×4 · d240 positions12495.4% [89.3, 100.0]121 nodes39.4% [32.9, 46.3]43 nodes40/40 · 40/40
4×4 · d340 positions1,27592.3% [85.1, 98.1]1,245 nodes26.4% [21.0, 32.1]254 nodes40/40 · 40/40
5×5 · d240 positions322100.0% [100.0, 100.0]322 nodes38.2% [30.7, 45.4]103 nodes40/40 · 40/40
5×5 · d340 positions5,59997.3% [94.0, 100.0]5,559 nodes25.5% [20.5, 31.0]1,116 nodes40/40 · 40/40
6×6 · d240 positions56695.6% [88.5, 100.0]552 nodes30.1% [25.1, 35.6]160 nodes40/40 · 40/40
6×6 · d340 positions13,61492.7% [84.8, 99.3]13,210 nodes16.8% [13.5, 20.3]2,247 nodes40/40 · 40/40