Skip to content

Methods

How every number on this site was made

Provenance, methods and evaluation design for the tournament, the A* study and the alpha-beta study; the assumptions and limitations behind them; and the decisions, agent card and AI use statement that go with them.

Where the data comes from

Data provenance

There is no external dataset. Every result is generated by running code in this repository with a fixed seed, and the generating script is named next to it. The original 2022 Python in coursework/ is never modified; the TypeScript port is checked against it by parity tests (DR-003).

ArtefactProduced byScriptRun details
Original benchmark: agent vs random, 160 gamesOriginal Python, unchanged, via the subject's refereescripts/benchmark_agent.pyPython 3.12.7, PYTHONHASHSEED=0
Parity fixtures (A*, rules, evaluation, minimax, moves)Original Python, unchanged; shuffle replaced by a sort and bias fixed at 1scripts/generate_parity_fixtures.pyCPython 3.12
Reference round robin, 1,200 gamesTypeScript port (parity-tested against the original Python)web/scripts/generate-reference-studies.tsseed 2022; Apple M4; Node v26.10.0; 2026-10-05
Python cross-check, 480 gamesoriginal Python (coursework/Project Part B), variants by patching one knobscripts/crosscheck_tournament.pyPython 3.12.7, PYTHONHASHSEED=0; seed 2022
A* paired heuristic study, 980 boardsTypeScript port (parity-tested against the original Python)web/scripts/generate-reference-studies.tsseed 2022; the notebook's board generator
Alpha-beta efficiency, 480 searchesTypeScript port (parity-tested against the original Python)web/scripts/generate-reference-studies.tsseed 2022; positions from greedy-vs-random games

What was done

Method

Agents and variants

The original agent plus variants that each change one knob: depth fixed at 2 or 3, and greedy one-ply (depth 1, no opening book). The same variants are rebuilt from the unchanged Python for the cross-check by patching dynamic_depth_allocation and the opening-book flag. Random is the subject's baseline.

Round robin

Every pairing plays the same number of games on each board size, as colour-swapped pairs: one seed, played twice with the agents swapping Red and Blue, so the first-move effect cancels within a pairing. The referee port validates every move; an illegal move or crash forfeits the game. Per-move seeds are derived from the game seed and move number, so any game can be replayed.

Strength and uncertainty

Win rates: Wilson score intervals. Strengths: Bradley-Terry fitted by the MM algorithm (draws count half, one virtual draw per pairing to keep estimates finite), reported on the Elo scale with random fixed at 0, and percentile intervals from a bootstrap that resamples games within each pairing and board size. Move times: bootstrap intervals over per-game means.

A* heuristic study

Paired design: both heuristics search the same random boards from the notebook's generator, and the unit of analysis is the board. Mean per-board difference with a paired bootstrap interval, a Wilcoxon signed-rank test (zero differences dropped, tie-corrected normal approximation), the matched-pairs rank-biserial correlation and Cohen's d_z. Optimality is checked against breadth-first search on the same board.

Alpha-beta study

The same position is searched at the same depth and move order three ways: plain minimax, the original alpha-beta, and a textbook alpha-beta. The pruning ratio is the share of the full tree visited, with bootstrap intervals over positions. Every search is also checked to return the same root value as plain minimax.

Statistical code

All helpers live in web/src/lib/stats with unit tests against reference values from scipy 1.17, statsmodels and R (Wilson intervals, Newcombe differences, Wilcoxon exact and approximate p-values, type-7 quantiles, Bradley-Terry maximum likelihood). Every interval shown records its seed and number of resamples.

What counts as evidence

Evaluation design

  • Questions asked: how strong is the original relative to controlled variants and a baseline; does dynamic depth help; does the A* heuristic choice trade speed for optimality; how much work does pruning save; and (with your own key) how does a language model compare to the random and greedy agents on identical seeds.
  • Sample sizes are stated with every result. The reference round robin has 40 games per pairing per board size (120 per pairing overall), which resolves differences of roughly 15 to 20 percentage points in a head-to-head win rate; smaller differences are reported as not detectable rather than as ties.
  • Paired comparisons where two methods are compared: heuristics on the same boards (paired bootstrap and Wilcoxon for node counts, paired bootstrap and exact McNemar for shortest-path rates), search variants on the same positions, colour-swapped game pairs on the same seeds, agents compared through Elo differences taken from the same bootstrap refits (never by whether two intervals overlap), and LLM versus each baseline on exactly the games the model finished (games won by only one of them, with an exact McNemar test).
  • LLM answers are scored per turn, one rule for both providers: the rate reported is the share of turns whose first answer was rejected, split into illegal moves and replies with no usable move (malformed, cut off or refused), because retries on the same turn are not independent trials. Turns within a game are not independent either (same model, prompt and position history), so the Wilson interval is on the effective number of turns after a design effect estimated from how much the rate varies between games, never below 1 (DR-005). With a handful of games the design effect is itself rough.
  • Effect sizes, not just p-values: differences in win rate and Elo with intervals, mean node differences, rank-biserial r and d_z.
  • No self-scored metrics. Results are measured against the referee and against other agents, never against a score the author chose.
  • Multiple comparisons: ten pairings are reported side by side without correction. Each interval is valid on its own; reading the table as a whole, expect about one in twenty 95% intervals to miss.

Read before you quote a number

Assumptions and limitations

Assumptions

  • Games are independent given their seeds; colour-swapped pairs share a seed by design, and the bootstrap resamples games, not pairs.
  • Bradley-Terry assumes strength is one number per agent (transitive, no style matchups) and pools board sizes in the “all sizes” view.
  • The TypeScript port behaves like the original in random play as well as in the deterministic cases the parity tests cover; the Python cross-check supports this within its resolution.

Limitations

  • Board sizes 4 to 6 only for the round robin; the original supports up to 15 × 15, where depth 1 is likely to be weaker still.
  • Move times depend on hardware and are only comparable within one run.
  • The Python cross-check has 80 games per pairing and can only detect large discrepancies (about ±15 points).
  • The A* study uses the notebook's random boards (barriers at random, up to half the board); other board distributions may behave differently.
  • LLM-as-a-player runs are a handful of games against one opponent, with one prompt; intervals are wide by design.

Looking back

What I'd change

  • Fix the one-line beta update and replace ratio-based depth with iterative deepening under a time budget, then test the change as a new variant against the original (DR-002).
  • Add a connection-distance feature and tune weights against data, not by hand (DR-001).
  • Pre-register comparisons and sample sizes before running the next tournament, and extend it to larger boards.
  • Record the original Python's random draws so whole random games can be compared move for move across languages (DR-003).

Decision records

Decisions, as made and as measured

Transparency

AI use statement

This site has two optional features that call a large language model. Everything else, including the agent you play against and every statistic on the site, is ordinary code that runs without any AI service. This statement is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act, and the NIST AI Risk Management Framework. It is not a claim of compliance or certification with any of them.

What the AI features do

  1. Move commentator (on /play and /spectate). When the minimax agent has searched for a move, you can ask a model to explain it in plain English. The model receives only the agent's own numbers for that move: board size, turn, the move, search depth, nodes visited, the top candidate scores and the six evaluation features with their weights and contributions, plus a sentence saying what the score means at that depth: the features describe the position right after the move, so they explain the score only for a one-ply search, and a forced win found by a deeper search is the reason for the move, not the features. It is instructed to use nothing else.
  2. LLM as a player (on /llm-arena). A model plays a few short games against the original agent. Each turn it receives the rules, the board, the move history and the list of legal moves, and returns one move as JSON. Answers are scored the same way for both providers, under the same output-token budget: an illegal move, or a reply with no usable move (malformed, cut off at the token limit, or refused), is rejected and asked again, up to three times. The headline is the share of turns whose first answer was rejected, reported next to random, greedy and scripted baselines on the same seeds.

What they never do

  • They never choose the agent's moves, change its evaluation, or feed into the tournament, the A* study or any other statistic on the site.
  • They never run unless you add your own API key and press a button.
  • They never send anything to this site's operators. The site is static and has no server; calls go directly from your browser to the provider you chose.
  • They never store your key anywhere except your own browser (session storage by default, local storage only if you tick "remember on this device"), and never write it to the audit log or to any file in the repository.

Data sent to the provider

Only the prompt shown in the audit log for that call: game state, rules, legal moves and, for the commentator, the agent's evaluation breakdown. No personal information is collected or sent. Your API key goes in the request header to the provider (Anthropic or OpenAI), which processes the request under its own terms and bills your account.

Human in the loop and transparency

  • Every model output on the site is labelled AI-generated with the model name.
  • Commentary is checked automatically against the facts it was given: cited feature directions must match the sign of their contribution (a feature shared by both players may be called neutral), every number and move it mentions must appear in the input, if it says which player moved that must be the player who did, and a forced win must be reported when the search found one and never claimed when it did not. The result of that check is shown next to the text, opened when a check failed, and stored with the call in the audit log.
  • You review each commentary and record a decision: accept, edit (your edit is stored next to the original) or reject. Commentary you did not review where it appeared (for example because the game moved on) stays "awaiting review" and can be decided later on /ai-log, where the stored check is shown above the decision buttons.
  • Every call, successful or not, is recorded in the AI audit log at /ai-log with its timestamp, feature, provider, model, exact input, output (including the raw reply when it failed validation, was cut off or was refused), latency, token usage when the provider reports it, the grounding check result, and your decision, so the record shows when commentary was accepted despite a failed check. Automated evaluation calls (the LLM player) are marked "n/a". The log stays in your browser and can be exported as JSON or CSV, or cleared.

Known limitations

  • A model can still misdescribe the facts in ways the automatic check does not catch (for example, emphasis or causal language). The check is a guard, not a guarantee; that is why the human decision exists.
  • LLM-as-a-player results from a handful of games have wide intervals and depend on the model, the prompt and the day. Treat them as a demonstration of the evaluation method, not as a ranking of models.

Open the AI audit log to see every call made from this browser.