Methods
How every number on this site was made
Provenance, methods and evaluation design for the tournament, the A* study and the alpha-beta study; the assumptions and limitations behind them; and the decisions, agent card and AI use statement that go with them.
Where the data comes from
Data provenance
There is no external dataset. Every result is generated by running code in this repository with a fixed seed, and the generating script is named next to it. The original 2022 Python in coursework/ is never modified; the TypeScript port is checked against it by parity tests (DR-003).
| Artefact | Produced by | Script | Run details |
|---|---|---|---|
| Original benchmark: agent vs random, 160 games | Original Python, unchanged, via the subject's referee | scripts/benchmark_agent.py | Python 3.12.7, PYTHONHASHSEED=0 |
| Parity fixtures (A*, rules, evaluation, minimax, moves) | Original Python, unchanged; shuffle replaced by a sort and bias fixed at 1 | scripts/generate_parity_fixtures.py | CPython 3.12 |
| Reference round robin, 1,200 games | TypeScript port (parity-tested against the original Python) | web/scripts/generate-reference-studies.ts | seed 2022; Apple M4; Node v26.10.0; 2026-10-05 |
| Python cross-check, 480 games | original Python (coursework/Project Part B), variants by patching one knob | scripts/crosscheck_tournament.py | Python 3.12.7, PYTHONHASHSEED=0; seed 2022 |
| A* paired heuristic study, 980 boards | TypeScript port (parity-tested against the original Python) | web/scripts/generate-reference-studies.ts | seed 2022; the notebook's board generator |
| Alpha-beta efficiency, 480 searches | TypeScript port (parity-tested against the original Python) | web/scripts/generate-reference-studies.ts | seed 2022; positions from greedy-vs-random games |
What was done
Method
Agents and variants
dynamic_depth_allocation and the opening-book flag. Random is the subject's baseline.Round robin
Strength and uncertainty
A* heuristic study
Alpha-beta study
Statistical code
web/src/lib/stats with unit tests against reference values from scipy 1.17, statsmodels and R (Wilson intervals, Newcombe differences, Wilcoxon exact and approximate p-values, type-7 quantiles, Bradley-Terry maximum likelihood). Every interval shown records its seed and number of resamples.What counts as evidence
Evaluation design
- Questions asked: how strong is the original relative to controlled variants and a baseline; does dynamic depth help; does the A* heuristic choice trade speed for optimality; how much work does pruning save; and (with your own key) how does a language model compare to the random and greedy agents on identical seeds.
- Sample sizes are stated with every result. The reference round robin has 40 games per pairing per board size (120 per pairing overall), which resolves differences of roughly 15 to 20 percentage points in a head-to-head win rate; smaller differences are reported as not detectable rather than as ties.
- Paired comparisons where two methods are compared: heuristics on the same boards (paired bootstrap and Wilcoxon for node counts, paired bootstrap and exact McNemar for shortest-path rates), search variants on the same positions, colour-swapped game pairs on the same seeds, agents compared through Elo differences taken from the same bootstrap refits (never by whether two intervals overlap), and LLM versus each baseline on exactly the games the model finished (games won by only one of them, with an exact McNemar test).
- LLM answers are scored per turn, one rule for both providers: the rate reported is the share of turns whose first answer was rejected, split into illegal moves and replies with no usable move (malformed, cut off or refused), because retries on the same turn are not independent trials. Turns within a game are not independent either (same model, prompt and position history), so the Wilson interval is on the effective number of turns after a design effect estimated from how much the rate varies between games, never below 1 (DR-005). With a handful of games the design effect is itself rough.
- Effect sizes, not just p-values: differences in win rate and Elo with intervals, mean node differences, rank-biserial r and d_z.
- No self-scored metrics. Results are measured against the referee and against other agents, never against a score the author chose.
- Multiple comparisons: ten pairings are reported side by side without correction. Each interval is valid on its own; reading the table as a whole, expect about one in twenty 95% intervals to miss.
Read before you quote a number
Assumptions and limitations
Assumptions
- Games are independent given their seeds; colour-swapped pairs share a seed by design, and the bootstrap resamples games, not pairs.
- Bradley-Terry assumes strength is one number per agent (transitive, no style matchups) and pools board sizes in the “all sizes” view.
- The TypeScript port behaves like the original in random play as well as in the deterministic cases the parity tests cover; the Python cross-check supports this within its resolution.
Limitations
- Board sizes 4 to 6 only for the round robin; the original supports up to 15 × 15, where depth 1 is likely to be weaker still.
- Move times depend on hardware and are only comparable within one run.
- The Python cross-check has 80 games per pairing and can only detect large discrepancies (about ±15 points).
- The A* study uses the notebook's random boards (barriers at random, up to half the board); other board distributions may behave differently.
- LLM-as-a-player runs are a handful of games against one opponent, with one prompt; intervals are wide by design.
Looking back
What I'd change
- Fix the one-line beta update and replace ratio-based depth with iterative deepening under a time budget, then test the change as a new variant against the original (DR-002).
- Add a connection-distance feature and tune weights against data, not by hand (DR-001).
- Pre-register comparisons and sample sizes before running the next tournament, and extend it to larger boards.
- Record the original Python's random draws so whole random games can be compared move for move across languages (DR-003).
Decision records
Decisions, as made and as measured
Each record states the decision first, then context, options, reasons, what happened (weak numbers included) and what I would change. Records are never edited after the fact; a new record supersedes an old one.
- DR-001Evaluation function features and weightsRead
- DR-002Dynamic depth allocation for minimaxRead
- DR-003Port to TypeScript and run in Web Workers, verified by parity testsRead
- DR-004Compare two methods by their paired differenceRead
- DR-005Report uncertainty at the unit that was sampledRead
- DR-006Ground commentary in what the search found, and keep the check with the decisionRead
- Agent cardThe _4399 agent: intended use, evaluation with intervals, known weaknessesRead
Transparency
AI use statement
This site has two optional features that call a large language model. Everything else, including the agent you play against and every statistic on the site, is ordinary code that runs without any AI service. This statement is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act, and the NIST AI Risk Management Framework. It is not a claim of compliance or certification with any of them.
What the AI features do
- Move commentator (on
/playand/spectate). When the minimax agent has searched for a move, you can ask a model to explain it in plain English. The model receives only the agent's own numbers for that move: board size, turn, the move, search depth, nodes visited, the top candidate scores and the six evaluation features with their weights and contributions, plus a sentence saying what the score means at that depth: the features describe the position right after the move, so they explain the score only for a one-ply search, and a forced win found by a deeper search is the reason for the move, not the features. It is instructed to use nothing else. - LLM as a player (on
/llm-arena). A model plays a few short games against the original agent. Each turn it receives the rules, the board, the move history and the list of legal moves, and returns one move as JSON. Answers are scored the same way for both providers, under the same output-token budget: an illegal move, or a reply with no usable move (malformed, cut off at the token limit, or refused), is rejected and asked again, up to three times. The headline is the share of turns whose first answer was rejected, reported next to random, greedy and scripted baselines on the same seeds.
What they never do
- They never choose the agent's moves, change its evaluation, or feed into the tournament, the A* study or any other statistic on the site.
- They never run unless you add your own API key and press a button.
- They never send anything to this site's operators. The site is static and has no server; calls go directly from your browser to the provider you chose.
- They never store your key anywhere except your own browser (session storage by default, local storage only if you tick "remember on this device"), and never write it to the audit log or to any file in the repository.
Data sent to the provider
Only the prompt shown in the audit log for that call: game state, rules, legal moves and, for the commentator, the agent's evaluation breakdown. No personal information is collected or sent. Your API key goes in the request header to the provider (Anthropic or OpenAI), which processes the request under its own terms and bills your account.
Human in the loop and transparency
- Every model output on the site is labelled AI-generated with the model name.
- Commentary is checked automatically against the facts it was given: cited feature directions must match the sign of their contribution (a feature shared by both players may be called neutral), every number and move it mentions must appear in the input, if it says which player moved that must be the player who did, and a forced win must be reported when the search found one and never claimed when it did not. The result of that check is shown next to the text, opened when a check failed, and stored with the call in the audit log.
- You review each commentary and record a decision: accept, edit (your edit is stored next to
the original) or reject. Commentary you did not review where it appeared (for example because
the game moved on) stays "awaiting review" and can be decided later on
/ai-log, where the stored check is shown above the decision buttons. - Every call, successful or not, is recorded in the AI audit log at
/ai-logwith its timestamp, feature, provider, model, exact input, output (including the raw reply when it failed validation, was cut off or was refused), latency, token usage when the provider reports it, the grounding check result, and your decision, so the record shows when commentary was accepted despite a failed check. Automated evaluation calls (the LLM player) are marked "n/a". The log stays in your browser and can be exported as JSON or CSV, or cleared.
Known limitations
- A model can still misdescribe the facts in ways the automatic check does not catch (for example, emphasis or causal language). The check is a guard, not a guarantee; that is why the human decision exists.
- LLM-as-a-player results from a handful of games have wide intervals and depend on the model, the prompt and the day. Treat them as a demonstration of the evaluation method, not as a ranking of models.
Open the AI audit log to see every call made from this browser.