Skip to content
Methods & decisions

Agent card

The _4399 minimax agent

A model card for a game-playing agent with no learned parameters. Numbers are from the reference studies on the website (/tournament), which list seeds, sample sizes and how each interval was computed. Last updated 6 October 2026.

Overview

ItemDetail
What it isMinimax with alpha-beta pruning over a hand-tuned six-feature evaluation, plus a two-move opening book and an instant-win check
AuthorsTeam _4399: Sunchuangyu "Rin" Huang and Wei Zhao (COMP30024, University of Melbourne, Semester 1 2022)
VersionsOriginal Python (coursework/Project Part B/code/_4399), unchanged. TypeScript port (web/src/lib/agent), verified move for move against the original by parity tests.
Learned parametersNone. Six weights chosen by hand (weights.json). See DR-001.
Search depth1 ply while at least 15% of cells are empty, then 2, 3, 4. See DR-002.

Intended use

  • Teaching and demonstration: playing Cachex in the browser, watching agents play, and stepping through how a classical search agent scores moves.
  • A reproducible baseline for experiments on Cachex (the tournament harness, the LLM-as-a-player evaluation).

Not intended for: competitive play claims, or as a reference for strong Hex or Cachex play. It is a student project and, as measured below, a modest one.

Training data and provenance

There is no training data. The evaluation weights were set by the two authors by hand during development in 2022 ("team empirical experience" in the code); the records of how are not in the repository. The opening book (always cell (1, 1), stealing it as Blue) is likewise hand-written.

Evaluation data on this card is synthetic: games between agents generated by seeded random number generators, produced by web/scripts/generate-reference-studies.ts (TypeScript port) and scripts/benchmark_agent.py and scripts/crosscheck_tournament.py (original Python).

Evaluation

All intervals are 95%. Win rates use Wilson score intervals; strengths are Bradley-Terry fits on the Elo scale with the random agent fixed at 0, with percentile intervals from 1,000 bootstrap resamples of games within each pairing and board size (one virtual drawn game per pairing as a regulariser).

Original benchmark (Python, unchanged). Against the subject's random agent on 4 × 4 to 7 × 7, 20 seeded games per size and colour: 144 wins in 160 games, 90.0% (84.4% to 93.8%).

Reference round robin (TypeScript port). 1,200 games on 4 × 4, 5 × 5 and 6 × 6; every pairing played 40 games per size as 20 colour-swapped pairs; seed 2022.

AgentElo vs random (95% CI)Win rate over all its games (95% CI)
Fixed depth 3 (variant)490 (432 to 559)76.0% (72.0% to 79.6%)
Fixed depth 2 (variant)369 (321 to 427)58.3% (53.9% to 62.7%)
Original (dynamic depth)339 (290 to 396)53.8% (49.3% to 58.2%)
Greedy one-ply (variant)324 (274 to 386)51.5% (47.0% to 55.9%)
Random0 (reference)10.2% (7.8% to 13.2%)

Head to head, the original beat random in 106 of 120 games (88.3%, 81.4% to 92.9%), drew level with greedy one-ply (61 to 59; 50.8%, 42.0% to 59.6%) and lost to fixed depth 3 in 84 of 120 (its win rate 30.0%, 22.5% to 38.7%). Compared directly, from the same bootstrap refits (the two strengths come from one fit, so their separate intervals are correlated and overlap is not a test): the original minus greedy is +15 Elo (95% CI −24 to +56), so no difference is detectable, while fixed depth 3 minus the original is +151 Elo (106 to 195). Red, the first mover, won 50.3% of all games (95% CI 47.5% to 53.2%): no first-move advantage was detectable in these games, which all allow STEAL; a no-STEAL condition was not tested.

Cross-check. The pairings feasible in Python (boards 4 and 5, 80 games each) were re-run from the unchanged original; all six win-rate differences against the TypeScript run have Newcombe 95% intervals that contain 0. That is no detectable disagreement, not proof of agreement: with 80 games per pairing the check can only detect differences larger than about ±15 percentage points.

Search efficiency. On 120 mid-game positions, each searched at depths 2 and 3 in two move orders (480 searches), the original alpha-beta visits on average 92% to 100% of the nodes plain minimax visits, per board size, depth and move order (bootstrap intervals on /tournament); a textbook beta update would visit 17% to 45% on average. The averages hide the shape: 444 of the 480 searches visited every node, and the 36 that pruned anything visited between 0.6% and 92% of the tree. Because of the beta bug, a cut-off can only happen once the search finds a line it scores as a forced win for Red (DR-005). All 480 searches returned the same root value as plain minimax.

Known weaknesses and failure modes

  • Effectively one ply deep. The deeper search triggers on only 0.7% of searched moves (34 of 4,602), because most games end before 85% of the board is full. In practice the agent is greedy one-ply plus an opening book, and plays like it.
  • Does not block one-move wins. While at least 15% of cells are empty it searches one ply, and its instant-win check looks only for its own wins, so it never considers the opponent's next move. A scripted Blue that just plays the first empty cell in (r, q) order (filling row 0 left to right) beats it as Red in 50 of 50 seeded games on 4 × 4, 43 of 50 on 5 × 5 and 34 of 50 on 6 × 6 (seeds 0 to 49). In 124 of those 127 losses the agent had left Blue a one-move win that a single Red move could have blocked. Reproduce with cd web && pnpm test (the "first-legal-cell" tests in src/lib/ai/ai.test.ts). The LLM arena includes this scripted line as a baseline, so a model's wins can be judged against it.
  • Pruning bug. The minimising branch's if beta <= min_score: beta = min_score never narrows the window (beta stays at +infinity), so alpha-beta prunes only after finding a forced win for Red and otherwise saves nothing. Correct results, wasted work.
  • No notion of connection. None of the six features measures progress towards linking the agent's edges, which is the goal of the game. A path-distance feature exists in the code but is switched off.
  • Rim-favouring positional weights. The score matrix rates the rim above the centre, against standard Hex opening advice; untested either way.
  • Rigid opening. Always (1, 1) (or steal (1, 1)); easy to anticipate. On 3 × 3 the book plays fixed cells.
  • Tie-breaking noise. A random ×(1 + 10⁻⁵) bias picks among equal scores, so play varies between seeds even in identical positions.
  • Scale of evidence. Strength estimates come from 4 × 4 to 6 × 6 boards. Nothing here supports claims about larger boards, where depth 1 is likely to be relatively weaker still.

Ethical and practical considerations

  • Low-stakes domain: a board game, no personal data, no decisions about people.
  • Academic integrity: the original is preserved for reference; current COMP30024 students should not copy it.
  • Compute: the agent itself is light (about 0.1 ms per move in the browser). The optional LLM features use the visitor's own API key and are billed to them; the site shows token usage for every call.
  • Honest reporting: variants exist to explain the original, not to replace it. The site keeps the agent exactly as submitted and documents its flaws instead of silently fixing them.

How to reproduce

uv run scripts/benchmark_agent.py           # original Python vs random (144/160)
uv run scripts/crosscheck_tournament.py     # original Python, feasible pairings
cd web && pnpm gen:reference                # TypeScript reference studies
cd web && pnpm test                         # parity, statistics and harness tests

Source: docs/model-card.md. Live numbers with downloadable data: /tournament.