- Decision: re-implement the original Python (referee rules, Part A A*, Part B agent) line by line in TypeScript, run all search in Web Workers in the browser, and prove equivalence with parity fixtures generated by the unchanged Python.
- Status: accepted on revival of the project (2026, Rin). Recorded on 6 October 2026.
- Supersedes: nothing. Superseded by: nothing yet.
Context
The 2022 code ran only from a terminal against the subject's referee. To make it something people can play, inspect and evaluate, it needed to run on a website. The site should be static (no server to pay for or secure), stay responsive while the agent thinks, and keep the original results faithful: a port that quietly plays differently would make every number on the site meaningless.
Decision
- Port each module with the original's structure and quirks preserved (for example
game_endignoring STEAL, the tie-break bias, and the inverted beta update described in DR-002), not a cleaned-up rewrite. - Generate fixtures with
scripts/generate_parity_fixtures.py, which imports the unchanged Python and records its outputs, and assert in Vitest that the port reproduces them exactly. - Where the original depends on CPython internals, simulate them: A* iterates neighbours out of a
set, sopython-set-order.tsreproduces CPython 3.12's tuple hashing and set probing to get the same tie-breaking and node counts. - For randomness, the port uses a seeded generator (mulberry32). Parity tests run the Python with the move shuffle replaced by a canonical sort and the bias fixed at 1, which removes the only difference.
- Run the agent, the A* study and the tournament in Web Workers, with a main-thread fallback.
Options considered
- Run the Python in the browser with Pyodide. Byte-for-byte faithful, but a large download, slow start, and still single-threaded and slow (deep copies at every node).
- Host the Python behind an API. Faithful, but needs a server, adds latency and cost, and makes the static, inspectable site impossible.
- Port to TypeScript with parity tests (chosen).
Why
A port gives speed (the tournament plays 1,200 games in about three minutes on a laptop), a static deployment, and code readers can step through. The risk of a port is silent divergence, and parity fixtures turn that risk into a test: if the port differs from the original on any recorded case, CI fails.
What happened
- The port reproduces the original exactly on everything recorded: paths and node counts for 184 A* runs, every board state and capture across 40 random refereed games, evaluation features on 237 positions, 179 minimax searches, 143 agent moves and 5 full self-play games.
- The tournament harness uses the port because the original is far too slow for it: in Python one 6 × 6 self-play game at fixed depth 2 takes about 40 seconds. To check the harness itself,
scripts/crosscheck_tournament.pyreplays the pairings that are feasible in Python (boards 4 and 5, 480 games) from the unchanged code. All six pairings agree within sampling error (differences in win rate from −8.8 to +3.7 percentage points, every Newcombe 95% interval containing 0). With 80 games per pairing that check can only rule out large discrepancies, about ±15 points. - Faithfulness has a cost: the site plays the original's bugs too (the alpha-beta beta update, rim-favouring positional weights). That is deliberate; they are documented instead of fixed.
- Random play cannot be compared game for game across languages, because the random streams differ. Comparisons are distributional.
What I'd change
- Record the original Python's random draws (or patch its RNG to mulberry32) so whole random games, not just deterministic ones, can be compared move for move.
- Add a property-based test layer (random positions, compare rules and evaluation against the Python on the fly in CI) on top of the fixed fixtures.
- Keep any improved agent strictly separate from the "as submitted" port, as a named variant with its own tests and its own row in the tournament.