Guided tour
Cachex Arena in three short walkthroughs
Each video follows one workflow from start to finish, with the step shown on screen and as captions. They were recorded from this site by a Playwright script that also checks every step (the steal, both captures, the recorded A* output, the seeded tournament), so the same seeds reproduce the same moves and numbers.
Walkthrough 1 of 3 · /play
Play the agent
A short game on 5 × 5 against the ported minimax agent: the opening steal, a capture, a recapture, and the agent's own explanation of its move.
Steps (transcript)
- 1Open Play: a 5 × 5 board against the minimax agent (seed 4399), coordinates on
- 2Red opens on the strong cell (1, 1)
- 3Blue steals: Red's opening tile is mirrored across the diagonal and becomes Blue's
- 4Red plays (0, 1); the agent answers at (4, 1)
- 5Red plays (1, 0), leaving two red tiles inside a diamond
- 6Capture: Blue closes the diamond at (0, 0) and removes both red tiles
- 7Red retakes (0, 1); the agent plays (1, 4)
- 8Recapture: Red plays (1, 0) and removes Blue's (0, 0) and (1, 1)
- 9Why that move? Pick the capture in the move log: search depth, candidate scores, features
Walkthrough 2 of 3 · /astar
A* Lab
The original Part A sample input, searched with both heuristics and animated expansion by expansion, then the paired study that compares them on 980 boards.
Steps (transcript)
- 1Open the A* Lab: the Part A search, ported line by line from Python
- 2Load the original sample input (code/sample_input.json)
- 3Manhattan, as in the recorded output: animate the expansions
- 4An 8-cell path that matches sample_output.txt from the original repo
- 5Switch to Euclidean and animate the same board
- 6Same board, both heuristics: path cost, nodes expanded, queue pushes
- 7Paired study on 980 boards: paired bootstrap CI and Wilcoxon test on expansions
- 8Optimality against breadth-first search: paired difference and exact McNemar test
Walkthrough 3 of 3 · /tournament
Tournament and LLM evaluation
A small seeded round robin with interval estimates, then the bring-your-own-key settings and the LLM-as-a-player harness, shown with a mocked model reply (no real key is used).
Mocked AI response for illustration. Steps 6 to 10 use a placeholder key and the model id mock-for-illustration. Requests to the provider are intercepted in the browser and answered by a mock that plays the first legal cell; no model was called, so the LLM row shows the mock, not a real model's results.
Steps (transcript)
- 1Open the tournament harness: a seeded, colour-swapped round robin
- 2Keep 4 agents and seed 4399; 4 × 4 only, 2 colour-swapped pairs: 24 games
- 3Run all 24 games in parallel Web Workers
- 4Win rates with Wilson 95% CIs, Bradley-Terry strengths with bootstrap CIs
- 5AI settings: bring your own key, kept in this browser and sent only to the provider
- 6For this demo: a placeholder key and the model id “mock-for-illustration”, no real keyMocked AI response for illustration
- 7LLM Arena: 2 games on 4 × 4 against the minimax agent, on the baselines' seedsMocked AI response for illustration
- 8Every reply is checked against the legal moves and labelled AI-generatedMocked AI response for illustration
- 9Side by side with random, greedy and scripted baselines, Wilson 95% CIsMocked AI response for illustration
- 10Every call is in the AI audit log, with JSON and CSV exportMocked AI response for illustration
- 11Forget key: the placeholder is removed from this browser
Screenshots
Every key feature at a glance
Captured by the same script, in light mode at 1440 × 900 (the landing page also in dark mode) and on a 390 px phone. Select one to enlarge it; use the arrow keys to step through.
Desktop · 1440 × 900
Mobile · 390 × 844
How these were made
pnpm showcase runs web/e2e/showcase.spec.ts on the system Chrome: it plays each journey at a human pace with an on-screen caption and cursor, asserts what it shows, and records it at 1280 × 800. ffmpeg then encodes the H.264 videos here and the GIFs in the README. The captions on this page are the same text as the on-screen steps.
No real API key is used anywhere in these recordings. Where an AI feature appears, the key is a placeholder, every request to the provider is intercepted in the browser, and the reply is a labelled mock.