Pogofish

Pogo is played on nine squares. Each side has six pieces, stacked two by two on its back row; you move one, two or three pieces taken from the top of a stack, and a stack belongs to whoever’s piece is on top. A player with no stack left to call their own has lost. That is all, and it is a game almost everyone has forgotten.

The question I wanted to ask was simple, almost naive. Does a machine that learns alone, by playing against itself, really learn anything? And above all: how would I know?

What was measured

I had already answered that question once, with a great deal of confidence. More on that below. For the second round, I did what I should have done in the first: write down before training what would count as success, then forbid myself from touching it.

Four criteria, then, fixed in advance. One: the final AI beats a player who moves at random, a “greedy” player (it wins when it can, otherwise it takes as many stacks as possible), a more classical learning agent, and its own version at 10 % of training, each time with the lower bound of the confidence interval above 50 %. Two: its progress over training rises and then levels off, rather than oscillating. Three: the same thing happens when the whole run is repeated with three different random seeds. Four: I play at least ten games against it myself and write down what surprises me.

Each match is played over 200 randomly drawn openings the AI has never seen, each played twice with colours swapped: 400 games. The AI is an AlphaZero written in Rust, trained on 10,000 games against itself for each seed.

The rule had to change along the way. With no limit, the trained AIs learned to stall: hold, go round in circles, never conclude. The switch rule, also written in advance, fired, and the game is now played in a variant where a position that comes back for the third time loses for whoever produced it. I was not beyond reproach either: I had written that switch rule, but not wired it in, and two training runs carried on past the point where they should have stopped. The code stops itself now.

What broke in the first round

In the spring I had published a long article on this same project, with a table of verdicts, a solver that answered whether the first player wins, and one AI beating another. Everything seemed to work, which should have set off an alarm.

Reread to prepare the second round, the code turned up nine defects. A few are enough to set the tone. The rival agent my AI “beat” had never been evaluated, since the program loaded its weights into the wrong network and silently dropped the ones that did not fit: I was making it play with a random head. The evaluation replayed the same game over and over, fixed start and no noise, so my “game balance” measured a single game. The filter that decided whether to keep a new version of the AI played its games at weighted random, when it asked for them to be almost deterministic, and it had a 21.7 % chance of adopting a version of equal strength. And the solver’s table stored as exact some values that were only bounds, which invalidates its answers.

Put another way, no first-round claim about the AI’s strength holds. The rules of the game, for their part, were right. That is something.

The results

The first criterion is met, on all three seeds. The final AI’s score over 400 games, with its 95 % confidence interval:

OpponentSeeds 1, 2 and 3
random1.000 [1.000, 1.000]
1.000 [1.000, 1.000]
1.000 [1.000, 1.000]
greedy0.855 [0.823, 0.887]
0.887 [0.858, 0.917]
0.863 [0.831, 0.894]
the classical agent0.785 [0.750, 0.820]
0.738 [0.703, 0.772]
0.743 [0.708, 0.777]
itself at 10 %0.792 [0.758, 0.827]
0.790 [0.756, 0.824]
0.772 [0.737, 0.808]

So is the third: against greedy, the three seeds land between 0.855 and 0.887.

The second is not, and I am leaving it that way. The curve does level off, but the test of rising required the last version’s interval to sit entirely above the first one’s. The last version rates 153 to 175 Elo points higher, and yet the two intervals overlap by 15 to 27 points. I have an explanation, which I think is a good one: each interval is measured against the random player, a long way off, and carries all of that distance’s uncertainty. But I found it after seeing the result, and an explanation found afterwards does not change a verdict set beforehand. Not met as pre-registered.

The fourth is pending. To be completely honest, I had started: I lost every game I played at the normal level. So we measured what the same AI is worth when it thinks for longer. At 400 simulations per move against itself at 100, it scores 0.640 [0.596, 0.684]. More interesting for you: it is deterministic. A line that beats it once beats it every time, and a win replayed from memory says nothing about the player.

All these figures come from a machine in the cloud, not from my computer. The repository holds the criteria as they were written before the training runs, and the raw results.

Your move

The opponent below is the seed-1 network at the end of training. At the normal level, it is the one that scores 0.855 [0.823, 0.887] against greedy; the easy and hard levels were not measured against it. One piece moves one square, two pieces move two squares, three pieces move one or three squares; a position repeated three times loses for whoever produced it. White starts. Play both colours: a single game says very little.

The board runs in your browser; the network loads on the AI’s first move.

The same network also plays in a terminal: the program to download, for macOS only (Apple or Intel chip). It is not signed, and macOS will block it the first time it runs; you have to allow it in the settings. That is where I still owe my ten games.

If you beat it with nobody’s help, write to me. I would like to know how it is done.

Elenchus read this story

Reasoning strong

The author critically evaluates their own AI experiment, acknowledging earlier methodological flaws and presenting new results while strictly adhering to pre‑registered success criteria.

What holds Pre‑registers success criteria before training and adheres to them.

My newsletter

In French: the books I read, the exhibitions I see, the AI I practise, the biases that catch me — and right now, a writer’s apprenticeship in public.