The Night a Program Beat the World’s Best Backgammon Player — and Nobody Realized What Had Just Happened

In 1979, a program called BKG 9.8 defeated Luigi Villa, the reigning world backgammon champion, in a match played in Monte Carlo. It was the first time a computer program had beaten a human world champion at any board game. The moment passed without much fanfare. Backgammon, after all, involves dice. People assumed the machine got lucky.

They were wrong about what they were looking at. And the story of how that dismissal got corrected — slowly, then all at once — is one of the cleanest windows we have into the full arc of AI capability.

Hans Berliner built BKG 9.8 at Carnegie Mellon. It used hand-crafted evaluation functions, painstakingly tuned heuristics, and a lot of Berliner’s own expert knowledge about the game baked directly into code. It was impressive engineering, but it was brittle. The program couldn’t explain its reasoning in any general sense. Move it to a slightly different problem and it would collapse. Berliner himself suspected luck had played a role in the championship win. He was being honest.

Then came TD-Gammon.

In 1992, Gerald Tesauro at IBM Research trained a neural network to play backgammon using temporal difference learning — reinforcement learning against itself, with almost no human knowledge seeded in. TD-Gammon started from random play and, through millions of self-generated games, converged on strategies that shocked human experts. It didn’t just play well. It discovered moves that top players had never considered, including a controversial opening preference that was initially dismissed and later accepted as genuinely strong. The machine had found real knowledge that humans had missed.

This is the part that deserves to stop you cold. A program in the early 1990s, running on hardware that would embarrass a modern smartwatch, was capable of genuine strategic discovery. Not retrieval of known patterns. Discovery. The gap between BKG 9.8 and TD-Gammon wasn’t mainly about compute — it was about the shift from encoding human knowledge to learning structure from experience. That conceptual move, from hand-crafted to learned representations, is the same move that would eventually produce AlexNet, AlphaGo, and everything downstream of them.

TD-Gammon also demonstrated something that took years to fully appreciate: self-play as a route to superhuman performance. You don’t need a labeled dataset of expert decisions if you can construct an environment where the signal of winning and losing is clear. The program generates its own curriculum, always playing at its own frontier. AlphaZero made this idea famous in 2017, teaching itself chess, shogi, and Go from scratch in hours. But Tesauro was running the same core experiment twenty-five years earlier with a fraction of the resources and a fraction of the attention.

What makes the backgammon story so illuminating isn’t just the technical lineage. It’s the recurring pattern of underestimation. Villa’s win was luck. TD-Gammon’s insights were probably artifacts. AlphaGo’s victory over Lee Sedol was supposed to be a curiosity limited to a narrow domain. At each step, observers found reasons to contain the implications. And at each step, the implications turned out to be larger than anyone comfortable admitted in real time.

We are in the same position now, looking at systems that reason across text, images, code, and scientific data in ways that keep outrunning the explanations we reach for. The backgammon story is a useful corrective. When a machine does something that seems like it shouldn’t be possible — and then someone proposes a mundane explanation for why it doesn’t really count — it’s worth remembering Monte Carlo in 1979, and asking whether we’re making the same mistake again.

The trajectory has never stopped accelerating. We just keep getting surprised by where it goes next.