In 1994, a small group of protein scientists invented a competition nobody expected to matter. They called it CASP — Critical Assessment of Structure Prediction — and the premise was almost aggressively humble: every two years, labs around the world would try to predict the three-dimensional shape of proteins from their amino acid sequences alone, and the results would be scored against structures determined by the slow, expensive grind of X-ray crystallography. The best teams might shave a few points off their error scores. Progress would be measured in fractions. The whole enterprise felt like a long-running seminar on how hard the problem was.
What CASP actually was, in retrospect, was a thirty-year countdown clock.
The protein folding problem had been open since Christian Anfinsen’s Nobel Prize work in the early 1970s showed that a protein’s sequence contains all the information needed to determine its shape. Knowing that was true and being able to compute it were entirely different things. The conformational search space is astronomically large — Cyrus Levinthal famously pointed out in 1969 that a protein couldn’t possibly find its native fold by random sampling, because the universe isn’t old enough. For decades, the field threw physics-based simulations, energy minimization, and fragment assembly at the problem. CASP scores crept upward. Slowly.
Then the 2018 competition happened. DeepMind entered a system called AlphaFold, and it didn’t just win — it won by a margin that made the second-place team look like it was working in a different century. AlphaFold’s median score on the hardest target category was around 58 on the GDT scale, where anything above 90 is considered comparable to experimental accuracy. The field was stunned. Two years later, at CASP14, AlphaFold 2 posted a median GDT score above 92. The problem, for practical purposes, was solved.
What makes this history so instructive isn’t just the size of the leap. It’s what CASP reveals about how scientific fields develop readiness for an AI breakthrough. The competition did something that most domains don’t have: it created a clean, standardized benchmark with ground truth labels, refreshed every two years with genuinely new targets, resistant to overfitting in a way that many modern AI benchmarks are not. It also created a community that was disciplined about what “solved” meant. There was no ambiguity, no room for a team to claim victory on a cherry-picked example. When AlphaFold 2 showed up, everyone recognized it immediately.
That infrastructure mattered as much as the algorithm. The structural biology community had spent decades depositing experimentally determined structures into the Protein Data Bank, which by 2020 held nearly 170,000 entries. Those structures became AlphaFold’s training signal. The crystallographers who spent careers solving structures one painstaking crystal at a time were, without knowing it, building the dataset that would eventually make their primary technique feel optional for many applications.
The trajectory from CASP’s founding to AlphaFold 2 is a template worth studying. Thirty years of careful measurement, honest scoring, and community infrastructure created the conditions for a moment of radical discontinuity. The lesson isn’t that progress is always slow and then suddenly fast — though that pattern does recur. The lesson is that the benchmarks and datasets a field builds in obscurity are often the most consequential investments it ever makes.
Right now, analogous competitions exist for RNA structure prediction, protein-protein interactions, and molecular dynamics. Some of them have been running for years without a dramatic winner. Given what we now know about how these curves behave, that’s not a sign of futility. It looks more like the early rounds of another countdown, and the finish line tends to arrive faster than anyone expects.