In 1988, a group of researchers at IBM’s Thomas J. Watson Research Center made a choice that the linguistics establishment found somewhere between naïve and offensive. Instead of trying to encode the rules of language — the grammars, the idioms, the centuries of accumulated theory — they decided to ignore almost all of it. They would train a machine to translate French into English by doing nothing more principled than counting. Counting words, counting pairs, counting how often one word in one language appeared next to a particular word in another. The resulting system, known collectively as the IBM Models, didn’t understand French. It had never heard a sentence spoken aloud. It had no concept of a noun phrase or a subordinate clause. And it translated better than anything that had come before it.
The paper, “A Statistical Approach to Machine Translation” by Peter Brown, John Cocke, Stephen Della Pietra, Vincent Della Pietra, Fredjelka Jelinek, John Lafferty, Robert Mercer, and Paul Roossin, landed in Computational Linguistics in 1990 and immediately polarized the field. Noam Chomsky’s influence was still enormous. The dominant assumption held that language required structured, symbolic representation — that a machine would have to learn something like the deep grammar humans carry in their heads before it could do anything useful with words. The IBM team’s response was essentially empirical: here are the outputs, here are the scores, here is what actually works.
What made the IBM Models technically beautiful was their explicit probabilistic framing. Each model in the series (IBM 1 through 5, with successors developed afterward) added more structure to the same core idea: given a source sentence, what is the probability distribution over all possible translations? Model 1 treated word alignment as completely unconstrained, every target word equally likely to have come from any source word. Model 2 added position dependence. Model 3 introduced fertility, the idea that one source word might produce zero, one, or several target words. Each addition made the math harder and the translations better. The whole framework was trained on parallel corpora, aligned sentence pairs from sources like the Canadian Hansard, the bilingual proceedings of the Canadian parliament. Millions of sentences. No linguist in the loop.
The reverberations were slow at first, then sudden. Throughout the 1990s, statistical methods spread through the natural language processing community. Phrase-based statistical machine translation, which extended the IBM word-alignment intuition to chunks of text rather than individual words, became the dominant paradigm by the early 2000s. Google Translate launched in 2006 on a statistical foundation directly descended from the Watson lab’s work. The basic posture — trust the data, build probabilistic models, scale up the corpus — was already becoming a philosophy, not just a technique.
The deepest contribution of the IBM Models wasn’t any particular equation. It was the demonstration that language, the thing philosophers and linguists had treated as categorically special, was amenable to the same empirical, statistical treatment you might apply to a physics experiment or a financial time series. Once that was demonstrated convincingly, it became very hard to argue that the next domain — parsing, question answering, summarization, reasoning — would require a completely different approach. The door was open, and the people who walked through it eventually built the transformer.
We are now training models with hundreds of billions of parameters on trillions of tokens, and the philosophical lineage runs directly back to a group of mathematicians in Westchester County who decided that counting was enough. It turns out counting, done at sufficient scale with sufficient care, is a surprisingly deep thing to do. We’re still discovering how deep.