There is a version of recent history where artificial intelligence appeared out of nowhere around 2022, surprised everyone, and became the defining technology of the decade overnight. That version is comfortable and widely believed and completely wrong. The ideas behind what we now call AI are older than video games. Older than the internet. Older than the programming languages used to build the internet. They are, roughly speaking, as old as the computer itself — because the moment we built a machine that could compute, the first thing some people asked was whether it could think.
What follows is the actual timeline. Not a history of hardware — we've covered that elsewhere and will again. This is a history of compute as concept. Of what we knew was possible, how early we knew it, and the long strange road between knowing and doing.
I
Before Anyone Called It AI
In 1943, a neurophysiologist named Warren McCulloch and a self-taught logician named Walter Pitts published a paper called "A Logical Calculus of Ideas Immanent in Nervous Activity." It proposed the first mathematical model of a neural network — an artificial neuron that took inputs, applied weights, and fired a binary output. It was abstract. It ran on paper, not silicon. And it demonstrated that networks of these simple units could, in principle, compute any logical function. The idea that a machine could simulate the structure of thought was born thirteen years before anyone coined the term artificial intelligence.
Let that land for a second. The concept of an artificial neural network predates the first video game by nearly a decade. OXO, generally considered one of the earliest computer games, appeared in 1952. Tennis for Two hit an oscilloscope screen at Brookhaven National Laboratory in 1958. The neural network model came first. AI as a concept is older than gaming as a concept. That feels impossible until you realize what was actually happening in the 1940s and 1950s — people were not building computers and then slowly wondering what to do with them. They were building computers because they already knew what they wanted to do.
Alan Turing published "Computing Machinery and Intelligence" in 1950. The paper opens with five words that still haven't been fully answered: Can machines think? Turing didn't just ask the question — he designed a test for it. The Imitation Game, now called the Turing Test, proposed that if a human interrogator couldn't reliably distinguish a machine's written responses from a person's, the machine should be considered intelligent. It was a framework for a problem that wouldn't be technically approachable for another seventy years. He knew that. He wrote it anyway. Because the idea was already fully formed. Only the hardware was missing.
II
The Naming and the Dream
The term "artificial intelligence" was coined at a specific time and place: the Dartmouth Summer Research Project, June 1956, organized by John McCarthy, a young math professor at Dartmouth College. He brought together a small group — Marvin Minsky, Claude Shannon, Nathaniel Rochester, and others — for an eight-week workshop to explore whether machines could simulate human intelligence. The proposal stated plainly that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it." They believed this. They believed it could be solved in a few summers.
They were right about the principle and catastrophically wrong about the timeline. The hardware of 1956 was nowhere near capable of running what these ideas required. But the ideas were complete. The research community, the vocabulary, and the ambition all crystallized that summer. Everything that followed — every breakthrough, every winter, every hype cycle, every frontier model — traces a direct line back to that room in New Hampshire.
Within two years, Frank Rosenblatt built the Perceptron — the first neural network that could learn from data. His 1958 creation at Cornell could take inputs, adjust its own weights through a training process, and classify patterns. The press went wild. Headlines called it an electronic brain. The Navy funded it, imagining it could walk, talk, and reproduce itself. Rosenblatt was more measured than the headlines, but the gap between what the Perceptron could actually do and what people expected it to do was already enormous. That gap would define the next two decades.
III
The Winters
What happened next is the part most people skip, and it's the part that matters most for understanding where we are now. AI didn't follow a clean upward trajectory from the 1950s to today. It crashed. Twice. Hard enough that the term "AI Winter" was coined — deliberately referencing nuclear winter. The field didn't just slow down. It froze.
The first winter started in the early 1970s. By then, a decade of lavish DARPA funding had produced impressive demos and bold promises but limited practical results. The Mansfield Amendment of 1969 required military funding to have clear practical applications, and AI research didn't qualify. Then in 1973, the British government commissioned Sir James Lighthill to evaluate the state of AI research. His report was devastating — he concluded that AI had failed to deliver on its ambitious goals and that the problem of "combinatorial explosion" represented a fundamental barrier. The UK slashed funding. British AI labs shut down. Researchers scattered. Across the Atlantic, DARPA pulled back too. The chill lasted roughly from 1974 through the early 1980s.
But the real damage had already been done a few years earlier. In 1969, Marvin Minsky and Seymour Papert — two of the most influential figures in AI — published Perceptrons, a mathematical critique of Rosenblatt's neural network. They proved that a single-layer perceptron couldn't solve certain basic problems. The book was technically correct about single-layer networks. But it was interpreted as a death sentence for the entire concept of neural networks. Funding dried up. Researchers abandoned the approach. Rosenblatt himself died in a boating accident in 1971, at forty-three, and couldn't defend or extend his own work. Neural networks went dark for over a decade.
The field clawed back in the early 1980s with Expert Systems — software that encoded human knowledge as if-then rules to solve specific problems. Companies like Symbolics sold specialized hardware called Lisp Machines for six figures apiece. Japan launched a ten-year Fifth Generation Computer project. Over a billion dollars flowed back into AI. Wall Street adopted expert systems for trading. For a few years, it looked like the winter was over.
It wasn't. By 1987, the Lisp Machine market collapsed. General-purpose workstations from Sun and Apple could run the same software for a fraction of the cost. Expert systems turned out to be expensive, rigid, and terrible at handling anything outside their narrow rule sets. Japan's Fifth Generation project ended without meeting its goals. Symbolics — the company that registered the first .com domain in 1985 — filed for bankruptcy. The second AI winter set in and lasted into the early 1990s.
Here's the detail that matters: the second winter was so harsh that researchers stopped using the word "AI" entirely. The field survived by rebranding. The same techniques, the same mathematics, the same work — but called "informatics" or "analytics" or "machine learning." The ideas didn't die. The terminology did. AI collected so much baggage from two cycles of broken promises that the name itself became toxic. If you were doing AI work in the 1990s and wanted funding, you called it something else.
IV
The Quiet Revival
While the label was in hiding, the work continued. In 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams published a paper that would eventually reshape everything: "Learning Representations by Back-propagating Errors." Backpropagation — a method for training multi-layer neural networks by sending error corrections backward through the layers — wasn't technically new. It had been formulated in various forms before. But this paper demonstrated it working on real problems, and it solved exactly the limitation Minsky and Papert had identified seventeen years earlier. You could train networks with multiple layers. You just needed the right algorithm and enough compute.
The "enough compute" part is the whole story. Backpropagation worked in 1986, but computers were still too slow to train networks of meaningful size. Training a big neural network could take months. The excitement faded again — not into a full winter, but into a long, quiet period where the mathematics was sound, the theory was proven, and the hardware simply wasn't there yet.
Meanwhile, AI was doing real work under its new names. By the early 1990s, neural networks were running in banking — credit scoring, fraud detection, risk assessment. In 1991, the first academic literature appeared using neural networks for credit classification models. By mid-decade, these techniques were embedded in financial infrastructure, processing transactions and flagging anomalies. Nobody called it AI. They called it analytics. It was the same math. The same idea McCulloch and Pitts had sketched on paper in 1943, now running at scale in production systems that moved money.
Then came May 11, 1997. IBM's Deep Blue defeated world chess champion Garry Kasparov in a six-game match — the first time a computer beat a reigning world champion under standard tournament rules. Deep Blue could evaluate 200 million chess positions per second. It wasn't neural-network AI — it was brute-force computation. But the image of a machine beating the greatest chess mind alive hit the public consciousness like a freight train. IBM's market cap jumped $18 billion overnight. The word "AI" started becoming speakable again.
V
The Hardware Finally Arrives
The missing piece had always been compute. The ideas were seventy years old. The algorithms were proven. The data was starting to accumulate. What nobody had, until the mid-2000s, was a machine architecture that could train neural networks fast enough to matter.
In November 2006, NVIDIA released CUDA alongside its G80 architecture — the GeForce 8800 GTX. CUDA turned the GPU from a graphics-rendering device into a general-purpose parallel computing platform. For the first time, researchers could write standard code that ran across hundreds of processing cores simultaneously. This mattered because neural network training is fundamentally a massive parallel math problem — millions of matrix multiplications happening in concert. GPUs were accidentally perfect for it. The hardware that gamers used to render better explosions turned out to be the hardware that AI had been waiting sixty years for.
Between 2006 and 2012, a small group of researchers — many of them students of Geoffrey Hinton at the University of Toronto — began using CUDA-enabled GPUs to train neural networks. The results were promising but incremental. The rest of the machine learning community remained skeptical. Then came September 30, 2012.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton entered AlexNet into the ImageNet Large Scale Visual Recognition Challenge — an annual competition to classify images from a dataset of over a million photographs. AlexNet achieved a top-5 error rate of 15.3%. The second-place entry, using traditional methods, scored 26.2%. That's not an incremental improvement. That's a demolition. AlexNet was trained on two NVIDIA GTX 580 GPUs. The deep learning era started on gaming hardware.
The AI community pivoted overnight. Google, Facebook, and every major tech company recognized what had just happened: neural networks, trained on large datasets with GPU compute, could outperform every other approach to pattern recognition. The sixty-year bet placed by McCulloch and Pitts, by Rosenblatt, by Hinton during the decades when nobody was funding neural network research — that bet had just paid off.
In June 2015, a programmer named SethBling released MarI/O — a neural network that taught itself to play Super Mario World using a technique called NEAT (NeuroEvolution of Augmenting Topologies). It started knowing nothing about the game. No rules, no controls, no concept of what a Goomba was. Through evolutionary learning — randomly generating networks, keeping the ones that performed best, mutating and combining them — MarI/O completed an entire level in 34 generations. It was a demonstration you could hand to anyone, gamer or not, and they would immediately understand what they were seeing: a machine learning from scratch, in real time, through trial and error. It made the abstract concrete.
VI
The Architecture That Changed Everything
In June 2017, a team at Google published a paper with a title that reads like a thesis statement: "Attention Is All You Need." It introduced the Transformer architecture — a neural network design that replaced the sequential processing of earlier models with a mechanism called self-attention, where every element in a sequence could attend to every other element simultaneously. This made training massively parallelizable. It scaled with hardware. And it turned out to be the architecture that would power virtually every major AI system built afterward.
OpenAI ran with it. GPT-1 arrived in June 2018 with 117 million parameters, trained on a dataset of books. It was a proof of concept — pre-train a Transformer on a massive amount of text, then fine-tune it for specific tasks. GPT-2 followed in February 2019 with 1.5 billion parameters. GPT-3 landed in June 2020 with 175 billion parameters and could write essays, generate code, and learn tasks from a few examples placed in the prompt. Each version was essentially the same idea, scaled up — more parameters, more data, more compute.
On November 30, 2022, OpenAI released ChatGPT — a conversational interface wrapped around a GPT-3.5 model fine-tuned with reinforcement learning from human feedback. It reached one million users in five days. One hundred million in two months. It became the fastest-growing consumer application in history. And suddenly, the word "AI" was everywhere again — but this time, the technology behind it was real, it was running at scale, and it was doing things that no one outside the research community had expected to see in their lifetime.
The public response was approximately: Where did this come from?
It came from 1943. It came from McCulloch and Pitts sketching neurons on paper. From Turing asking if machines could think. From Rosenblatt building a learning machine and dying before the world caught up to what he'd built. From Hinton keeping the faith through two winters when neural networks were considered a dead end. From NVIDIA accidentally building the hardware AI needed while trying to make better video games. From seventy-nine years of people knowing exactly what was possible and waiting for the technology to catch up to the idea.
AI didn't pop up a couple of years ago. It popped up around the time compute itself popped up — because the possibility was visible from the start. Every advancement since has been the hardware slowly, painfully catching up to concepts that were already fully formed before the first person ever played a video game.