By 2017, everyone in artificial intelligence knew that neural networks were the future. The three godfathers — Hinton, Bengio, and LeCun — had proven it. Deep learning was working. But there was a problem with how language models actually ran.
They were slow. Painfully, structurally slow.
The dominant architecture was something called an LSTM — Long Short-Term Memory network. It processed language the way humans read: left to right, word by word, sequentially. The computer would read "The," then "cat," then "sat," then "on," then "the," building up understanding as it went. This seemed natural. It's how we read. But it was terrible for training AI.
Sequential processing meant you couldn't parallelize. You couldn't read the whole sentence at once. You had to wait for word one before processing word two. And worse, these networks struggled with long-range dependencies. By the time the model got to the end of a long sentence, it had often "forgotten" important context from the beginning. The architecture had memory, but it was imperfect, lossy, limited.
Engineers tried to solve this with increasingly complex mechanisms — gating functions, memory cells, attention bolted onto the side. The models got more complicated. Training got slower. And still, the fundamental problem remained: sequential processing doesn't scale.
Then eight researchers at Google decided to try something radical. What if you threw away the sequential architecture entirely? What if you let the model look at all the words simultaneously and figure out which ones matter?
The paper appeared in December 2017. The title was almost flippant: "Attention Is All You Need." It was bold, provocative, dismissive of everything that came before. It was also correct.
The core insight was something called self-attention. Instead of reading left to right, the model could look at every word in a sentence simultaneously and calculate which words were relevant to which other words. If you're trying to understand "The animal didn't cross the street because it was too tired," the word "it" needs to pay attention to "animal," not "street." Self-attention let the model figure that out automatically, regardless of how far apart the words were.
This solved both problems at once. First, you could parallelize everything — look at all words simultaneously, perfect for modern GPUs that excel at doing many calculations at once. Training became massively faster. Second, you could capture long-range dependencies perfectly because distance didn't matter anymore. Word one could directly influence word one thousand.
The architecture was called a Transformer, and it was elegant in a way that felt almost too simple. You stack layers of self-attention, add some feed-forward networks, include positional encoding so the model knows word order, and you're done. No recurrence. No convolution. No sequential processing. Just attention.
Dispensing with them entirely. Just throwing out the entire architecture everyone else was using.
The AI research community read the paper and initially reacted with skepticism mixed with curiosity. It seemed too simple. Surely you needed the complexity, the memory mechanisms, the sequential processing? But then people started implementing it. And it worked. And it worked better. And it scaled beautifully.
Within a year, Google released BERT — Bidirectional Encoder Representations from Transformers — which used the architecture to achieve breakthrough results on language understanding tasks. OpenAI released GPT-2, using Transformers for text generation. By 2020, virtually every major language model used Transformer architecture or variants of it. The eight researchers hadn't just published a paper. They'd made every other architecture obsolete.
The beautiful irony was that attention mechanisms themselves weren't new. Researchers had been adding attention to neural networks for years as a supplementary feature. But it was always attention plus something else. The radical move was making attention the only thing. All you need.
What happened to the eight authors tells you something about how value flows in artificial intelligence. Aidan Gomez co-founded Cohere. Llion Jones co-founded Sakana AI. Illia Polosukhin co-founded NEAR Protocol. They became successful entrepreneurs. Noam Shazeer became the exception — he left Google, co-founded Character.AI, and in 2024 Google bought the company back for $2.7 billion, with Shazeer reportedly receiving $2.5 billion personally.
The others remained in technical roles at major companies. Professionally successful, financially comfortable, respected in the field. But not capturing more than a tiny fraction of the value their invention created. Because the value of the Transformer wasn't in the paper. The paper was published openly, freely available. The value was in implementation, in scale, in the platforms built on top.
The Transformer architecture became infrastructure. And infrastructure, once published openly, belongs to everyone. The eight researchers created something worth trillions. The gap between creation and capture is enormous.
But here's what they did accomplish: they removed a fundamental bottleneck. Before Transformers, scaling language models was hard. After Transformers, scaling became almost straightforward. Bigger models, more data, more compute — and performance kept improving. GPT-3's 175 billion parameters would have been impractical to train with LSTMs. The Transformer architecture made the current AI boom possible by making scaling possible.
The paper has now been cited over 100,000 times, making it one of the most influential computer science papers ever written. The eight names — Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin — are known to anyone working in natural language processing.
They're not household names. Most people using ChatGPT or Claude have never heard of them. But they changed everything. They proved that attention was all you need. They threw away the complexity and found elegance. They published openly and created infrastructure.
The scientists wrote the score. And this time, we know their names: Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin. Eight researchers who changed everything with a paper titled, almost as a joke, "Attention Is All You Need."
It turned out that yes, it was.