The Paper That Changed Everything: How 8 Google Engineers Wrote 'Attention Is All You Need' in 6 Months โ And Accidentally Invented the Future of AI
In June 2017, a team at Google published a 15-page paper with a dismissive title. Within 18 months, it had killed RNNs, launched the GPT revolution, and made $2 trillion in market cap appear out of nowhere.
The Conference Room Where Deep Learning Broke
It was early 2017. Inside Google Brain's Mountain View offices, Ashish Vaswani was staring at a training curve that refused to cooperate. He was trying to build a translation model โ English to German, the kind of thing Google Translate had been doing for years. But the state-of-the-art approach, a complex beast called an LSTM (Long Short-Term Memory network), was taking weeks to train. Worse, it kept forgetting things. Ask it to translate a long sentence, and by the time it reached the end, it had forgotten the beginning.
LSTMs were sequential by nature โ they had to process words one at a time, like reading a book with a finger dragging across each word. You couldn't parallelize them. You couldn't throw more GPUs at the problem and make it go faster. You were stuck.
Vaswani had an idea. What if you didn't need to process words in order at all? What if you could look at the entire sentence at once and let the model figure out which words mattered to which other words?
He called it "attention."
His manager, Jakob Uszkoreit, was skeptical. So were most of the team. Attention mechanisms had been around since 2014, bolted onto the side of LSTMs like training wheels. But pure attention? No recurrence at all? That was heresy.
Vaswani pushed anyway. He assembled a small team: Noam Shazeer, Niki Parmar, Llion Jones, Aidan Gomez (a 20-year-old intern), ลukasz Kaiser, and Illia Polosukhin. They had six months before the NeurIPS deadline.
They started building.
The Architecture That Shouldn't Have Worked
The core insight was deceptively simple: instead of processing words sequentially, treat the entire input as a set. Then use a mechanism called "self-attention" to let each word "look" at every other word and decide which ones are important.
Here's how it worked:
Step 1: Embeddings and Positional Encoding
First, convert each word into a vector (a list of numbers). But since there's no sequential processing, the model has no idea about word order. So they added "positional encodings" โ sinusoidal functions that embed the position of each word directly into its vector. Word 1 gets a different positional signature than word 10.
Step 2: Self-Attention (The Magic)
For every word in the sentence, compute three vectors: Query (Q), Key (K), and Value (V). Think of Q as "what I'm looking for," K as "what I have to offer," and V as "the actual content."
Then, for each word, compute a score between its Query and every other word's Key. Softmax those scores to get weights. Multiply those weights by the Values. Sum them up. You've just computed a context-aware representation of that word.
Mathematically:
Attention(Q, K, V) = softmax(QยทK^T / sqrt(d_k)) ยท V
The division by sqrt(d_k) (the dimension of the key vectors) was a subtle trick Vaswani added to stabilize gradients. Without it, the softmax would saturate and stop learning.
Step 3: Multi-Head Attention
But one attention mechanism wasn't enough. Different words need different types of context. So they ran 8 attention mechanisms in parallel ("heads"), each learning different relationships. One head might focus on subject-verb agreement. Another on noun modifiers. Another on long-range dependencies.
Concatenate the outputs, pass them through a linear layer, and you've got multi-head attention.
Step 4: Feed-Forward Networks
After attention, each word's representation gets passed through a simple two-layer feed-forward network. Same network for every word, applied independently. This adds non-linearity and expressive power.
Step 5: Stack It 6 Times
They stacked 6 identical layers, each with multi-head attention + feed-forward. Each layer refines the representations further. By layer 6, the model has learned incredibly rich, context-aware embeddings.
The Decoder Side
For translation, they needed a decoder too. Same architecture, but with one twist: "masked" self-attention, so the model can't cheat by looking ahead at future words during training. Plus cross-attention, where the decoder attends to the encoder's output.
They called the whole thing a Transformer.
The Training Run That Proved Everyone Wrong
In May 2017, they started training. They used 8 NVIDIA P100 GPUs. The model had 65 million parameters โ tiny by today's standards, but massive for 2017.
The first results were... underwhelming. The baseline LSTM was still winning. Aidan Gomez, the intern, spent nights debugging. They tweaked learning rates, warmup schedules, dropout rates. Shazeer obsessed over initialization schemes.
Then, one morning, the numbers flipped.
On the WMT 2014 English-to-German benchmark, their model hit 28.4 BLEU โ a new state-of-the-art. It beat the best LSTM models. And it trained in 3.5 days instead of weeks.
But the real shock was how it parallelized. Because there's no sequential dependency, you could throw more GPUs at it and get linear speedups. LSTMs couldn't do that. Transformers could scale.
They had built the first architecture that could actually leverage modern hardware.
The Paper Nobody Understood
They submitted to NeurIPS 2017. The title was almost dismissive: "Attention Is All You Need."
The paper was 15 pages. Dense. Full of equations. The diagrams were clean but abstract. If you weren't deep in the weeds of sequence modeling, it was hard to see why this mattered.
The reviews came back: accept. But not as a spotlight. Just a regular poster.
At the conference in Long Beach that December, they set up their poster. A few researchers stopped by. Some asked questions. Most walked past. There were flashier papers that year โ GANs generating faces, RL agents playing Dota 2.
No one knew they were looking at the architecture that would power GPT, BERT, ChatGPT, Stable Diffusion, AlphaFold, and hundreds of billions of dollars in market cap.
The Explosion
January 2018: Researchers at Google and elsewhere start experimenting. They realize the Transformer isn't just good at translation โ it's good at everything. Text classification. Question answering. Summarization.
June 2018: OpenAI publishes GPT-1. It's a Transformer trained on books. 117 million parameters. It can generate coherent text. People start paying attention.
October 2018: Google publishes BERT. It's a Transformer trained to predict masked words ("The cat sat on the ___"). It crushes every NLP benchmark. The field goes insane.
February 2019: OpenAI publishes GPT-2. 1.5 billion parameters. They call it "too dangerous to release." The media has a meltdown. Transformers are now mainstream.
June 2020: GPT-3 drops. 175 billion parameters. It can write essays, code, poetry. It can do tasks it was never trained for. The concept of "emergent abilities" enters the lexicon.
Every one of these models is a Transformer. Same architecture from the 2017 paper. Scaled up.
The Technical Legacy
The Transformer didn't just win โ it became the only game in town. By 2023, RNNs and LSTMs were effectively dead. CNNs were hanging on in vision (but even there, Vision Transformers were taking over).
Why? Three reasons:
1. Parallelization
Self-attention processes all tokens simultaneously. You can train on thousands of GPUs in parallel. LSTMs couldn't do that. This unlocked scale โ and scale, it turned out, was everything.
2. Long-Range Dependencies
LSTMs had a "forgetting" problem. Transformers don't. Every word can attend to every other word directly. No information bottleneck.
3. Transferability
Pre-train a Transformer on tons of text, then fine-tune it on your task. This "transfer learning" paradigm became the foundation of modern AI. BERT for NLP. GPT for generation. ViT for vision. The architecture was universal.
But the real genius was in the details:
- Layer normalization (applied before each sub-layer, not after)
- Residual connections (skip connections around every layer to help gradients flow)
- Learned positional embeddings (later replaced with sinusoidal for better extrapolation)
- Scaled dot-product attention (that
sqrt(d_k)term that stabilizes training)
These weren't just hacks. They were carefully engineered design choices that made the architecture robust, scalable, and trainable.
The Team That Scattered
After the paper, the team went their separate ways.
Ashish Vaswani left Google in 2021 to found Adept AI (building "AI that can use software").
Noam Shazeer left in 2021 to found Character.AI (chatbots with personalities). Google later paid $2.7 billion to acquire it and bring him back.
Illia Polosukhin left in 2017 to found NEAR Protocol (a blockchain). He'd already moved on before the paper even published.
Aidan Gomez, the intern, left in 2019 to found Cohere (enterprise LLMs).
Jakob Uszkoreit left in 2021 to found Inceptive (AI for RNA design).
Llion Jones and Niki Parmar stayed at Google for a while, then quietly exited to startups.
ลukasz Kaiser moved to OpenAI in 2021.
Google trained them. Google published the paper. Google open-sourced the code.
Then every single one of them left to compete with Google.
The Legacy
Today, the Transformer is everywhere:
- GPT-4, Claude, Gemini โ all Transformers
- DALL-E, Midjourney, Stable Diffusion โ Transformers (image generation via latent diffusion)
- AlphaFold โ Transformer for protein folding
- Whisper โ Transformer for speech recognition
- GitHub Copilot โ Transformer for code generation
The architecture is so dominant that when people say "AI," they usually mean "a large Transformer trained on a lot of data."
The paper has been cited over 120,000 times. It's one of the most influential CS papers ever written.
And it almost didn't happen. If Vaswani had listened to the skeptics. If the NeurIPS reviewers had rejected it. If the team had stuck with LSTMs for another year.
But they didn't. They bet on attention. They wrote the code. They trained the model. They published the paper.
And they changed everything.
The Question Nobody Asks
Here's the wild part: the Transformer wasn't invented to build AGI. It wasn't designed for chatbots or image generation or protein folding.
It was designed to translate German sentences faster.
That's it. A practical engineering problem at Google. "Can we make translation training less painful?"
And the solution they found โ letting words attend to each other in parallel โ turned out to be the key to general intelligence.
No one saw it coming. Not the authors. Not the reviewers. Not the field.
The future of AI was hiding in a paper about machine translation. And eight engineers in Mountain View found it by accident.
Attention, it turned out, was all we needed.
Keep Reading
The 10-Day Bet That Broke Java: How Brendan Eich Built JavaScript in a Netscape Conference Room โ While Sun Microsystems Tried to Kill It
In May 1995, Netscape gave Brendan Eich 10 days to invent a new programming language or watch Java take over the web. He delivered a prototype that looked nothing like Java โ and accidentally created the most widely used language on Earth.
The 72-Hour Rewrite That Killed Oracle: How Michael Stonebraker Built Postgres in a Berkeley Lab โ And Accidentally Invented the Database That Powers Uber, Instagram, and Apple
In 1986, a Berkeley professor watched Oracle charge $50,000 for software that crashed every Tuesday. So he locked himself in a lab with 6 grad students and rewrote database history โ inventing the MVCC architecture that would change everything.
The 6-Week Hack That Killed Flash: How Jordan Walke Built React in His Spare Time โ And Facebook Accidentally Started a Revolution
A Facebook ads engineer couldn't sleep. Every time someone clicked a sponsored post, the entire page reloaded. So he built a side project that re-rendered DOM nodes in 16 milliseconds โ and broke the internet's understanding of how web apps should work.