In 2017, eight researchers at Google published a paper called “Attention Is All You Need.”
Almost nobody outside AI noticed.
Every major AI chatbot you use today is built on that one idea.
You don’t need math to understand it. You just need to understand one problem about language.
Words don’t have meanings. They have meanings in context.
Take the word “bank.”
“I sat by the bank.” “I deposited money at the bank.” “Don’t bank on it.”
Same word, three meanings. You didn’t have to think about it. Your brain looked at the surrounding words and picked the right one instantly.
The linguist J.R. Firth wrote in 1957: “You shall know a word by the company it keeps.”
The hard part here was teaching machines to do it.
The old way: reading one word at a time
Before 2017, models read like you read through a keyhole. One word at a time, left to right.
And this had two problems.
They’d forget things. For example, by the end of a long sentence, the beginning would’ve faded. If a pronoun referred to something twenty words back, the model often lost track of what it pointed to.
They were also slow to train. Word two had to wait for word one, and word three had to wait for word two. You couldn’t do this in parallel, so you couldn’t really scale it.
Researchers had already started bolting a trick called attention onto these models. The 2017 paper’s bold move was to drop the word-by-word reading entirely and keep only attention.
That’s why it’s called “Attention Is All You Need.”
The new way: every word looks at every other word
The transformer, the architecture that was described in that paper, does something at least conceptually simple. It reads the whole sentence at once. And for each word, it asks: which other words here matter for understanding this one?
That’s attention.
Here’s a classic example:
“The animal didn’t cross the street because it was too tired.”
What does “it” refer to? The animal.
Now change one word:
“The animal didn’t cross the street because it was too wide.”
And now “it” is the street.
One word changed. The meaning of “it” changed. The model handles this by letting “it” look at every other word and weigh what matters.
How the looking works
Every word carries three things. A question: what am I looking for? A label: what do I contain? A payload: what can I pass along?
Each word checks its question against every other word’s label. Strong matches get more weight. Then it gathers the payloads, weighted by those matches, and updates its own meaning.
The model runs many of these checks at once, each tracking a different kind of relationship. Some seem to track grammar, others reference. Then it stacks dozens of layers on top of each other, refining every word’s meaning again and again, until the math captures what each word means in this specific sentence.
And because the transformer reads everything at once, training can run in parallel. That’s what made it possible to scale these models up.
Then the model does one simple thing, after all that processing: it predicts the next chunk of text, what we call a token, which is usually a word or part of a word. That’s it. It predicts the next token. And then it does it again.
Every answer ChatGPT has ever given you was written this way. One token at a time.
So where does the intelligence come from?
Predicting the next word sounds kind of dumb, but it isn’t. It’s actually very hard.
To predict the end of a mystery novel, you have to track the plot. To predict the next line of code, you need to understand the logic. To predict the next word of an argument, you need to understand the argument.
These models are trained on enormous amounts of text. Every time they guess wrong, billions of internal numbers get updated slightly. Do that trillions of times, and to get the predictions right, the model has to build internal representations of grammar, facts, and reasoning patterns.
Then it’s fine-tuned, often with human feedback, to act like a helpful assistant instead of an autocomplete.
Nobody really taught the model intelligence. Something that looks a lot like intelligence turned out to be a requirement for predicting well.
What we still don’t know
We know how these models are built, but we know much less about what happens inside them once they’re trained.
Even asking the model doesn’t settle it. When a model explains its reasoning, there’s no guarantee the explanation matches what actually happened inside it.
Maybe these models “understand things,” or maybe they “just predict words.” The truth is probably something way more strange.
Dan Berges is the founder and managing director of Berges Institute, an online Spanish language school, and lead developer of Berges AI, a text assistant built on open-weight models that gives direct, concise answers. He also publishes content in Spanish about descriptive grammar, semantics, and pragmatics on Instagram, YouTube, and TikTok.

