The Paper That Created Modern AI

#Transformer architecture #Attention mechanism #AI language models #Recurrent neural networks #Google Translate #AI research paper #Programming
💬 Chat with this Video
Ask anything about this video…
Every time you talk to an AI, every question you type into ChatGPT, and every answer you get from Gemini or Claude, you're using an idea from one single paper. One research paper written by eight people who worked at Google who are mostly just trying to make Google Translate a little bit better. Today, I'm going to tell you the story of how the Transformer came to be invented, and it's pretty wild because I feel like nothing these days is really that groundbreaking. And it's not like we are inventing the lightbulb or the telephone every year. No, but I feel this is mentioned up there among the greats and will be referred to in many years to come. To understand why this is so significant, you have to understand how machines used to read language before it. Old systems would take in text one word at a time. Kind of like you would imagine words passing through a straw. So, one word passes through, gets saved to memory, then the next word passes, gets saved to memory, and so on. And this would happen from left to right. This process actually has a name, and it's called recurrent networks. And whilst vastly used at the time, I guess because there was no alternative and no one had questioned this system so far, this system definitely had its flaws. One of major one being that this system was forgetful. By the time the model reached the end of a long sentence, the beginning had already gone fuzzy. So, connecting an idea at the start of a paragraph to one at the end of the paragraph was genuinely hard. And the second issue is that each step depended on the one before it. This means that even if you wanted to, you couldn't run all the steps at once. You had to march through each step one by one. So, through the sentence in order. As you can imagine, this made the whole process very, very, very slow. Not only for reading sentences, but also for training these models. And it's not like you could just throw more computing power at it, either. It wouldn't make a difference. Word five genuinely can't be computed until word four is done. So, more GPUs don't speed up processing one sentence. So, anyway, this group of eight engineers and the whole field of linguistics really were kind of stuck using this system. And I'm not sure what happened exactly, but I guess they started questioning why this sequential approach existed. I mean, why does the model read one word at a time? Wouldn't it be better if the model could look at the entire sentence at once? So, every word simultaneously. So, you know, in just the all all the words and for each word figure out which other words the model should pay attention to. Now, what is great is that actually this mechanism to pay attention to other words already existed as a helper in these old word-by-word models. And pretty fittingly, it was called attention. So, let's understand this theory a little bit better by looking at a sentence. The animal didn't cross the street because it was too tired. Now, when you read the sentence, we as humans know what it refers to. We know it refers to the animal. But, a computer reading word-by-word might think it refers to street, purely because it's the closest noun to the word it. The mechanism of attention is what lets the model make the same decision as you, linking it to animal, no matter how far apart they are. So, these eight Google researchers had an idea. They had the mechanism, and now they just needed to develop their idea further. One of the Google researchers, a guy named Jacob Uszkoreit, decided to fully revamp the helper, or in other words, the mechanism of attention, and isolate it to use it on its own. This was essentially the right idea, but executing it wasn't that straightforward, it turns out. It was in fact Noam Shazeer, another at the time, and a little bit of a legend around Google from what I gather, that managed to turn the theory into practice. Now, I don't want it to sound like I'm giving Noam Shazeer any more credit than any of the other guys, because they certainly didn't. In fact, when it came to writing the paper on this, each of the eight researchers who worked on the idea gave themselves equal credit, and even scrambled the order randomly, working as a true collective. Even the intern at the time, a guy called Aidan Gomez, is listed as an equal contributor. And this paper, this paper that is now one of the most cited AI or machine learning papers of the century, the one that talks about the mechanism of attention, feeling all you really need to improve language models, has the best name ever. If you are a Beatles fan and are familiar with their songs, you will definitely appreciate how apt the title of this paper is. It's called Attention Is All You Need. Get it? Because of the attention mechanism. Anyway, so the paper is in the process of being written, and their researchers' deadline is looming, and essentially, they are describing how a specific architecture works. And this architecture needed a name, so Jacob Uszkoreit went with Transformer, for no other reason that he just liked how it sounds, from what I understand. I mean, fair, it does sound pretty cool. Maybe he was a Bumblebee fan, who knows. So, now they have the paper. They have described the architecture, time to put it to the test. Because these researchers, if you remember, were originally looking at making Google Translate faster, what better way to see if the Transformer was able to fix their problem. So, they put it to the test. They ran the old model against the new Transformer, and would you believe the Transformer didn't just beat it, it totally demolished it. Not only did it beat the old system, it beat the best systems in the world when put to the test and trained in just a fraction of the time. These researchers were pretty ecstatic. I mean, this was a major major breakthrough. So, they finished the paper and essentially published it to the world. Now, this part makes me a little bit sad because I feel their pain because honestly, no one really cared about this paper because on the surface, it was a translation paper and it was pretty boring. It was pretty technical and narrow and turns out talking about translating words faster isn't exactly sexy. But luckily, the researchers who did read it totally got it. They totally understood how running a system in parallel is much better than running a system one word at a time. Running a system in parallel meant you could scale it. You could make it bigger, feed it more of the internet, throw more computers at it, and it wouldn't crash. It would just keep getting better and better and better. So, this happened in 2017 and pretty much within a year, the Transformer had already paved the way for the first of this new kind of models to be created. These models could read and write language and one of them was called GPT or Generative Pre-trained Transformer. And yes, that T, the one you say every time you say ChatGPT, stands for Transformer. And it didn't just stop at chatbots. Nearly every major AI system you can name today, so GPT-4, Gemini, Cohort, Llama, you name it, is underneath a Transformer. So, think about it. The models that generate images, tick. The models that write code, them too. And even the AI that predicted the structure of nearly every protein known to science, look under the hood and again and again and again, you will find the Transformer architecture. But our The doesn't stop there, no, because in a major twist, all eight of the researchers ended up leaving Google. They founded or joined companies that today are recognizable names in the AI space. One co-founded Cohere, one co-founded Character AI, one went to OpenAI, others started Adept and Essential AI and Sakana AI in Tokyo, and even a blockchain platform. So, in other words, this paper, the one that was written by Google employees, essentially created Google's biggest rivals. Talk about being the creator of your own downfall. So, yes, this was the story of the Transformer and how it came about, as well as the research paper called Attention is All You Need.

Generated algorithmically for Search Engine Indexing.

Summarize Another Video