Another DeepSeek Moment

#AI
💬 Chat with this Video
Ask anything about this video…
It is simply incredible what just happened. Deep Seeks 2 latest AI systems appeared just about 3 months ago and the quicker flash model is getting an update already and this is unbelievable. I don't even know where to start. One, the new flash model is way better than the previous one. Many of these results more than doubled. One went 7x in just one revision, just a bit after the original release. Is this real life? And two, it gets better. Look, it not only beats the previous flash version, but also beats the pro version that is about five times larger. I mean, what? We often say that benchmarks are not everything and these are not even all of them. Yes, but a change this huge cannot be ignored. So what happened? How did they do this black magic? New model that was trained longer? Nope. The underlying model architecture and size is the same. So what changed? Excuse me. Is this some kind of joke? Just the post-training step changed. Dear fellow scholars, this is two minute papers with Dr. Koa Eer. Yep, that is the power of post training. You see, it has the same base brain, but a completely new playbook. The base model already contains the raw knowledge. Then the post-training step teaches it when to use which ability, how to plan, how to check its work, and how to recover from mistakes. Imagine the brain as a builder in a toy workshop. Before post-training, it can see the right pieces but uses them poorly and makes a huge mess. After post-training, it learns strategy. Build a base. Test each part. Fix mistakes early. Think ahead. Now we have the same builder, the same toolbox, the same raw knowledge, but now we have a much smarter sequence of actions. And that is why post-training can create a huge leap in performance even when the underlying model stays very similar. Remember this one because you are seeing one of the finest examples of it ever. And look, this is ridiculous. Yep, you can download and own the weights forever. No 5hour sessions, no weekly caps, no games. This is a true deepseek moment. a true open AI. It needs a beefy machine to run locally. Yes. Or you can run it on Lambda or use the API. It's dirt cheap compared to the frontier companies. By the way, I held off on making this video for a bit so I could test it myself and see how you genius fellow scholars use it in the wild. I think that made for a much better video for you. Now, hold on to your papers, fellow scholars, because if this pace continues, in less than a year, we might have a free open model close to what is now the billiondoll feeble level intelligence compressed enough to run on a beefy laptop. It sounded impossible just a few weeks ago, but now I no longer think that this is impossible. I am out of words. This is the power of open source and open science. What a time to be alive. So we get this for free and we will be able to own it forever. And that is the most important. Knowledge is incredible. But knowledge only changes the world when it reaches the people. So thank you so much to everyone who worked on this. You are heroes. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video. Easy peasy. Running a Deepseek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later results. Love it. Seriously, try it out now at lambda.ai/papers.

Generated algorithmically for Search Engine Indexing.

Summarize Another Video