Claude AI Failed 650 Times…Then Beat The Human Record

#LLMs #AI Tools
💬 Chat with this Video
Ask anything about this video…
What is happening? An unreleased version of Claude was asked to solve a long-standing mathematical problem. The Rayman hypothesis. Simplified. It is a statement about the distribution of prime numbers. Devilishly difficult. No human has ever been able to prove it. Did it solve it? Nope. [laughter] So, what did it do then? Well, it actually was able to improve a related bound beyond human record. It pushed humanity further sort of. I think that is incredible. Now there is the mathematical side. Many mathematicians I heard seem to be both surprised and impressed by the results. Some call it a massive leap forward. I ran the verification myself but I am just a student looking to learn and I am not qualified to speak more about the mathematical side of it. But the AI side is incredible. Three things that happened that I found super interesting. One, surely exquisite mathematical prompting was done, right? No, not really. Here's what happened. First, a non-mathematician person prompted this AI. Wow. Okay. The first 650 tries did not work. Now, hold on to your papers, fellow scholars. Quoting throughout this process, Jared's input was mostly limited to sending Claude messages of encouragement. Mostly varants of keep going and believe in yourself. What he had to give words of encouragement to an AI to keep going and it succeeded. Perhaps in the future, the most powerful mathematical proofs will not be written by geniuses. They will be written by life coaches. What a time to be alive. And get this, it's not the first time this has happened. Quoting a prompt including similar encouragement was used to help Claude disprove the Jacobian conjecture. Okay, but I was even more surprised about two more things. Dear fellow scholars, this is two minute papers with Dr. Koa Eher. Two, the full technical paper is available, but it's pretty tough, of course. So much so that they asked the AI to explain its findings. So it did and in the meantime a formalized version of the proof is also available which can be automatically verified. You can even run it yourself right now. Now something I don't think you hear too much about elsewhere. I went through more than a 100 pages of transcripts and some super cool tidbits from the journey. Claude had internet access in general but did not need to use it during the key breakthrough run. Claude went down many wrong roads first but was able to learn from them and recover. Then the first crucial result appeared after about 37 minutes of radio silence. Imagine how tense that must have been. And then pop. And the AI itself also thought the result is suspiciously good. It said, quoting, "Too strong to be new." Absolutely crazy. And three, even Claude was surprised about its finding and it was skeptical at first. They also say perhaps Claude itself also underestimates the rate of AI progress. But an AI that is surprised, of course, this does not mean a human-like surprise. It simply reflects patterns learned from us during training. But still, I feel like we are living in a new world where sentences that didn't used to make sense now suddenly do. And in a world where these AI systems built by human ingenuity are now pushing humanity forward, what a time to be alive. We need new tools for the era of LLMs and Weights and Biases now has Weave, a lightweight toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best. Try it out now at wnb.me/papers me/papers or click the link in the description below.

Generated algorithmically for Search Engine Indexing.

Summarize Another Video