Claude Is Now Leaving Invisible Fingerprints In Its Text

#AI Tools #LLMs
💬 Chat with this Video
Ask anything about this video…
Images can be watermarked for copy protection. Now, get this. Claude AI announced that they are watermarking the text you generate with it. From when? When does this start? Well, Anthropic is rolling this out right now. Yep. I am not talking about this because I agree with it, but because I think it's important that all of you fellow scholars know about this to inform the public. So, this piece of text is watermarked. Wait, what? >> [laughter] >> You can watermark an image by putting your logo on it, but text? How would you watermark text? You put hidden characters in it, right? Nope. This paper describes that it is a fingerprint in text that is invisible to humans, but is detectable for machines. It even survives copy-pasting and some editing, too. I'll tell you what it doesn't survive in a minute. So, when generating text, the AI decides what the next word should be, and there can be a few candidates. Here, you could say, "I saw a dog, a puppy, a cat, or a house." Based on context, each of these words gets a probability to be chosen. Now, with a fingerprinting algorithm, it secretly assigns a color to each word. Some are green, preferred, some are red, not preferred. And now comes a little nepotism. A little cheating, if you will. When choosing the next word, the green ones get a little nudge upwards. They will occur slightly more often. So, here's how to check for a watermark. In a piece of text, someone who knows the red and green words simply counts how many greens you have. This scheme has a mathematical property where, as you see more and more green, the probability of it being real human text is extremely small. Found 21 grains in a paragraph, suddenly the probability of that done by humans can be less than winning the lottery, much smaller. Note that the green words can be anything, no matter how inconspicuous, so you can't spot them. Which words are green can also change over time. Ouch. This is the simplified version of the algorithm. They are likely using the SynthID variant, which has context-dependent probabilities and a tournament system, too. But the heart of the algorithm is the same in most research papers I read. Some words are preferred and are given a slight edge in the generation, creating a unique fingerprint. Now, there are a lot of misconceptions about this out there. One, Claude-written text cannot be traced back to you, but it shows that Claude wrote the whole thing or heavily edited the text. I don't agree with this. I am making this video to let everyone know. Knowledge only changes the world when it reaches people. Second, some say, "Just edit a few words and it's clean." Nope, you can't get rid of it so easily. So, can you get rid of it? Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. With light editing, no. If you rewrite the whole thing, exchanging every word, yes, you can get rid of it. An open-weights LLM that works for you can also help. Okay, so who can check if there is a watermark in the text? Well, not you and not me. Some eligible organizations can, but that's it for now. So, what is the solution? Well, of course, use free and open-weights AI systems and run them yourself. These work for you, not against you. That [music] is the way of the scholar. We need new tools for the era of LLMs, and weights and biases now has weave a lightweight toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best. Try it out now at wnb.me/papers or click the link in the description below.

Generated algorithmically for Search Engine Indexing.

Summarize Another Video