So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet (from now on), yet people start complaining that "their" output is marked as spam.
Well, the solution is quite easy: start thinking on your own again and write the lines yourself.
I welcome this watermarking. Finally it's an easy detectable signal that someone just generated some request/answer to waste my time by forcing me thinking for the other person too. Now that time is over.
(It never was hard to spot this type of texts but now there is evidence)
> So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet
We are very, very far from that! So far we have a single ai vendor introducing a statistical bias to their generation that can make it simpler to identify genAI in some texts. We don’t know yet how effective that will be in practice, and if other models will follow suit
Isn’t this about turning a weakness into a feature? Claude is already unable to write original prose that does not trigger an AI detector like Pangram.
Not sure why Anthropic is getting all the attention here. Google been doing this since at least late 2024[1] and openai I believe is doing this around 9 days ago[2].
Realized your right about OpenAI but they plan to. In the link they state "our goal is to expand provenance signals to all modalities including text" so its comming up regardless. I pretty sure Google does it though.
So, it seems the watermark is less what it sounds like (a stamp) and more an "imperceptible statistical pattern woven into the choice of words and sentence structures." Does that mean AI responses will sound even more "AI?" Like, will it become even easier to detect on a read-through because of the word choices and patterning? I see the word "imperceptible" there, but what does this mean in this context? My non-tech brain is kinda breaking here.
> Like, will it become even easier to detect on a read-through because of the word choices and patterning?
It's turned on right now. Can you tell a difference? I can't.
How many ways could I write this paragraph and still convey the same idea? Way more than we're aware of. Hundreds? Thousands? Maybe a lot more? The number of semantically similar variants increases exponentially with each word.
I suspect anthropic could turn their fingerprinting up or down if they want. If it were turned way up, claude would use weird phrasing but it would take very little text to tell if something were AI generated. If they turned it down, it would seem imperceptible to humans, but you would need a large sample to determine (with high accuracy) that a passage was AI generated. There's probably a very large middle ground where humans can't tell, and where it doesn't take a large text sample to know (with high probability) that some text was AI generated.
The information being encoded (the watermark) is the _relative ranking of each token compared to other possibilities_. If our prompt was "Write a positive review for a restaurant" and the response began:
"The restaurant "
Our next set of predictions might be:
1. was \n 2. had \n 3. offers
So we append the rank of the next token (1, 2, or 3) onto the secret. Given a long enough response, that secret becomes unique enough to use as a watermark. This obviously relies on having full deterministic access to the LLM itself, i.e. I don't believe it will be possible for users to derive the fingerprint from text that they've generated, only Anthropic will be able to.
The immediate objection is that this runs the risk of degrading the quality of the response. I think that's totally valid and I'll be curious how Anthropic handles it.
That's my very rough understanding! If someone with more knowledge wants to expand, feel free.
Sure, but if my experience (I know, I know) is anything to go by, it's increasingly difficult to get Claude to write in anything other than it's own house style.
Will anyone even really care (outside of detecting cheaters in academia)? I see people talking about LLM based posts all the time even at work in Slack conversations where a manager verbatim pastes a clearly AI output as a response for some question and while people seem to care internally they don't actually seem to call them out.
The actual original dunk from that page is great enough to be reproduced here again verbatim:
> during the months of autumn [...] Work is left to feebler hands. ... In those months the great oracle becomes —what at other times it is not—simply silly. In spring and early summer, the Times is often violent, unfair, fallacious, inconsistent, intentionally unmeaning, even positively blundering, but it is very seldom merely silly. ... In the dead of autumn, when the second and third rate hands are on, we sink from nonsense written with a purpose to nonsense written because the writer must write either nonsense or nothing.
Sounds like "The Saturday Review of Politics, Literature, Science, and Art" where this came from was a wonderful little venture. Is there anything like it today?
Feels like "we sink from nonsense written with a purpose to nonsense written because the writer must write either nonsense or nothing" accurately describes HN sometimes :D
What would really be crazy is if there were a platform on which people provided commentary about other people's commentary on other people's commentary.
Legitimately we should have had this day one for a whole host of reasons (not the least of which is the issue of inadvertently using output from generative AI as training data for Generative AI).
But bigger than that is that while there are folks that want to be able to have a machine do their textual output for them with as little energy on their part as possible, the rest of us have to deal with the impact of that call on their part.
The purpose of communicating with other humans is not utilitarian, psychologically it is the pathway to connection and forms a large basis of how we are able to get along without eradicating ourselves as a species. Using generative AI to replace that humanity seems unthinkable to me, and having markers that help us differentiate serves as a way to keep others from gaslighting us, which undermines trust and communication.
I realize as a stereotype technical folks see communication as a means to a utilitarian end, but it's so much more important than that, so much so that even this small action of watermarking generative AI output can keep untold disasters from occurring because humans didn't know other humans were using it.
Well, of course. The stigma surrounding AI use will only get worse with stuff like this. I'm not interested in having people single me out for using AI. So glad I switched away from Anthropic.
The stigma surrounding AI use will get worse if... People know AI is being used? Sorry how does maintaining heightened paranoia by omitting clear signals help exactly?
Unfortunately the witch hunts don't stop whether you use AI or not. If you don't use it, false positives mean the anti-ai hunters come for you eventually anyway. Happened in plenty of art communities already, will happen in code communities as well. The people who get hurt are the people just trying to make things.
Hiding your use of AI will only make things worse when people find out. Look at how badly the CEO of Saber Interactive is crashing out after they lied about firing a writer to replace them with AI.
Amusingly, the 'histrionic' Reddit post linked to in the article smells very AI generated to me. [0]
"And it doesn't stop there. There's the power question."
So not only is the Reddit user in question writing impassioned angry screeds about being caught using AI in their work, but they are presumably asking for Claude's help in writing said prose. I wonder what that prompt looked like?
>Amusingly, the 'histrionic' Reddit post linked to in the article smells very AI generated to me.
The whole text was a smorgasbord of the worst, most boring AI clichés possible. It was hard to finish for me, even though it is not that long, just because it was so badly written.
The only reason why I think it was partly written by a person who is just way too used to AI (as opposed to a simple prompt) is that the gaps in its logic feel more human than artificial.
This is a dumb idea by Claude. It's currently a binary switch and doesn't differentiate between "fix the grammar on this paragraph" and two pages of AI Slop.
There should at least be some cut off where it applies and where it doesn't.
Well, um... I fail to see how this changes anything? Give me two random pieces of text, one written by Claude, one not. I'm pretty sure I will be able to tell with close to 100% accuracy which is which, as long as as there will be a non-trivial amount of text.
Claude outputs are already very easily distinguishable. The watermark is already pretty much there, even if it's not explicitly put in the outputs.
I understand their reasoning, but imagine at some point in the future, watermarks eventually lead to a level 9 vulnerability in your system. Theoretically, it is possible
So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet (from now on), yet people start complaining that "their" output is marked as spam.
Well, the solution is quite easy: start thinking on your own again and write the lines yourself.
I welcome this watermarking. Finally it's an easy detectable signal that someone just generated some request/answer to waste my time by forcing me thinking for the other person too. Now that time is over.
(It never was hard to spot this type of texts but now there is evidence)
> So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet
We are very, very far from that! So far we have a single ai vendor introducing a statistical bias to their generation that can make it simpler to identify genAI in some texts. We don’t know yet how effective that will be in practice, and if other models will follow suit
Isn’t this about turning a weakness into a feature? Claude is already unable to write original prose that does not trigger an AI detector like Pangram.
Not sure why Anthropic is getting all the attention here. Google been doing this since at least late 2024[1] and openai I believe is doing this around 9 days ago[2].
[1]https://www.nature.com/articles/s41586-024-08025-4 [2]https://help.openai.com/en/articles/8912793-provenance-signa...
They don't do it for text I think that is the core issue.
Somehow it is acceptable for us to have photo, video or audio watermarked but not text or worse code.
Realized your right about OpenAI but they plan to. In the link they state "our goal is to expand provenance signals to all modalities including text" so its comming up regardless. I pretty sure Google does it though.
So, it seems the watermark is less what it sounds like (a stamp) and more an "imperceptible statistical pattern woven into the choice of words and sentence structures." Does that mean AI responses will sound even more "AI?" Like, will it become even easier to detect on a read-through because of the word choices and patterning? I see the word "imperceptible" there, but what does this mean in this context? My non-tech brain is kinda breaking here.
> Like, will it become even easier to detect on a read-through because of the word choices and patterning?
It's turned on right now. Can you tell a difference? I can't.
How many ways could I write this paragraph and still convey the same idea? Way more than we're aware of. Hundreds? Thousands? Maybe a lot more? The number of semantically similar variants increases exponentially with each word.
I suspect anthropic could turn their fingerprinting up or down if they want. If it were turned way up, claude would use weird phrasing but it would take very little text to tell if something were AI generated. If they turned it down, it would seem imperceptible to humans, but you would need a large sample to determine (with high accuracy) that a passage was AI generated. There's probably a very large middle ground where humans can't tell, and where it doesn't take a large text sample to know (with high probability) that some text was AI generated.
https://arxiv.org/html/2510.20075v6
It is quite counterintuitive, but you can hide texts the same size as the original text in imperceptible statistics of a text.
Compared to that feat, hiding a watermark is very easy.
The information being encoded (the watermark) is the _relative ranking of each token compared to other possibilities_. If our prompt was "Write a positive review for a restaurant" and the response began:
"The restaurant "
Our next set of predictions might be:
1. was \n 2. had \n 3. offers
So we append the rank of the next token (1, 2, or 3) onto the secret. Given a long enough response, that secret becomes unique enough to use as a watermark. This obviously relies on having full deterministic access to the LLM itself, i.e. I don't believe it will be possible for users to derive the fingerprint from text that they've generated, only Anthropic will be able to.
The immediate objection is that this runs the risk of degrading the quality of the response. I think that's totally valid and I'll be curious how Anthropic handles it.
That's my very rough understanding! If someone with more knowledge wants to expand, feel free.
Is it possible for them to sound even more AI?
Don’t forget the survivor bias is at play, you don’t see the ones that are good at passing for human written
Sure, but if my experience (I know, I know) is anything to go by, it's increasingly difficult to get Claude to write in anything other than it's own house style.
Unlikely, but I really hope so!
Anthropic seems to be on a mission to drive away as many users as possible from Claude
They are probably working on a $500 subscription without watermarks. The majority of their users have an IQ below 100 and need the crutch.
Will anyone even really care (outside of detecting cheaters in academia)? I see people talking about LLM based posts all the time even at work in Slack conversations where a manager verbatim pastes a clearly AI output as a response for some question and while people seem to care internally they don't actually seem to call them out.
Summarising a reddit discussion counts as news nowadays?
https://en.wikipedia.org/wiki/Silly_season
The actual original dunk from that page is great enough to be reproduced here again verbatim:
> during the months of autumn [...] Work is left to feebler hands. ... In those months the great oracle becomes —what at other times it is not—simply silly. In spring and early summer, the Times is often violent, unfair, fallacious, inconsistent, intentionally unmeaning, even positively blundering, but it is very seldom merely silly. ... In the dead of autumn, when the second and third rate hands are on, we sink from nonsense written with a purpose to nonsense written because the writer must write either nonsense or nothing.
Sounds like "The Saturday Review of Politics, Literature, Science, and Art" where this came from was a wonderful little venture. Is there anything like it today?
Feels like "we sink from nonsense written with a purpose to nonsense written because the writer must write either nonsense or nothing" accurately describes HN sometimes :D
What would really be crazy is if there were a platform on which people provided commentary about other people's commentary on other people's commentary.
Been the case for more than a decade
0 upvotes too
Sauce for the goose.
Legitimately we should have had this day one for a whole host of reasons (not the least of which is the issue of inadvertently using output from generative AI as training data for Generative AI).
But bigger than that is that while there are folks that want to be able to have a machine do their textual output for them with as little energy on their part as possible, the rest of us have to deal with the impact of that call on their part.
The purpose of communicating with other humans is not utilitarian, psychologically it is the pathway to connection and forms a large basis of how we are able to get along without eradicating ourselves as a species. Using generative AI to replace that humanity seems unthinkable to me, and having markers that help us differentiate serves as a way to keep others from gaslighting us, which undermines trust and communication.
I realize as a stereotype technical folks see communication as a means to a utilitarian end, but it's so much more important than that, so much so that even this small action of watermarking generative AI output can keep untold disasters from occurring because humans didn't know other humans were using it.
This is aimed at academia with all the cheating kids and youngsters. I really dont have a big problem with it but I doubt it will be successful.
Well, of course. The stigma surrounding AI use will only get worse with stuff like this. I'm not interested in having people single me out for using AI. So glad I switched away from Anthropic.
The stigma surrounding AI use will get worse if... People know AI is being used? Sorry how does maintaining heightened paranoia by omitting clear signals help exactly?
I'm not interested in "helping" people call me "slop fetishist" or "clanker lover", thanks.
Unfortunately the witch hunts don't stop whether you use AI or not. If you don't use it, false positives mean the anti-ai hunters come for you eventually anyway. Happened in plenty of art communities already, will happen in code communities as well. The people who get hurt are the people just trying to make things.
Did you write this with AI? jk
Hiding your use of AI will only make things worse when people find out. Look at how badly the CEO of Saber Interactive is crashing out after they lied about firing a writer to replace them with AI.
I don't hide my AI use at all.
Then why care about a watermark?
Do these watermarks show up in code as well?
Amusingly, the 'histrionic' Reddit post linked to in the article smells very AI generated to me. [0]
"And it doesn't stop there. There's the power question."
So not only is the Reddit user in question writing impassioned angry screeds about being caught using AI in their work, but they are presumably asking for Claude's help in writing said prose. I wonder what that prompt looked like?
[0] https://www.reddit.com/r/artificial/comments/1vlzjc8/about_t...
>Amusingly, the 'histrionic' Reddit post linked to in the article smells very AI generated to me.
The whole text was a smorgasbord of the worst, most boring AI clichés possible. It was hard to finish for me, even though it is not that long, just because it was so badly written.
The only reason why I think it was partly written by a person who is just way too used to AI (as opposed to a simple prompt) is that the gaps in its logic feel more human than artificial.
This is a dumb idea by Claude. It's currently a binary switch and doesn't differentiate between "fix the grammar on this paragraph" and two pages of AI Slop.
There should at least be some cut off where it applies and where it doesn't.
Well, um... I fail to see how this changes anything? Give me two random pieces of text, one written by Claude, one not. I'm pretty sure I will be able to tell with close to 100% accuracy which is which, as long as as there will be a non-trivial amount of text.
Claude outputs are already very easily distinguishable. The watermark is already pretty much there, even if it's not explicitly put in the outputs.
Then a whole industry will rise at tracking and removing those watermarks.
Already happening. Just have a quick search for llm watermark remover and discovered multiple repos already available.
It’s very easy, just change each third word into something different than what you primarily wrote.
It’s very simple, just change every third word for something different to what you originally wrote.
It’s very straightforward, just change each and every third word for something different instead of what you actually wrote.
And so on ad infinitum.
Yet another (accidental) win for the open-weight models.
I understand their reasoning, but imagine at some point in the future, watermarks eventually lead to a level 9 vulnerability in your system. Theoretically, it is possible