It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
> The main question is why aren't they leaking it to AA themselves?
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
The problem is that you probably do little research or read very few old books.
There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.
You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.
Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.
Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.
I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.
I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
> We've been seeing that headline for a few weeks now and I really don't understand the problem.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
> because copyright law forces them to do stupid things.
Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.
> because copyright law forces them to do stupid things
This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?
Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?
Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.
It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯
I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.
Yeah, which is completely fine. There's a major difference between:
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
More importantly: Once Anthropic is gone, all knowlege is lost.
It will probably be actioned off in the bankruptcy proceedings.
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.
Obviously. Copyright infringement is settled law.
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
we also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up
Nobody will ask you, they'll be taken from you, in case you missed what happens around.
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
I agree with you for the same reason I think McDonald's is the best restaurant in the world!
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
> The main question is why aren't they leaking it to AA themselves?
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
Because that’s illegal.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
Project Unica is an initiative by the University of Illinois libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
What often gets missed is that they are buy one physical copy and turning it into a digital copy.
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
> turning it into a digital copy.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
They are. They can’t release it; that would be copyright infringement. But they’re absolutely planning to make further use of the book later.
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
https://www.bbc.com/news/articles/cp3rprx2wl4o
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
Rare and out of print does not mean important or valuable.
Usually it means the opposite. Books that are old and valuable tend to be out of copyright, so they do see new printing runs.
Presumably, being from the 18th Century, copyright law wouldn't apply?
Detestable if they're doing it anyway to prevent competitors getting hold of it.
The problem is that you probably do little research or read very few old books.
There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.
You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.
Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.
Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.
I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.
I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.
The problem is we don't know what we're losing, due to lack of transparency.
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
I’m confused - are you intending to say Anthropic and Google are buying ALL the books?
> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. This isn't about legality, this is about what's right.
> We've been seeing that headline for a few weeks now and I really don't understand the problem.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
Since when books have become a supply limited asset ?
Try and read a book that’s been burnt and find out
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs.
At least there will be a copy left for us. The AI companies won't share these books in their original form.
Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things.
Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.
The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place
> because copyright law forces them to do stupid things.
Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.
Copyright law didn't force them to torrent terabytes of books what they did and got caught doing so.
I'm not convinced they destroy the books to obey the law, lol.
> because copyright law forces them to do stupid things
This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?
Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?
Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.
It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯
I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.
Google probably wanted to sell the whole of Google Books.
Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.
Yeah, which is completely fine. There's a major difference between:
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
I much prefer B.
B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to.
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.