The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
That's a bit of a wild interpretation of copyright law.
I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.
But would this argument would have flown in the past?
It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
Well in this case the original owner doesn’t know, the purchaser doesn’t know, the reporter doesn’t know, you and I don’t know. So what actually makes these books rare?
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
Do you really think they aren't? Do you think Amazon are carefully culling rare books from the list?
Although, we have to face the fact that the 'indie' book industry isn't necessarily preserving stuff either. I was recently in an old fashioned book shop in Charing Cross road, and they were selling old pictures as well as books. I was going to buy one as a present until I realised that they were pages that had been cut out of some book, because they could get more selling them separately.
The figure I've heard quoted is that 640,000 tons of used books (not 640,000 books, 640,000 tons of books) are shredded/pulped every year for lack of finding a buyer. I assume that very few of these are rare books.
I think in the situation where the background rate is 640,000 tons of books being destroyed per year, the burden of proof is on the one making the claim that Amazon is destroying rare books, and doing so at a higher rate than they're being destroyed already.
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
Well, I hate to be the one to break it to you but they're not releasing digital versions, they're just destroying them.
"rare" is used in these headlines/articles to incite and generate clicks
The headline implies at some level that Amazon loves or cares about ..... stuff.
This is what the copyright laws dictate no ?
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
They could just not do the evil thing.
(This is why I will never be a billionaire)
Isn’t it being destroyed because it makes the scanning process easier?
Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
Non-destructive scans are at least 10x as much, and tend to lower quality.
https://software.annas-archive.gl/AnnaArchivist/annas-archiv...
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
Perhaps this is a use case that's worth investigating, to invent better, less-destructive scanning processes! Or some way to rebind them afterwards.
No.
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
No. A copy is still a copy even if you destroy the original.
Until about a year ago this would have been a reasonable and respectable argument, but at least in California you are arguing against current legal precedent: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
That's a bit of a wild interpretation of copyright law.
I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.
But would this argument would have flown in the past?
It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
Well in this case the original owner doesn’t know, the purchaser doesn’t know, the reporter doesn’t know, you and I don’t know. So what actually makes these books rare?
Its somewhat ironic.
The worst part of Amazon destroying priceless old books in the quest to build the torment nexus is the hypocrisy
They obviously aren’t priceless if Amazon is buying them for pennies.
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
You're not really responding to the post, which is about the hypocrisy.
"They're allowed to be a hypocrite" doesn't mean they aren't a hypocrite.
[dupe] https://news.ycombinator.com/item?id=49330742
and previously:
https://news.ycombinator.com/item?id=49310725
https://news.ycombinator.com/item?id=49068738
Another discussion on this https://news.ycombinator.com/item?id=49336050
There's nothing in this article supporting the claim that rare books are being destroyed, it just has a link to a paywalled article from 404 Media.
Do you really think they aren't? Do you think Amazon are carefully culling rare books from the list?
Although, we have to face the fact that the 'indie' book industry isn't necessarily preserving stuff either. I was recently in an old fashioned book shop in Charing Cross road, and they were selling old pictures as well as books. I was going to buy one as a present until I realised that they were pages that had been cut out of some book, because they could get more selling them separately.
The figure I've heard quoted is that 640,000 tons of used books (not 640,000 books, 640,000 tons of books) are shredded/pulped every year for lack of finding a buyer. I assume that very few of these are rare books.
I think in the situation where the background rate is 640,000 tons of books being destroyed per year, the burden of proof is on the one making the claim that Amazon is destroying rare books, and doing so at a higher rate than they're being destroyed already.
Tech crunch used to be tech news, now it’s ai slop with a biased agenda.
It used to be breathless startup hype and gossip, which it might still be in some capacity.
no one uses ai for writing at tc. you're gonna end up chicken-littling yourself
Carnegie was building libraries and concert halls, Silicon Valley CEOs buy girlfriends with big tits, cheat in computer games and destroy books.
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.