The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.
I like this and think it could make a lot of sense but you would need to thoughtfully tie it to volume or something similar. I could see publishers gaming the system. I would also add that once it drops out of print that it should belike generic drugs. Anyone can use it. IMO making it quicker free use stops most of what this article is describing.
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
I worked in a bookshop. I know all about this. I've seen pallets of books with covers torn off so they can be reported destroyed.
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
This was my issue as well. It seems like they tip toed around it in one paragraph saying usually these are books with isbn numbers so not truly collectible or very rare and then in another paragraph plainly stated that it could be foreign language or low volume books that are rare. Rare expresses a different meaning to me and I think it amounts to exactly what you said, manufactured outrage.
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
if we are gonna pearl clutch about this problem then the headline should be about how libraries -- the entities who preserve rare books -- don't have enough funding to do this. why are we now expecting booksellers to preserve rare books? they are booksellers, in trade
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
Im with you, my take-away from the article was of sadness for all the books that are going extinct because of this.
Real humans sharing their unique knowledge, packaged in a book.
The thing is, I doubt this is even in the top 10 causes of books going extinct. I feel like people often over-romanticise the medium, especially.
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
Please scan this book and put it into Anna's library, or keep it for the future.
I would be also interested in participating in your costs.
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
I wonder if there's any way this can be construed to get favorable tax treatment. That'd actually get them doing it.
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
Google faced a decade of litigation for making books searchable. It's all downside and little upside for a business.
I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.
The only way to gatekeep a book behind payment should be, if it is in print.
I like this and think it could make a lot of sense but you would need to thoughtfully tie it to volume or something similar. I could see publishers gaming the system. I would also add that once it drops out of print that it should belike generic drugs. Anyone can use it. IMO making it quicker free use stops most of what this article is describing.
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.
This is why we read the comments first :)
Thank you.
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
Not even the title of one of those rare books?
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
Yes this is disingenuous to the extreme.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
Makes you wonder how many packets with rare books.containing airtags said recipient receives. My guess is: one so far.
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
In what way is this stealing?
The rare qualifier is used precisely because no reasonable person thinks this is stealing.
The rare qualifier is used to highlight the destruction of rare items.
It is still stealing even if the book is common.
https://jskfellows.stanford.edu/theft-is-not-fair-use-474e11...
https://styleblueprint.com/everyday/the-quiet-theft-ai-steal...
Hoarding.
Nothing good comes from hoarding whether it being toilet paper, money or knowledge.
> Hoarding.
No reasonable person thinks buying 1 of something is "hoarding".
What about 1 of everything?
As I understand it, "rare" in this context could mean anything—even a washing machine manual from the 1980s...
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
Complete stab in the dark: AWS "training as a service" (further split into multiple microservices) for companies looking to train their own models.
AWS has nova and titan models that nobody seems to use
The dataset is valuable on its own without the model.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
I worked in a bookshop. I know all about this. I've seen pallets of books with covers torn off so they can be reported destroyed.
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
This was my issue as well. It seems like they tip toed around it in one paragraph saying usually these are books with isbn numbers so not truly collectible or very rare and then in another paragraph plainly stated that it could be foreign language or low volume books that are rare. Rare expresses a different meaning to me and I think it amounts to exactly what you said, manufactured outrage.
You don't know the books are being destroyed. There are automated scanners with page turners.
Most reporting says they are being destroyed. Page flipping is slow compared to cutting the spine and scanning the pages.
I watched Short Circuit yesterday. All I can think now is "need input".
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> Then LLMs killed stackoverflow.
I think the common consensus is that stackoverflow killed stackoverflow, quite a few years before LLMs became entrenched.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
https://www.authorsalliance.org/2025/09/07/the-anthropic-set...
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
Discussions:
2 days ago https://news.ycombinator.com/item?id=49310725
21 days ago https://news.ycombinator.com/item?id=49068738
if we are gonna pearl clutch about this problem then the headline should be about how libraries -- the entities who preserve rare books -- don't have enough funding to do this. why are we now expecting booksellers to preserve rare books? they are booksellers, in trade
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
Bad enough to steal IP but destroying books…that’s just evil
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
> "Rare" meaning inclusive of high-quality items,
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.