I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?
You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.
No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.
Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.
Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.
The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.
So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.
That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved. It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.
Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?
To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
Like I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
Well, at least afterwards the book is scanned and preserved digitally in their archives of training data.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
A little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight.
This is just one example, but it has become unfortunately common across all social media platforms.
Not under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
When you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.
I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
And the solution to the water issues can be found in any introductory textbook on the subject: a water price.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
More and more I feel like anti-AI is a bigger bubble than AI. It seems like every week it expands into a new dimension - anti-Flock protesters tearing down years-old traffic cameras that were used for research into auto accidents, etc.
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
>In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
The human tradition is storytelling. The idea that storytelling was something a company could own and other's weren't allowed to tell is very, very new in our history.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
> Why would a reduction in human writers cause a complete reduction in motivation to read?
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models.
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
They are ordering books with ISBN. So I take that they are tracking what books they have scanned or pirated already and only picking up what they are missing. As just ordering mass bulk and getting 20 of the same encyclopaedia would be waste.
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
Public libraries destroy unsold book donations all the time. I often tried to give away some old books I have online and nobody wants them. Some of these books have some nostalgic value to me so I hate to see them just get destroyed so they just lie in my shed.
I have a copy of Michael Abrash's Graphics Programming Black Book (it's like 1k+ pages) with DESTROY written in red on the sides. I appreciate that someone saved it and sold it to me for cheap :)
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.
Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?
You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.
Right, that's my point.
These generally are not books people care about. The information contained therein was doomed.
Now the information has been digitally preserved and a digestion of the information will be made publicly available.
They’re not supposed to be storing a copy. What they’re doing is destroying their copy after training a model on it.
No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.
They're absolutely keeping the digitized copies. They're not going to just train a single model.
Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.
Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.
AIs are weirdly bad at quoting stuff in my experience.
But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.
Is there a specific book from 40 years ago you have in mind? Asking out of curiosity.
Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.
There's just a few copies in a single library worldwide which is probably a national or a university library
The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.
They are useful as a source of non synthetic data, replacing synthetic data consisting of rephrasings of common topics I assume?
Also, a lot of these "rare books" are stuff like TV programming magazines from October 1994.
So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.
That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved. It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.
Around a million books are destroyed each day in the US.
Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?
To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
Whose job is it to dictate what is and isn't worth saving?
Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
My personal waste book at home is so rare, it's unique. That doesn't mean it needs preservation.
Your what now? Made from your personal waste? That does sound unique.
https://en.wiktionary.org/wiki/wastebook
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
https://en.wikipedia.org/wiki/Anecdotal_evidence
Like I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
Well, at least afterwards the book is scanned and preserved digitally in their archives of training data.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
> I doubt AI companies would use a single physical book if they could avoid it
They just don't want to pay what the copyright holders want to charge
Ai companies don't use ebooks, because they are more expensive than second hand books
They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.
I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
The entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.
i imagine its because the doctrine of first sale does not apply to ebooks.
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.
They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
If that's true, then why did they pirate so many books?
Since when do AI companies care about copyright law? They're destroying them so their competitors can't use them.
A little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight.
This is just one example, but it has become unfortunately common across all social media platforms.
Since they had to pay 1.5 billion for it
> They're destroying them so their competitors can't use them.
That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
To do what?
To not destroy rare books.
Not under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
When you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.
I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
If only this complaint was being posted by an organization ideologically opposed to copyright itself!
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
> Anthropic thought it was just a technical problem
Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
You can donate with over 20 different payment methods after making an anonymous account.
https://annas-archive.gl/donate
It's an excellent play by them, using moral outrage to the benefit of the project. When life gives you lemons...
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
> Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time?
No. They also use lots of other methods to get training data.
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
The outraged people don't care. They hate AI, and so anything surrounding AI that can be evil is evil. Books are good, AI destroys books, AI is bad.
Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.
I think people genuinely don't get that book destruction is like a pretty natural part of the book lifecycle.
And the solution to the water issues can be found in any introductory textbook on the subject: a water price.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
More and more I feel like anti-AI is a bigger bubble than AI. It seems like every week it expands into a new dimension - anti-Flock protesters tearing down years-old traffic cameras that were used for research into auto accidents, etc.
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
>In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
The human tradition is storytelling. The idea that storytelling was something a company could own and other's weren't allowed to tell is very, very new in our history.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
> Why would a reduction in human writers cause a complete reduction in motivation to read?
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
So with this one copy BS are you not allowed to have backups of the data?
I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models.
No data => No models => No competition.
IA was in hot water already even before ChatGPT came out.
> ChatGPT […] originally released on November 30, 2022
https://en.wikipedia.org/wiki/ChatGPT
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive
Yeah, great logic, that way they are sure there are no extra copies around.
I'm sure the AI companies will retain scans of the books for training on newer models
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
Yes, we know they're destroying the books. Whether this is actually a problem or another convenient "AI bad" trope remains to be seen.
https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
They are ordering books with ISBN. So I take that they are tracking what books they have scanned or pirated already and only picking up what they are missing. As just ordering mass bulk and getting 20 of the same encyclopaedia would be waste.
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
Most old books aren't worth protecting, and the publishing industry destroys oodles of them as waste that no one wants.
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
Is this a reference to Fahrenheit 451?
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
> It is 100% the consequence of stupid laws.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
Public libraries destroy unsold book donations all the time. I often tried to give away some old books I have online and nobody wants them. Some of these books have some nostalgic value to me so I hate to see them just get destroyed so they just lie in my shed.
I have a copy of Michael Abrash's Graphics Programming Black Book (it's like 1k+ pages) with DESTROY written in red on the sides. I appreciate that someone saved it and sold it to me for cheap :)
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
"You own a particular physical copy; you don't possess an abstract transferable 'one-copy license."
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
Since when do AI companies care about the law? Most of their training data is pirated.
The headlines about destruction came not soon after they got rapped on the knuckles and told "no more pirating".
> Since when do AI companies care about the law? Most of their training data is pirated.
I imagine since the law recently cost one of them truckloads of money for their violations of it?
It's giving Vishnu, but the world cannot exist without Shiva.
Aren't AI companies all about the rare book auctions now?
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
No it is not. https://reddit.com/r/Annas_Archive/comments/1vrvt9e/athome_s...
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
Use a vpn
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.