AI companies destroy physical books – let's scan rare books before it's too late (annas-archive.gl)

479 points by Cider9986 20 hours ago

824 comments:

by thread_id 9 hours ago

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.

https://en.wikipedia.org/wiki/Google_Books

https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

by probably_wrong 8 hours ago

Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.

I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.

by ndiddy 7 hours ago

At one point Google Books was supposed to act as a clearinghouse for scans of out-of-print books. You could have purchased a scan of any book on the site for a reasonable price, and libraries could subscribe to a service where the full text of all books was available. This settlement then got shot down because some research libraries and authors argued that this was anti-competitive, and instead wanted Congress to pass a law to free up the rights to orphaned books so anyone could start a competing service. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. It's not a flashy issue that would make a ton of people vote for you to get re-elected, and it won't create a ton of new jobs. The result is that now nobody outside of Google gets to see the full Google Books scans.

by allturtles 7 hours ago

Yes, this was a great tragedy. I was very sad to see academics at the time arguing against Google providing what would have been one of the greatest storehouses of readily available knowledge in the world, in favor of an imaginary alternative that didn't exist and never would.

by cube00 5 hours ago

> Google providing what would have been one of the greatest storehouses of readily available knowledge in the world

I'd be worried about how much they'd be charging for access once they had the monopoly on so many rare books.

by allturtles 3 hours ago

This is exactly the kind of pointless concern I'm talking about. How many legal digitial providers of those rare books are there now? Zero, I believe.

It's the logic of cutting off one's nose to spite one's face, which prefers that no one benefit rather than Google see any benefit.

by anon84873628 5 hours ago

Then at least there would be outrage to drive the passing of the needed legislation which otherwise hasn't come to pass anyway.

by avidiax 5 hours ago

Copyright law desperately needs a production requirement or allowance.

The copyright owner must make new copies of the work available; the price must be no greater than the original price (not inflation adjusted). And if they fail to do so, anyone may produce copies and escrow the original price (less the cost of production) for collection by the copyright holder.

That means that orphan works are effectively in the public domain. Calculus professors can ask students to get the cheaper 2nd edition, not the latest 22nd edition. And a company like Google could make scanned works available in their entirety for a small amount of money for each work. And the copyright holder still gets their end, without having to arrange a printing or hold stock.

by shagie 4 hours ago

As a photographer, do I have to make every photograph that I've ever sold available to anyone to buy forever more? Can I refuse to sell a print to someone? I wasn't famous when I sold one for $20 back in the 90s... if I became famous, would I still need to sell that at $20 (inflation adjusted)?

What happens to limited editions of print runs? Can I not make a run of 200 prints anymore because the 201st will be something that someone could request?

Does a musician have to license any song they made to anyone who asks? Can they refuse to license a song to some organization they disagree with and not have it fall into the orphan works category?

---

Amending copyright to the way you describe requires a renegotiation of the TRIPS agreement ( https://en.wikipedia.org/wiki/TRIPS_Agreement ) with all the nations of the WTO (or withdrawing from the WTO).

by avidiax 3 hours ago

Under my proposal, if someone wants to make prints of one of your old photographs, you are getting ~$20, which is ~$20 more than you are getting now. So it doesn't seem like a damage to you. If you get famous and now you could sell a print for $2,000... how is this helping society since the work has already been produced, so no new incentive is necessary?

I suppose the rule could make reference to a rival good, i.e. the 22nd and 23rd editions of the calculus textbook. It should not be reasonable for the publisher to make the 22nd edition only available for $1,000, and the 23rd edition for $250. But that definition would invite many lawsuits and chilling litigation in general. A clear definition based on the historic price is much simpler.

You can still number your limited runs, and your limited runs still have increased numismatic value over some other reproduction.

I don't envision licensing of performance rights in this system, just recordings or reproductions.

Is this a reduction of the rights granted by copyright? Yes, intentionally so. It is stripping copyright holders of the right not to copy, against the interests of society in granting that copyright in the first place.

by bigbadfeline 3 hours ago

> It's not a flashy issue

It's a very serious issue, very well known to the people with the connections and power to affect it.

> [ it wouldn't ] make a ton of people vote for you to get re-elected

The tons of people are moved by the media, people are oblivious to the tricks of that trade, for the same reasons, obviously. In other words, this issue isn't something that happens to slip below the radar, it's kept stealthy by well organized engineering and considerable expense.

> and it won't create a ton of new jobs

Nothing ever creates tons of new jobs, the "tons" are reserved for promises and other useless noise.

by p0w3n3d 7 hours ago

Quod licet Iovi, non licet bovi

Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

by misnome 7 hours ago

> Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.

by ndiddy 7 hours ago

That's why they got sued, but the suit is mainly over whether controlled digital lending is legal at all rather than their "emergency library". Archive.org lost the case on summary judgment, meaning that they could not come up with a single fair use argument for CDL that the judge found compelling enough to let the case go to trial. The full judgment is here https://storage.courtlistener.com/recap/gov.uscourts.nysd.53... but here's a couple excerpts:

> The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers.

> In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."

by SideQuark 5 hours ago

> suit is mainly over whether controlled digital lending is legal at all

No, it was not, even supported by the quotes you pulled. Libraries right now, with publisher blessing, offer all manner of controlled digital lending. The suit was because IA did it buy undercutting the publishers copy rights to that legal market. Had IA simply done what every other library has done to provide controlled digital lending, there would be no suit.

by ndiddy 3 hours ago

"Controlled digital lending" is not a generic term for "lending digital items". It specifically refers to the practice of a library digitizing physical materials in its collection, then lending them digitally based on a 1:1 owned-to-loaned ratio. The idea is that the library should be able to treat digitized versions of a book the same way it treats the physical book, and the total number of physical and digital copies of the book that are lent out at once should never be more than the number of physical copies that the library has.

In contrast to this, the e-book lending practiced by most libraries with publisher blessing involves the library purchasing special library-specific e-book licenses from the publisher. These licenses contain various contractual restrictions, such as the library having to re-purchase the e-book after a certain amount of time or after a certain number of borrows.

by allturtles 7 hours ago

There is so much misinformation/confusion about this... they go sued after lending "unlimited" copies, but they were sued (and lost) for lending exclusive copies (controlled digital lending):

> “At bottom, [the Internet Archive’s] fair use defense rests on the notion that lawfully acquiring a copyrighted print book entitles the recipient to make an unauthorized copy and distribute it in place of the print book, so long as it does not simultaneously lend the print book,” Judge John G. Koeltl of the U.S. District Court in Manhattan wrote. “But no case or legal principle supports that notion. Every authority points the other direction.” [0]

[0]: https://www.insidehighered.com/news/tech-innovation/teaching...

by palmotea 5 hours ago

> Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.

> I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.

Can you still search the restricted parts? If so there's still value to it: it helps you identify the book so you do an inter-library loan to get at the full content. Sure, it's not frictionless, but I wouldn't be all or nothing about it.

by jacekm 7 hours ago

> The project was met with significant legal challenges from authors and publishers which was eventually overcome

I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.

by chungusamongus 6 hours ago

IIRC the courts ruled that because it would be implausibe for a person to use google books previews to read an entire work (you'd have to make a whole bunch of separate accounts to do so), it could not plausibly impact the market for that work.

by doctorpangloss 5 hours ago

The consensus among copyright lawyers is Google lost Authors Guild v Google, but for some reason the media and Wikipedia do not clearly report it that way.

by titzer 9 hours ago

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space.

It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

by greyw 8 hours ago

It's data for their AI pipeline. Basically digital gold.

by dblohm7 6 hours ago

> It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

They never should have been trusted in the first place.

by ktm5j 6 hours ago

I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually necessary, and that the AI companies are just doing it out of laziness.

by shagie 5 hours ago

If it isn't scanned destructively, is it a liability?

Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain?

These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold.

It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.

by ktm5j 4 hours ago

Questions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).

by shagie 4 hours ago

The physical books that were scanned by Google were returned to the libraries.

https://btaa.org/library/programs-and-services/book-search/f...

> Will scanning harm the books?

> No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections.

That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling.

https://cases.justia.com/federal/district-courts/california/...

> Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco).

> Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).

---

The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.

by ktm5j 2 hours ago

> The AI training isn't borrowing from libraries and returning from libraries

Perhaps that's an important distinction here, I was really just trying to explain what I thought the commenter was trying to say. Calm down with the copy pasta walls.

But I still think that AI companies could make a very similar argument to the one Google made in your copy-pasta. The laziness I refereed to earlier is the fact that they haven't even tried. They could donate the books afterwards which would go a long way towards helping that argument in court.

by shagie 38 minutes ago

My copying and pasting is twofold.

First, in today's environment people will state AI hallucinations as fact or "google it yourself" as the reference. If someone doesn't know where to find that information they're left with the "someone on the net said XYZ".

Secondly, I'm not always certain that if I do provide a link to a large document that people will find the relevant section in there. And second and a halfly, if someone comes back to it in a year or two or five that the link will still be live. I've had situations in the past where I've provided a link and then the domain changes hands and the new owners of the site put up a retroactive robots.txt and make it inaccessible on the wayback machine.

---

The question for an AI company that can't return, sell, or donate the material after it has been scanned - are they then to pay to store it in a warehouse indefinitely?

After they've format shifted the content and retained the format shifted content for continued use in training, it becomes at least a very gray area to donate the original.

And even if the books were scanned like Google Books did, and they could donate them to libraries - the libraries don't want those books. These aren't rare books in that people would expect you to wear cotton gloves while handling them... they're books that have mostly disappeared from availability.

https://old.reddit.com/r/books/comments/1vugion/the_federal_... provides an example of what is being seen in a used book store:

> I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades.

Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 is technically a rare book. https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...

If you had a copy of it, I would challenge you to find a library that would accept it as a donation.

In the event that you wanted to read a copy of it, there is a copy of it in the Library of Congress. https://search.catalog.loc.gov/instances/b220c9bd-63ad-5a8b-...

by jujube3 an hour ago

The answer to his question is "yes." If the book isn't destroyed, it becomes a liability to Anthropic, because then it's no longer format-shifting.

by zardo 3 hours ago

If you're going to destroy the book anyway you'll go with the easier method for a good scan(separate the pages from the spine)

by nazgulsenpai 8 hours ago

They are also an AI company now. Why would they stop?

by JKCalhoun 8 hours ago

Also, why would they ever share their collection?

Book scans, secreted away, are worthless to the public.

by al_borland 8 hours ago

They could use them as training data, without providing access the actual books.

by jimmaswell 6 hours ago

At least we would all benefit from the books this way, so long as legal nonsense keeps the scans unavailable to the public.

by al_borland 4 hours ago

I don’t think locking the content of rare books away in the hands of corporations who only give us access to tools trained on the books, and not the actual book, is a good path forward.

This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.

by janpeuker 5 hours ago

The irony of countries blocking Anna's Archive (UK, Italy, Netherlands etc) but then it's the hackers who conserve and steward the books. Because governments can't stop AI companies from shredding history like they did with Google Books _because_ they actually tried to do it _by the book_.

by spandrew 5 hours ago

Amazon also did this for their "Look Inside" feature. To do this they had to spin up massive digital infrastructure——then realized that they could sell that infrastructure and make more money than from the books they were scanning.

This is was what ended up spinning up AWS as a business.

I dislike the idea of destroying rare books--but how rare? A digital copy has a lot more benefit.

by wrathofquan 6 hours ago

I've worked in academic libraries since 2010. I've always felt like it was a mistake for our digital library leadership to put so much trust in Google Books despite their promises to maintain the integrity of libraries (this mostly happened by the way).

At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.

by toomuchtodo 7 hours ago

Internet Archive version: https://openlibrary.org/

Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...

Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...

Web app: https://archive.org/want/?mode=donation_book

For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.

If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.

by shagie 7 hours ago

As an aside... https://www.google.com/books/edition/The_Systems_Bible/mrOsb... (and I haven't hit any "you can't read this" limits yet).

While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle. https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-...

by b112 6 hours ago

If you read the judgement against (I think) openai, the judge said that it was OK to scan the books for LLMs, if they were destroyed afterwards. That is, only one copy of the data existed.

by cladopa 10 hours ago

It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.

By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.

If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.

by nloomans 10 hours ago

> a digital copy with the ability of doing millions of copies is stored somewhere

somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”

the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.

> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.

archivists keep everything, because we don't know right now what will be important 100 years from now.

by radu_floricica 9 hours ago

> somewhere were we can't access it

By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.

Plus having the info part of a LLM makes it immediately available to literally billions.

I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.

by jjulius 9 hours ago

>Plus having the info part of a LLM makes it immediately available to literally billions.

Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.

If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".

And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.

by SkyBelow 9 hours ago

Depends upon what you want.

For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.

The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.

But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.

That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.

by chefandy 6 hours ago

> But, that little bit of data is a bit more data than existed before,

No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.

> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.

That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.

I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.

by nuancebydefault 5 hours ago

> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost

I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?

by chefandy an hour ago

It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.

by SkyBelow 5 hours ago

>all of these things are permanently lost

A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).

That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.

>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.

My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.

>That’s a false dichotomy.

I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.

by choo-t 2 hours ago

> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. Or be useful and scan book for Anna archive and other shadow libraries.

Citizen's lobbying against megacorp is a mirage.

>But it's _closer_ to being widely available, not farther.

By what metric ? The copy is now guarded by a company instead of being on the second hand market.

by winterismute 6 hours ago

> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.

I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).

by darkwater 9 hours ago

> Is it "locked"? Yes, by copyright laws,

> Plus having the info part of a LLM makes it immediately available to literally billions.

Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".

by gnfargbl 9 hours ago

> permanently locking human knowledge inside private corporate servers

History tells us that very few "permanent" situations are truly permanent.

Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.

by darkwater 9 hours ago

> History tells us that very few "permanent" situations are truly permanent.

If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.

by pixl97 3 hours ago

I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.

So physical books are rather non-permanent themselves.

by mohamedkoubaa 6 hours ago

Exactly. Calling anything on an SSD permanent is criminal.

by mvlipwig 6 hours ago

I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.

by TFNA 6 hours ago

> any important book has been duplicated by thousands, tens of thousands or even million of units.

Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.

The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.

by afpx 9 hours ago

I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.

by quietsegfault 5 hours ago

Do you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today.

I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.

by mannyv 8 hours ago

I have books, but they are just objects. They're nice objects, but just objects.

Fetishizing books isn't going to help.

In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.

by sajithdilshan 10 hours ago

Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.

by brightball 10 hours ago

Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.

by shiandow 10 hours ago

Somehow I don't think they're looking for the books that have been copied over and over.

by quietsegfault 4 hours ago

Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?

by shiandow 4 hours ago

For one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.

by pibaker 6 hours ago

> any important book has been duplicated by thousands, tens of thousands or even million of units.

It is common for academic books to have publication runs in the low three digits.

You may argue these books are not important. But how do we know if we fail to preserve it?

by quietsegfault 5 hours ago

Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?

I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.

by arttaboi 5 hours ago

With all due respect, I would say it wouldn’t hurt not to downplay this.

by jll29 9 hours ago

Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.

by GreenLightGo 6 hours ago

Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...

by mtkd 5 hours ago

>It is not a big deal

have you ever held and read an old book?

by rvz 9 hours ago

First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.

Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.

Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?

[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...

by kccqzy 9 hours ago

Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.

by quietsegfault 4 hours ago

Why is it a big deal?

Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?

by voidhorse an hour ago

Why would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.

by pshirshov 10 hours ago

Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.

by wasmperson 5 hours ago

I was also skeptical of this claim but managed to find someone who explains it:

https://downtownbrown.substack.com/p/five-fallacies-ai-and-d...

It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.

by demibabs 3 hours ago

Good article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.).

If anyone has an example, I’d love to hear it.

by voidhorse 43 minutes ago

Since we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do.

The distributed nature of the problem and the utter lack of transparency are huge factors here too.

by pfdietz 7 hours ago

Here we have another entry in the long list of "things described on the Internet that never happened".

by mistercow 9 hours ago

Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them.

But where did you hear that they’re buying “all copies”? And to what end?

by dbspin 9 hours ago

This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.

To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.

Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.

by famouswaffles 9 hours ago

>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.

I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.

by dbspin 8 hours ago

OK... I'm going to assume good faith even though your wording makes it somewhat unlikely.

Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.

Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.

Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.

by famouswaffles 7 hours ago

I think there are two issues here:

1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.

Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.

2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.

Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.

Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.

by voidhorse 41 minutes ago

How would film be any different in any capacity whatsoever?

Believe it or not, what matters here is the message and access to the message, not the medium.

by dataflow 10 hours ago

Where did you see they tend to buy all the copies? This comment is the first time I've heard of this.

by wmeredith 10 hours ago

I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?

by dataflow 9 hours ago

I could see buying multiple copies being useful to mitigate problems, like damage.

But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?

by p-e-w 9 hours ago

It’s just another lie of the type these threads tend to be filled with nowadays.

Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.

I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?

by diseasedyak 9 hours ago

It really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.

by wongarsu 8 hours ago

I wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are

by vavos 5 hours ago

I think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated

by sajithdilshan 10 hours ago

what is your source?

by ezfe 19 hours ago

I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.

Instead, they enforce the copyright and force AI companies to shred books they want to ingest.

edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.

by RajT88 19 hours ago

The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.

These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.

by card_zero 19 hours ago

Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.

by scarmig 18 hours ago

The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.

by asaddhamani 17 hours ago

But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.

by scarmig 17 hours ago

Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.

Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.

That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.

by fmajid 15 hours ago

The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.

by ralferoo 13 hours ago

> I really don't know what people objecting to this imagine typically happens to old, unwanted books.

In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.

Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.

by theshrike79 13 hours ago

So trashed, but with a bunch of extra steps then?

by tentacleuno 2 hours ago

I would presume that the extra steps incrementally diminish the possibility of the book remaining unsold, and thus destroyed or sent elsewhere.

by theshrike79 an hour ago

And it also adds costs in every step. Someone needs to move thousands of unwanted books from high end stores to lower and lower end stores. Someone needs to store them in the proper environment etc.

I do get the _idea_ of preserving books, but... people don't care. I just threw out well over a thousand books from my grandparents house this spring.

There were ~6-10 "valuable" books there. Two because I personally knew someone who wanted old war-time books and a few 100+ year old bibles. And maybe two dozen books worth saving, mostly because they were from big-name authors or had stuff that nobody would print anymore (I have detailed instructions how to make laughing gas and how to build an underground chemical lab - hobby books in the 50s were ... interesting :D )

I literally couldn't give away the rest. And I tried. It was all just "interesting, but..." - no way to justify using the shelf space for books that, realistically, nobody will actually ever read again.

by soperj 17 hours ago

They're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist.

by skeledrew 16 hours ago

Dumpster is where they go when the resellers fail to complete sales.

by scarmig 17 hours ago

The magical library in the countryside, to be painfully explicit, does not exist.

by Ekaros 17 hours ago

And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.

by scarmig 17 hours ago

Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).

by fmajid 15 hours ago

They should send them to The Internet Archive instead.

by FeloniousHam 8 hours ago

Why aren't we storming the Library of Congress? They are the real villains here.

by soperj 17 hours ago

How many times can you post the same thing in a thread?

by exe34 15 hours ago

The internet archive

by dukeyukey 16 hours ago

If it were legal they may well do that as a public branding exercise. Google already tried and got punished for it!

by skeledrew 16 hours ago

It was never available to you/us in the original form either.

by red75prime 17 hours ago

...because it is illegal to copy copyrighted material. 70 years later they might do it.

by rhdunn 16 hours ago

95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70.

There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.

In short, it's a mess.

by mrweasel 16 hours ago

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.

At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.

by novok 15 hours ago

This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.

by fmajid 15 hours ago

Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.

by svachalek 8 hours ago

Would it even be legal to share? I don't think it would be.

by mejutoco 13 hours ago

In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.

by bulbar 18 hours ago

> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?

I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.

by margalabargala 18 hours ago

Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?

by scarmig 18 hours ago

You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.

by margalabargala 18 hours ago

Right, that's my point.

These generally are not books people care about. The information contained therein was doomed.

Now the information has been digitally preserved and a digestion of the information will be made publicly available.

by halsafar 18 hours ago

Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.

by skeledrew 16 hours ago

Funnily the argument made just a few months ago by many rights holders who wanted their pound of flesh was that, if prompted a certain way, exact text could be retrieved.

by margalabargala 17 hours ago

Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.

by fluoridation 15 hours ago

>If you have a translation, do you have the book?

No, you have a translation.

>Even if, like the Odyssey, there are hundreds of wildly varying translations?

Precisely why translations are not considered equivalent to the original text.

>What about an abridged copy? What about the Sparknotes version?

An abridged copy is not a copy of the unabridged version.

>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?

No.

I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?

by margalabargala 8 hours ago

My opinion is different from yours.

If I have a translation of a book, I think I have more of that book than if I had nothing at all. It's fuzzy.

by fluoridation 8 hours ago

You don't have the book, you have someone else's interpretation of the book's contents, re-expressed into a language you can read. Both steps can involve a loss or distortion of information, either because the translator doesn't fully grasp the original language or context, or because in the re-expression they chose to leave out details that were relevant to you. The more distant the original language to yours, the more translation involves interpretation, too. A translation is really not too different from a commentary; it's just a different text.

by margalabargala 8 hours ago

That sounds to me like fuzzily having a version of the book. That is, it's more like having the book, than having nothing would be.

You seem to be arguing that a translation, etc is not literally having the book, which is something that has always been my stance as well.

by fluoridation 7 hours ago

You're trying to use the Socratic method to show that the havingness of the book is a spectrum, and I'm taking the position of a hardliner who considers that "having the book" means having the original string of symbols from beginning to end, and anything besides that is not having the book. I'm trying to show you that your line of argumentation is uncompelling to such a person.

by margalabargala 7 hours ago

I was never using the Socratic method. I was asking rhetorical questions, and then gave the (my) answer at the end.

We just have different opinions about what it means to have a book. I think that having a copy of "Pride and Prejudice and Zombies" is more like having a copy of "Pride and Prejudice", than having no book at all is like having a copy of that book. You disagree and that's fine.

by fluoridation 6 hours ago

Asking rhetorical questions as a form of argumentation is the Socratic method.

by margalabargala 6 hours ago

If one is using the Socratic method, rhetorical questions are one tool they might employ.

The reverse is not true. Just because someone asks a rhetorical question, does not mean they are using the Socratic method.

by close04 13 hours ago

The scan and destroy method is what a judge allowed to do in order to have a copy of the book in the training dataset. With the physical copy destroyed there's still only 1 copy in circulation. Once "inside" an LLM I don't know if anyone decided unequivocally that it's copyright infringement or not, and if that counts as a second copy.

There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?

by GPerson 18 hours ago

They’re not supposed to be storing a copy. What they’re doing is destroying their copy after training a model on it.

by derektank 17 hours ago

No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.

by monocasa 17 hours ago

They're absolutely keeping the digitized copies. They're not going to just train a single model.

by Natsu 17 hours ago

AIs are weirdly bad at quoting stuff in my experience.

But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.

by margalabargala 17 hours ago

Bad at quoting, good at digesting and regurgitating. The concepts are preserved even if quotes aren't.

I'd rather a digital copy exist in someone's hands than a rotting physical copy.

by subscribed 15 hours ago

But you don't have access to this copy. The digital copy is removed and all you get is paraphrased content. Some frontier models have been explicitly forbidden from recalling exact quotes in system prompt.

It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?

by sharpshadow 17 hours ago

On a similar topic are all those artifacts kept in museum storages for literally eternity. Maybe AI money can crack open access to it.

by qingcharles 16 hours ago

Many are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it.

One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.

by willy_k 18 hours ago

Is there a specific book from 40 years ago you have in mind? Asking out of curiosity.

by ipaddr 17 hours ago

Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.

by ErigmolCt 16 hours ago

I suspect rare here often means out of print or commercially obscure, not unique

by alightsoul 19 hours ago

There's just a few copies in a single library worldwide which is probably a national or a university library

by tptacek 18 hours ago

The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.

by alightsoul 18 hours ago

They are useful as a source of non synthetic data, replacing synthetic data consisting of rephrasings of common topics I assume?

by runarberg 19 hours ago

At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.

by eru 19 hours ago

My personal wastebook at home is so rare, it's unique. That doesn't mean it needs preservation.

by card_zero 18 hours ago

Your what now? Made from your personal waste? That does sound unique.

https://en.wiktionary.org/wiki/wastebook

Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.

by eru 18 hours ago

The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books.

So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.

P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?

by fmajid 15 hours ago

The AI companies’ working assumption is that if someone found it worth printing, it has enough information content to help train a model. That assumption might be invalid with some of the more rambling self-published books, however.

by asdfsa32 18 hours ago
by eru 13 hours ago

You are giving me too much credit: it's made up evidence. It's an illustration that rarity doesn't equal value, and doesn't depend on whether I actually own a wastebook or ten.

by asdfsa32 13 hours ago

Your point is understood but the crux of the issue is a bit similar to capital punishment, the argument is that risk of losing even one innocent person or useful book isn't worth taking; specially considering the value created for society as part of such risky undertaking, whatever it is exercising capital punishment or scanning and destroying books.

by eru 13 hours ago

Whenever you build a highway or a bridge or a power plant, your engineers have to put a money value on human life, or at least the worth of a statistical human life. Just to make ordinary engineering decisions.

And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.

So I don't quite agree with these taboos in the absolute.

(I'm still against capital punishment on practical grounds.)

For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.

In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.

The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.

by asdfsa32 11 hours ago

I am well aware of the Statistical Value of Life and how it impacts projects. But that is generally used for allocating preventative measures. No road is going to get approved if it is going to result in random deaths by design.

by runarberg 19 hours ago

Like I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed.

Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.

My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.

by eru 18 hours ago

Well, at least afterwards the book is scanned and preserved digitally in their archives of training data.

If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.

by enraged_camel 19 hours ago

Also, a lot of these "rare books" are stuff like TV programming magazines from October 1994.

by card_zero 18 hours ago

So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.

by tptacek 18 hours ago

That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved. It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.

by mslt 17 hours ago

Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?

by svachalek 8 hours ago

Preservation would be cool but what does that have to do with anything? The system is that the remaining copies have private owners and the private owners can do anything they want with them, including sell them or shred them.

Most of these owners aren't doing anything particular to preserve them, they're stacked up with 10,000 other TV Guides in a hoarder's moldy basement. Anyone interested in keeping October 1994's TV Guide pristine has had over 30 years to procure and protect their copy.

These particular buyers are converting their copy to a digital one that's getting some kind of use, which is better than the fate of 99.9% of the other copies.

by tptacek 5 hours ago

I don't even think ownership rights are the principal component of this situation. It's true that in a different copyright regime, Anthropic could simply publish an archive of every book they scanned (and I think they likely would do that, if they could). But the primary reason these books get destroyed is that nobody cares about them. Old, rare books are a garbage disposal problem, not a cultural one. To me the most important thing to know about this story is that the AI companies are a tiny, tiny blip in the big picture of what happens to old physical books.

Everybody imagines libraries hoping against hope to get their hands on all these rare books so they can shelve them and preserve them for generations to come. But if you donate a box of old books to a public library, there's a good chance the clerical staff handling that donation is going to roll its eyes and cart the books out to a dumpster. Large public library systems have stopped taking book donations, for this reason. "Bring them to a thrift store" is what they'll tell you.

by dukeyukey 16 hours ago

It's more that, we don't need to preserve a million copies of the same mundane book. Losing a few copies to AI training is fine.

by pfdietz 8 hours ago

When applied to ordinary mundane objects, this is the mindset that leads to pathological hording behavior. Ultimately I think it's rooted in a fear of, an attempt to deny, mortality and the passage of time.

by pfdietz 18 hours ago

Around a million books are destroyed each day in the US.

by mslt 18 hours ago

To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?

I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.

Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?

by tgsovlerkhgsel 7 hours ago

A reasonable opinion, but I'd personally strongly choose the former.

A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.

Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.

And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.

If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.

by skeledrew 16 hours ago

A museum - or any other building - can only hold so much physical stuff. How much of it do you really want preserved? How do you choose what is preserved (it's an eventually inevitable choice)? Do you save the 1980s stuff but not the 90s? Or save every even/odd year? Some other method? How much direct access do you think people need to pre-digital history?

by throwaway219450 17 hours ago

The BBC has both in some cases, but we know for sure what was broadcast when:

https://genome.ch.bbc.co.uk/about

Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.

by unleaded 17 hours ago

Whose job is it to dictate what is and isn't worth saving?

by azan_ 17 hours ago

Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?

by unleaded 8 hours ago

I don't think I have enough money or room to buy all the books I think are worth saving, and my list no doubt has some overlap with someone else's. If I did buy them, what if someone else wants to read them? If I lose it or it gets destroyed in some way, the chance it will be gone forever increases (assuming there are multiple copies). Entrusting the availability of knowledge to individuals like that sounds like a bad idea. There could be some kind of publicly funded organisation that can take care of a big collection of books, afford to keep them safe, and make them accessible to anyone.

by zmmmmm 19 hours ago

It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.

I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.

by hn_throwaway_99 19 hours ago

There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.

The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):

From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...

> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”

by ivell 17 hours ago

Can they keep backup of the digital copy?

by hn_throwaway_99 7 hours ago

It's a good question.

This site, https://copyrightalliance.org/education/copyright-law-explai..., states "It is important to note that this exception for backup copies only applies to computer programs and not to other copyrighted works, such as digital movies, music, or photographs or ebooks." But it seems unbelievable to me that they would have spent millions copying all these books and not have backups.

by freejazz 18 hours ago

> I doubt AI companies would use a single physical book if they could avoid it

They just don't want to pay what the copyright holders want to charge

by alightsoul 19 hours ago

Ai companies don't use ebooks, because they are more expensive than second hand books

by breezybottom 19 hours ago

They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.

by hn_throwaway_99 19 hours ago

The entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.

by alightsoul 19 hours ago

I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks

by warkdarrior 16 hours ago

On Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.

by steelframe 14 hours ago

This is exactly why Amazon has also been doing this acquisition and destructive scanning of millions of books for some time now.

by bawolff 18 hours ago

i imagine its because the doctrine of first sale does not apply to ebooks.

by lkbm an hour ago

With many old books, a big part of the problem is that it's non-trivial to determine who owns the copyright. Sometimes the contract would say the copyright reverts to the author after a certain amount of time out of print,but you have to go dig through old contracts to figure out whether that's the case for any given book.

by cm2012 19 hours ago

Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.

by wesleywt 15 hours ago

They are scanning "rare" books. I presume there are not a lot of copies left to throw out.

by joshstrange 11 hours ago

Rare by whose definition?

I’m not aiming this at you directly by: ISBNs or STFU

Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.

People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.

It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.

by squidbeak 10 hours ago

> Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this

BBC good enough for you?

https://www.bbc.com/news/articles/cp3rprx2wl4o

"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.

"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.

"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."

by joshstrange 10 hours ago

It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good. As for the 18th century books, that’s pure speculation. It’s like a museum saying, “yes we have they prints for sale that they keep buying and destroying but wouldn’t it be a shame if someone destroyed the actual Mona Lisa?”.

And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.

If these books are so rare and important, then why has no one cared until now to actually preserve them?

by squidbeak 9 hours ago

> It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good.

You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?

> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.

It's good to see you agree any digitization of this category of book should be non-destructive.

> If these books are so rare and important, then why has no one cared until now to actually preserve them?

(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.

by joshstrange 9 hours ago

You are conflating 2 parts of the article to make it say something it's not. No one has put forward any examples of a "lone surviving 18th century edition" being bought up my AI companies and destroyed. That was just an example of "wouldn't it be terrible if", not a "this has actually happened". It's a complete hypothetical, it's a made up scenario, it's a boogeyman. You've hung your entire 2 comments on something that has no evidence of happening.

> It's good to see you agree any digitization of this category of book should be non-destructive.

I don't. Digital or physical, it's the same, there is no difference in my mind. I have many paper books but they are art, not functional, I also have the ebooks which is what I actually read. The _ideas_ are what's important, not a dusty, decaying shell in which the ideas are contained. I wouldn't shed a tear over every library digitizing their books and destroying the physical versions, nothing is lost. More importantly, once you've bought something it's yours, yours to read, yours to display, yours to destroy. On hacker news, of all places, the people fighting _against_ first sale doctrine is appalling.

> Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.

Our definitions probably differ here but preserving is not storing a book, preserving is ensuring that even if this copy is destroyed the ideas inside live on. I think that people that hoard (actually) rare books without a thought or care to making sure the text inside is preserved for future generations out of some desire to simply own something rare are the actually monsters here. And let's dispose with the notion that booksellers are "preserving" books, they are holding inventory, inventory they were happy to sell to Amazon/etc. If AI companies were raiding museums at gunpoint we'd be having a different discussion. They are buying books for sale, if they keep them in a library at corporate or scan and destroy them it makes no difference.

by ctm92 13 hours ago

They also only scan books that are easily and cheaply available, which means they are either not rare or have no significance.

Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.

by tgsovlerkhgsel 8 hours ago

Not the copyright holders, "we the people": Copyright is an artificial legal construct that was repeatedly ratcheted up over and over again.

Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly.

Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years?

I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).

by asdefghyk 13 hours ago

RE "...Instead, they enforce the copyright and force AI companies to shred books they want to ingest...."

Why are AI companies forced to shred books?

by jeroenhd 13 hours ago

They want to have a digital copy and the judge ruled they can only keep one copy.

by drtgh 11 hours ago

You do not need to shred books to scan them. This only happens if you don't care about preserving the integrity of the books and you want to scan more cheaply.

by jeroenhd 10 hours ago

But their legal framework for being permitted to scan the books en masse (they are "transforming" the book from physical to digital) requires destruction of the original. Otherwise it wouldn't be transforming, it would be duplicating.

by ErigmolCt 16 hours ago

Copyright holders are certainly responsible for keeping unavailable works inaccessible, but shredding is mostly an industrial scanning decision, not a copyright requirement

by ranit 12 hours ago

> Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.

The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.

by breezybottom 19 hours ago

They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.

by parineum 19 hours ago

They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.

by freejazz 18 hours ago

If that's true, then why did they pirate so many books?

by scarmig 17 hours ago

Is your complaint that they follow copyright law, or that they don't follow copyright law?

by freejazz 12 hours ago

I'm not complaining

by Planktonne 13 hours ago

A company with no respect for copyright law can't pretend they respect it when it is convenient for them; it's disingenuous, and shows that there is clearly another explanation.

by parineum 8 hours ago

Because they hadn't been sued for pirating so many books yet.

by freejazz 7 hours ago

Well its not like copyright was invented last year...

by breezybottom 19 hours ago

Since when do AI companies care about copyright law? They're destroying them so their competitors can't use them.

by gpt5 19 hours ago

A little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight.

This is just one example, but it has become unfortunately common across all social media platforms.

by alightsoul 19 hours ago

Since they had to pay 1.5 billion for it

by freejazz 7 hours ago

You'd think they could've negotiated something better than $3k per work if they had actually gone to the table first.

by eru 19 hours ago

> They're destroying them so their competitors can't use them.

That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.

by tptacek 19 hours ago

To do what?

by runarberg 19 hours ago

To not destroy rare books.

by fenomas 18 hours ago

Not under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won.

So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.

by tptacek 19 hours ago

When you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.

by runarberg 18 hours ago

I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale.

Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.

by tptacek 18 hours ago

You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.

This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?

You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.

by frm88 10 hours ago

You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.

Source? You state that in a tone that implies you have verifiable knowlege of this. All the information I found says that the exact number, titles and authors are under NDA.

by svachalek 7 hours ago

Given no information, would you assume companies that are buying "huge quantities" of "rare" books are getting first edition Jules Verne by the crate, or 1989 Highlights magazines recovered from dentist offices?

by runarberg 6 hours ago

Both. I think they are just buying whatever they can get their hands on most of it is going to be junk, but one in a million is going to be a rare and valuable book which we didn’t even know existed.

After scanning and destroying 10 million books, AI companies will have destroyed 10 such books.

(I am obviously assuming a Poisson distribution here where I pulled the parameter p = 1/1000000 out of thin air; point is p is non-zero; and at this scale the undesirable event is bound to happen a bunch of times).

by tptacek 7 hours ago

It's in the 404 story.

by runarberg 8 hours ago

I assume public libraries know what they are doing and know which book they are destroying, that they keep an accurate inventory and hire professionals maintaining said inventory and marking which books can safely be destroyed.

I assume no such things of AI companies.

by tptacek 7 hours ago

Obviously that's not true. Newark is literally telling people to talk all their old rare books to thrift stores, which will immediately throw them in dumpsters.

by runarberg 6 hours ago

This is a bad comparison. Public libraries are trying their best to preserve rare copies, despite being overwhelmed and underfunded. If they come across a rare copy chances are it will be preserved. When an AI company comes across a rare book, it will not recognize it as such.

There is also a fundamental difference in intention. AI companies are seeking out old books and will destroy them. Public libraries are trying their best but simply don‘t have the resources to find rare books in every collection.

And finally there is a fundamental difference in scale. Public libraries are not buying millions of copies to destroy, completely eliminating the chance they will ever be discovered.

---

PS. I don‘t like the anti-academic tone of your post. Librarians know what they are doing, their expertise is valuable, and when they act according to their specialized knowledge it does have a positive effect on the world.

by tptacek 6 hours ago

What comparison? I'm talking about libraries as they actually exist and you're talking about what you think libraries should be doing. I'm arguing positively, and you're arguing normatively.

I'm not "anti-academic". I'm just very involved in my local community and I read the annual reports from our library system. We accept book donations, like a lot of suburban library systems do (from our extremely book-y community), and we are open about the fact that most of those books get trashed. We do not have the personnel on staff you claim libraries generally do.

I would say that between the two of us, I'm the one arguing more respectfully about libraries. I see them as real institutions with real pressures and constraints that do an important public service, and you see them as an instrumentality in your argument against AI.

by runarberg 4 hours ago

You were the one who brought public libraries into this conversation, in an attempt to legitimize this practice by AI companies. I assumed this attempt was you saying the two behaviors are comparable.

I pushed back on that comparison. These behaviors are in fact not comparable. What the AI companies are doing is bad actually, and what libraries are doing is, while unfortunate, acceptable, given the limited resources they have.

by tptacek 4 hours ago

Yes, I brought public libraries into this conversation, because they daily destroy hundreds of thousands of books, most of them without any professional evaluation whatsoever. I don't have a problem with this, because I understand a little how the book trade works, and whatever romantic idea people have about the preservation of books, it has little to do with what books actually are.

What I don't understand is your response to this. An AI company digitizes a book before destroying it: "bad actually". A public library takes a cartload of books to a dumpster without so much as opening the front cover of any of the books: just fine.

by dukeyukey 16 hours ago

Can you explain why AI companies destroying these is worse than a library or bookseller destroying these?

by runarberg 8 hours ago

Public libraries staff professional archivist, they take careful inventory and know exactly which book they are destroying.

AI companies buy boxes and boxes of books with unknown content staff anybody who can operate a scanner and have no idea which books they are destroying.

If a public library comes across a book they didn’t know they had, it is very likely that somebody will notice and know how to continue, to find out if this book is worth saving etc. AI companies will treat this book exactly like any other and destroy it to feed the plagiarist machine.

by kasey_junk 7 hours ago

This is not what most public libraries do with donation books. They mostly just recycle them.

The very act of cataloging donation books takes a huge amount of resources that the collection managers don’t have.

Many public libraries don’t have a collection management department at all! They outsource that function entirely to vendors and those vendors mostly source new books and spend resources preparing them for the hard use of a library (changing bindings, uploading and cleaning metadata, tagging, etc).

Decommissioning is done by volunteers who just look at the condition of books and chuck the bad looking ones in a bin for recycling.

My wife worked in this industry on the vendor side, your model of public libraries is closer to a small subset of certain big city libraries narrowed to their rare and research departments. The median book bought by a library is a Danielle Steele romance novel packaged for library consumption and sent to a small town library system that will be lucky to have a single professional librarian for the whole system.

by tptacek 7 hours ago

At this point I believe you're just sort of making up a fantasy of what you think a public library should do, not what they actually do.

by dukeyukey 6 hours ago

Rare and valuable books are not going to be in the giant pallets AI companies are buying. Most public libraries do not have professional archivists or anything like that, not so charity shops, or booksellers.

by runarberg 6 hours ago

No, but they are going to be intermixed in boxes from an old estate, or a used books store which is closing up, etc.

> Most public libraries do not have professional archivists

Most reasonably sized library systems do in fact. Even small ones have trained librarians with an some degree in library science who took at least an introductory class in archiving as a part of their degree, and very likely has some idea how to prevent rare books from being destroyed.

by tptacek 5 hours ago

Right, and what happens to those boxes from old estates are that they get put into dumpsters. The only difference here are:

* It's happening on a much smaller scale

* They're actually digitizing the books before they destroy them

This obviously isn't about the books. It's about people want reasons not to like AI companies.

by runarberg 4 hours ago

The scale matters. It is disingenuous to just remove it from the equation as if it doesn’t.

As does the intention matter. The AI companies are doing this because they want to profit off of it. They don’t have to do this, and the fact that they do is part of what makes them bad for humanity.

by tptacek 4 hours ago

Right: the scale matters, which is why it's very weird to take AI companies to task for something the public library system does on a far greater scale. And, in fact, the AI companies do have to do this: they're required by law to destroy these books.

by breezybottom 10 hours ago

"Can you explain why Soviets persecuting Jews is any worse than Nazis persecuting Jews?"

by dukeyukey 6 hours ago

If someone was ok with Nazis killing Jews but not Soviets, I am going to think they don't believe it's the killing that is the problem.

by HedonicEscal8r 19 hours ago

If only this complaint was being posted by an organization ideologically opposed to copyright itself!

by raincole 19 hours ago

> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.

What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.

Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.

by remus 11 hours ago

> ...AI companies will still do scan'n'destroy because it's just cheap.

There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.

by bondarchuk 14 hours ago

We the people in the society who have the power to make laws through democratic means are the ones locking these books up.

by wotamess 17 hours ago

"want to ingest"

Not "need to ingest"

Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred

So in the end it's really a Congress problem as usual

by sophacles 19 hours ago

If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".

by signa11 15 hours ago

mr. vernor-vinge's "Rainbows End" is oddly prescient ! Highly recommended nevertheless.

by customguy 17 hours ago

> force AI companies to shred books they want to ingest.

Nothing forces them to shred books, they do it because it's slightly cheaper that way.

by hparadiz 17 hours ago

There was a court case where they said that if they copied the books it's not fair use because they didn't own it but if they bought physical copies and then destroyed them then somehow it was fair use because it fell into the niche of personal backups. I forget the details but basically they buy one time prints and destroy them immediately.

by customguy 12 hours ago

I had no idea about that, or how fucking bad this actually is:

https://en.wikipedia.org/wiki/Project_Panama

So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.

by skeledrew 16 hours ago

They do it because, for each work, they bought one copy, which they scan and no longer need the physical version of, and would be in copyright violation if they keep more copies than they bought.

by annapanna 15 hours ago

>and would be in copyright violation if they keep more copies than they bought.

They can contact the copyright holder and ask/buy a license to make multiple copies.

by hparadiz 15 hours ago

yes and they can ask for a pony and then a unicorn too

by skeledrew 14 hours ago

For what?

by postepowanieadm 17 hours ago

By destroying them they don't copy only convert them into another format.

by ajsnigrutin 13 hours ago

It's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money.

On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.

by watwut 15 hours ago

Copyright allows you to sell book you have and does not force you to shread it.

They are not forced to shread them by copyright.

by wesleywt 15 hours ago

I was wondering what the pro book shredding take was going to be. Why destroy the book after scanning? You can create a beautiful library of rare books with all the AI debt bubble.

by jacobo37 19 hours ago

this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...

by rpdillon 19 hours ago

Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.

Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.

The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.

Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?

by alightsoul 19 hours ago

This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case

Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule

by rpdillon 11 hours ago

> publishers and probably authors no longer exist

Who is going to claim a copyright violation?

by Sha1rholder 18 hours ago

> Anthropic thought it was just a technical problem

Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?

by dukeyukey 16 hours ago

I imagine they know the public would react like the public is reacting. Like obviously this is not worse than what libraries and thrift shops do daily, but it is bad optics.

by Sha1rholder 13 hours ago

I don't think so. Libraries and thrift shops usually don't consider it a good idea to damage those precious, hard-to-reprint books. None would say anything if Anthropic only destroyed 1 million Harry Potter or Foundations

by footydude 14 hours ago

> "ai" companies are simply burning our cultural context ...

Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.

Example for the UK - British Library guidance: https://www.bl.uk/services/legal-deposit

by Ekaros 17 hours ago

Many of them don't have current publisher because no one wants to buy them. I would bet that vast majority of these books have no commercial market.

by jscd 18 hours ago

Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?

Also, who’s forcing AI companies to “ingest” books in such a destructive way?

Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.

by scarmig 17 hours ago

Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.

Anna's Archive, for one, would be more than happy to host at no cost to the author.

by jeroenhd 13 hours ago

AI companies are buying the physical books, they can turn them into confetti if that's what they want to do. If the physical books are running out, the authors can print and sell more. Or they can sell digital copies so the information is not lost.

The law is currently forcing these companies to destroy the books after scanning them.

by jscd 9 hours ago

> The law is currently forcing these companies to destroy the books after scanning them.

No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)

by NishanStepak 6 hours ago

Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.

by qarl2 6 hours ago

Given all the bad PR around this issue - you'd think they'd do the 10 seconds of examination "is this rare" before putting it in the cutter.

I'm genuinely surprised they don't.

by hakanensari 5 hours ago

The price is more or less an indicator. If they can shred a $1000 book, kudos to them or their budget.

The idea they're simply shredding human history is too one-sided and alarmist. There is a huge long tail of rare-ish old and scrappy books out there that are definitely not the last copy of anything.

That said, I'm pretty sure a very small % of these books fall into the mistake category. But life!

by qarl2 4 hours ago

I've been noticing a lot of alarmist stuff in the anti-AI community, too.

I love how Oregon showed how to solve the electricity problem and is being entirely ignored. THEY'RE STEALING OUR ELECTRICITY!!!

by akudha 6 hours ago

Amazon has been getting bad press for years now - union busting shenanigans, poor working conditions, low quality products, customer reviews being gamed etc. Why would they care? Unless bad press directly and measurably affects their bottom line, they wouldn't care. Just this week there was news about a massive datacenter they are planning whose power generators would become single biggest polluter in the country, if it gets built.

If they ever cared enough about bad PR, they have much bigger items to think about, like the proposed data center. Rare books probably wouldn't even make it their top 5 bad-press list

by kfrzcode 5 hours ago

I'm surprised you're confident they don't.

You think someone's picking up a first edition Steinbeck and destroying it so it gets into the next training corpus?

Or is it like, technical manuals and incredibly dry almanac content?

I think the details actually matter here. I'd love to see a realtime list of these titles being scanned and processed. Then we'd know if anyone actually gives a care.

That said I'm willfully ignorant of most "headline news" so I'd be curious to see how heinous the issue really is.

by ForHackernews 5 hours ago

> You think someone's picking up a first edition Steinbeck and destroying it so it gets into the next training corpus?

I doubt the low-paid workers stripping, scanning and shredding the books are checking what anything is. I'd imagine they are dumped with a big crate of old books and an hourly quota they have to rapidly process.

Arguably old technical manuals with limited print runs are even rarer or more valuable to historians. We have lots of copies of Steinbeck, albeit not first editions.

by qarl2 2 hours ago

> I doubt the low-paid workers stripping, scanning and shredding the books are checking what anything is.

You don't think they are, at the very least, checking to see if the book has been scanned before?

A search like that could very easily be extended to check for rare books too.

Why do you assume it's not being done?

by voidhorse 36 minutes ago

Why assume it is?

We don't know, and the fact that the process is so opaque and would have been totally undisclosed if not for journalists doesn't exactly make me confident they are "doing things by the book" on this one.

by cube00 5 hours ago

> Given all the bad PR

Companies are immune to bad PR in this fast news cycle. Everyone has already moved on to the next injustice in less then 24 hours.

Surprisingly you'll even get downvoted for "living in the past" if you attempt to bring it up later.

by lkbm 43 minutes ago

Companies discovered that they can pour tends of billions into clean energy and still be hated for destroying the planet. They can use a small amount of water and be blamed for draining the Great Lakes.

Why would they avoid doing things people dislike if the PR outcome is the same?

by Teever 5 hours ago

I'm kinda surprised too, but also not, because they tried to hide what they were doing.

So that tells you that they know what they're doing is wrong / controversial and that they needed to shape their actions based on that fact.

The course of action they took tells you a lot about the people making these decisions.

by ziyadb 10 hours ago

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

by hypendev 10 hours ago

>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

How tho?

Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.

Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.

And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.

by Mudbugs 9 hours ago

Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today)

Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.

The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.

by WarmWash 9 hours ago

I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.

The internet made this even worse by celebrating people who only have to say the right thing.

by p-e-w 8 hours ago

This would be more convincing if there was any agreement on what the right thing supposedly is.

The major lesson I myself learned as I got older is that “you should do the right thing” is a rephrasing of “you should do what I want you to do”.

by andsoitis 8 hours ago

> This would be more convincing if there was any agreement on what the right thing supposedly is.

The internet has amplified voices who say things to signal rather than do things to change.

That's orthogonal to knowing what "the right thing" is.

by Lerc 7 hours ago

The increase in people who believe that perception is reality has the logical consequences of people shouting their opinions loudly enough so that it becomes 'the right thing'

by bix6 8 hours ago

That’s just not true. There is a clearly a right thing to do in many circumstances.

by nmz 7 hours ago

Knowing and enacting are two distinct things, "you should do what I want you to do" means you're already doing the wrong thing, and yes, there is an agreement on what the right thing is, its what is ethical, that is, what non profits and archivists are forced to do.

by p-e-w 6 hours ago

I can’t believe that in the multipolar world of 2026, where hundreds of conflicting world views coexist, where universalism is being disproven on a daily basis, it’s still possible to read things like “there is agreement on what the right thing is”.

This is as blatantly false as claiming that the Earth is flat, and the fact that there is no such agreement (descriptive moral relativism) has been firmly established in philosophy for well over a century.

by nmz 4 hours ago

If you think ethical arguments like "murder is bad" or "human knowledge should be preserved" and a demonstrable falsity like "the earth is flat" are equivalent arguments then you are so far gone you might as well be a flat earther.

Blind futurists and AI cheerleaders scare me on how cavalier and how many crimes against humanity they ignore.

by areoform 7 hours ago

    > The example book of Old books of agriculture is probably not that important today
If I may be flippant, not to you but to the sentiment, skill issue.

We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.

And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era?

What if... you just want to learn what your ancestor was doing at such and such time?

What if you want to find clever techniques for robot arms to work with food crops in space?

Or, heck just the alpha from a hedge fund point of view of finding old climate patterns and... :)

Your ability to make the most of knowledge is only limited by your imagination.

    > Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
I don't understand what you're trying to say here.

    > The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
https://en.wikipedia.org/wiki/No-kill_shelter

re: saving books, at a personal level, I try to use the excuse of work to find, read, and do stuff with old books,

https://1517.substack.com/p/powder-and-stone-or-why-medieval

And yes, people still care. And people who care do things.

by smallerize 8 hours ago

Well if you're appealing to librarians:

"Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on.

Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...."

https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...

by Mudbugs 8 hours ago

I don't get the point of this one. We not talking about "retiring books", we talking about "rare but still fine books".

I don't think many Liberians like the idea that they have to do it; it is just one of those things that has to be done, sadly.

by smallerize 8 hours ago

I don't think that changes the message at all. If a librarian doesn't like routine book destruction, they would still feel worse about this different thing that is happening.

by tingletech 7 hours ago

I don't know what Liberia has to do with it.

Weeding (deselection works) is a fundamental part of collections management. Every trained librarian is going to understand this.

The interlibrary loan system has mechanisms in place to make sure the member libraries keep two copies of each work in each region. Collection managers consult these databases during weeding to make sure they don't deaccession the last copy.

If the monograph was never collected by a library and it gets caught up in a destructive scanning project then I guess it was pretty "rare" in a literal sense. "Rare Book" in library land is sort of a term of art and I'm not sure if the books in these destructive scanning projects meet the criteria.

by 1234letshaveatw 7 hours ago

I'm not a bsky person so I didn't click though but library books aren't "retired" to a farm upstate. My SO works at a library, they are mostly shredded. This is a nothingburger

by demibabs 6 hours ago

Also, why are people acting like Anthropic digitizing and destroying a rare book is what makes it inaccessible to the public?

If they were simply buying the books and doing nothing with them, the public wouldn’t be able to access their copies anyway.

by voidhorse 31 minutes ago

Because destroying the book makes it possibly inaccessible permanently because we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that.

If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up. Hopefully someone will at least be able to purchase their digital scan and hopefully the scan is of decent quality and clearly indicates the provenance of the text.

by pwdisswordfishq 9 hours ago

They should be happy to have a job at all in the Republic of Liberia.

by n4r9 9 hours ago

A quick search suggests that most "weeded" books are sold on or donated rather than burnt or sent to landfill. Do you have sources that say otherwise?

by Mudbugs 8 hours ago

Just go to a local library and ask a librarian or anyone who has a lot of books and tries to give them away; sadly, in most cases, they pick the valuable ones, and the rest just get sent for destruction (Burning).

It is pretty standard procedure.

by alightsoul 6 hours ago

The difference everyone is missing, is that now there is commercial incentive to burn books aka destroy them after scanning. It's now a profitable business to do so, not something that only has to be done to clear up space

by 1234letshaveatw 7 hours ago

Ours has a volunteer run used bookstore that sells a few for pennies on the dollar. Vast majority are shredded

by vin047 6 hours ago

>it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge

This sentence reminded me of “The Things” by Peter Watts. The Thing in that short story believes it’s actually the good guy and decides to commit “violent integration” for the sake of humanity.

I’m not sure I buy the links premise that Anthropic is doing this with malice, but if it were, would be an eerie parallel.

by starkd 9 hours ago

Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.

by sergimansilla 9 hours ago

No, they are destroying them because that’s the way that they can keep a digital copy legally.

by beering 8 hours ago

This is not true. Google Books does not destroy books. No court has ruled that you must destroy books to legally keep a digital copy.

by tingletech 6 hours ago

I think under first sale doctrine you have a much stronger case with destructive scanning. Google Books, HathiTrust, and Internet Archive's book scanning project have had a lot of legal expenses.

by voidhorse 8 hours ago

You don't know that. A major part of the problem here is the lack of transparency.

Keep in mind we wouldn't even know this was happening were it not for investigative journalism.

by andsoitis 8 hours ago

> we wouldn't even know this was happening were it not for investigative journalism.

Pretty feeble investigative journalism if they cannot give us the names of even 5 such books we are supposed to be outraged about.

by shagie 8 hours ago

While not an authoritative source... https://old.reddit.com/r/books/comments/1vugion/the_federal_... (and not actual book titles)

> Correct. I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades. It would have eventually found its way to the discount rack and then the dumpster. At least this way we're getting some money in that we can use to buy actual cool books/collections, pay salaries and bills, etc.

So... something like Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 - https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...

I make no claims that that's a such a book that's been ordered, but that's the type of book that the reddit post references.

by alightsoul 6 hours ago

That data is useful as a source of scientific knowledge even if it's not current. Although it's probably already online, they probably don't want to download it and risk getting another copyright lawsuit

by fortran77 9 hours ago

They are destroying them because it’s easier to scan them if you slice the binding.

by eloisant 9 hours ago

I didn't fully understand why but apparently there is also a legal reason to destroy the books, it makes it them less likely to be considered copyright infringement.

by beering 8 hours ago

Not really. Selling the book onward does seem legally dubious but legally nothing (yet) prevents you from storing the book in a warehouse. obviously it’s cheaper to dispose of them.

by cdkmoose 9 hours ago

And very hard to re-assemble after you have done that. They are not in the book binding business.

by voidhorse 8 hours ago

> This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.

But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.

Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.

If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.

by joshstrange 8 hours ago

Are you under the impression they scan the book, train on it, then destroy the digital copy? Because that's not what's happening. They scan it, and hold it forever to train future models on. The scan still exists, not available to the general public but that's no different than if they had bought the books and kept them in a private library closed to the public.

by anjel 8 hours ago

Destroying the physical copy inflates the value of the retained digital copy

by chrisjj 9 hours ago

> This way, it will live forever instead, scanned

But unreadable by humans, right?

by mjdv 8 hours ago

They're scans. Those are human-readable. They probably won't make them available to the public, which is the exact same state they would be in if they bought the books and just put them on a bookshelf in a warehouse.

by alightsoul 6 hours ago

That was due to negligence. Now it is profitable to keep the books private, and actively deny others access to them.

by ceasesurthinko 5 hours ago

Ah, yes, because we all know OpenAI and Anthropic are famously just front operations for the high-end antiquarian book trade.

by merely-unlikely 2 hours ago

High-end antiquarian book traders is a genuinely hilarious way to describe them and I shall refer to them as such henceforth.

by alightsoul 4 hours ago

That's an in house operation that benefits them, not a front.

by RIMR 9 hours ago

Those books at the flea market aren't "rotting", they are available for sale.

You seem perfectly fine living in a world where your flea market is devoid of books.

by otherme123 8 hours ago

A lot of them are rotting. We are not talking "one of the three living copies of the first edition of Joyce's Ulysses". Rather "1956 statistics of the cultive of yuca in 'some small village from Mexico': a boring analysis". Those books have value to train LLMs as they are 100% free of AI text, but has been collecting dust (or rotting) in someone's room for decades, and no human is buying them even for 10 cents.

Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.

by hypendev 7 hours ago

No, they are not. A lot of them are rotting and get thrown or given at the end of the day.

by dumb_notsmart 8 hours ago

Oho you think they're buying common books? Books they can find online already?

by ro_sharp 9 hours ago

You’re equating something being ubiquitous and affordable with it not being valuable.

Yes, there are plenty of books, many were printed, many have lasted a very long time (plenty over 100 years!).

That says more about the success and utility of the technology than it does about whether individual books should be shredded.

by Aurornis 9 hours ago

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,

These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?

Your local library throws out books every year and nobody thought twice about it.

by mindcandy 9 hours ago

They are ridiculous because the reality of the situation is boring. Boring doesn’t drive engagement. So, all the clickbait headlines and ragebait comments imply scandal. If they didn’t, they wouldn’t get attention.

After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.

This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.

by alightsoul 6 hours ago

Does china developing ai powered missiles sound ridiculous? It sounds ridiculous to me because the US is surely doing the same, It would be boring if we knew the us was doing the same. It is emotional to think about china destroying the US when it is just the antithesis of American exceptionalism. It reminds me of what Dario said about their stance on open source.

by merely-unlikely 2 hours ago

That doesn't sound ridiculous at all (maybe reductive but not ridiculous). And the US doing the same doesn't make it any less interesting either. In fact the US's use of AI in developing target banks in a recent conflict caused quite a lot of consternation when publicly revealed. Ukraine has also recently used AI enabled autonomous drones to make an entire area one big "kill zone", no humans required. Not so ridiculous or boring if you ask me.

by snickerbockers 9 hours ago

>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?

Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?

>Your local library throws out books every year and nobody thought twice about it.

When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.

by BeetleB 8 hours ago

No. Almost none of the books you donate to the library get to the shelves. If they can sell it, they will. Otherwise they are thrown away.

I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".

by merely-unlikely 2 hours ago

> "patrons should know how to throw away their own trash"

That makes me a lot sadder than AI companies destroying the books but at least preserving (most of) their knowledge.

by BeetleB an hour ago

It's the reality. If you're old enough to remember when people subscribed to magazines ... how often did you throw out a magazine?

How often have you thrown out manuals for products you don't use any more?

by Goronmon 9 hours ago

As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices.

What happens when the books don't sell?

by pfdietz 8 hours ago

Our local library-connected biannual book sale puts a large dumpster, about the size of an 18-wheel truck trailer, outside the warehouse where the sale is conducted. At the end of the sale it gets filled to the brim with waste inked cellulose, and sits there open in the weather until it's taken away for disposal.

by merely-unlikely 2 hours ago

Is that intended as an implicit threat - buy or we trash?

by p-e-w 8 hours ago

They are thrown away, which has been happening since forever, and nobody ever gave a fuck until AI got involved.

by snickerbockers 6 hours ago

I think it hits different when you're buying large quantities of used books with an eye towards ones which aren't readily available online, and with the specific intention of destroying them.

If all these companies were doing was burning star wars tie-in novels and harry potter sequels nobody would care. That's not their goal because they already have those in their training set. The whole point here is to find rare or underappreciated books from the pre-digital era which nobody ever made publicly available in a digital format.

BTW destroying them isn't even necessary for scanning. It's the easiest way because removing the binding and turning it into a flat stack of papers solves many problems but there are actually dedicated book scanners designed to hold open the book while its photographed, and un-curling pages in post-processing was already a solved problem long before people were using AI to correct images.

by ceasesurthinko 5 hours ago

Realistically the vast majority of those books haven't been touched for half a century and won't ever be read by a human again. It is a price to destroy these, but it's a price I would pay for progressing models towards AGI and the billions of lives it will save.

by johannes1234321 9 hours ago

> When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.

Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.

Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.

by infecto 9 hours ago

As someone who has gone to many many used book sales over decades… many of the books at a sale never get sold, guess what they usually get tossed in the dump. This includes your local library book sales. Books are heavy and worthless. Cheaper to throw away the ones that nobody picks up in a sale.

I know it comes at a shock but truly most books are absolutely worthless.

by jonhohle 9 hours ago

The average lifespan of a library book is 26 loans. For a popular book that could be less than a year.

by brookst 9 hours ago

No, that’s the hyperbolic reaction clickbait wants from you.

Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.

Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.

Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.

by dd8601fn 8 hours ago

I don’t think any of these people know.

All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.

Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.

And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.

by andsoitis 8 hours ago

> Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.

Then they should list the names of these books otherwise I say they're alarmist.

by dd8601fn 7 hours ago

If they have, I haven’t seen it.

It’s very possible I’ve just missed deeper reporting, obviously.

But otherwise I agree. I’m neither losing sleep over it or just trusting that these (historically kinda scummy) businesses are actually behaving.

If there’s a serious problem I’d like to see something more concrete. Same for hand-waving the question.

by pessimizer 8 hours ago

I heard from the first stories that virtually all of these books being ordered have ISBN numbers. Books that are rare that have ISBN numbers are rare because no one wanted them 99.9% of the time. Somebody wants every book, but you'd spend many, many years finding that somebody.

by VanTheBrand 8 hours ago

If they have no value why are they being acquired and scanned?

by tempestn 26 minutes ago

Any original text has value for the purpose of AI training. The argument is that they have no other value.

by merely-unlikely 2 hours ago

A flagship LLM today is trained on tens of trillions of tokens, the equivalent of hundreds of millions of 100k word books. No human has that kind of appetite.

As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.

by joshstrange 8 hours ago

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).

> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.

by pmarreck 8 hours ago

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Hyperbole much?

Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!

by Leynos 9 hours ago

If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.

by leonidasrup 10 hours ago

A small change in the copyright law would fix this problem. Something like:

If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

by infecto 9 hours ago

Puts the burden on government to store what is probably 90% worthless material.

Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.

There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.

by ceasesurthinko 5 hours ago

You don't need a massive government program for this.

Just post on r/DataHoarder: "Free 16TB NVMe SSD to anyone who indexes and mirrors the entire out-of-print 20th-century physical archive." The problem would be solved by next Tuesday. With probably 10x redundancy and people willing to do it for free for fun.

by pessimizer 8 hours ago

> Puts the burden on government to store what is probably 90% worthless material.

That's what governments are for.

by shagie 5 hours ago

That's what governments already do.

https://www.copyright.gov/mandatory/

> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).

> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.

> Mandatory deposit applies to any work published in the United States. This requirement does not apply to works first published in a foreign country until they are published in the United States. Copyright registration is optional, but it provides additional legal benefits and fulfills the mandatory deposit requirement with the submission of the required copies.

by infecto 8 hours ago

Since we are talking about a US perspective do you have evidence that backs this up? It just comes across as an empty statement. The government is the will of the people and I personally like the idea of fixing copyright instead of making the government store how to use windows 95 books.

Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.

by brookst 9 hours ago

Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?

How does this help anything, except create more work to throw in the trash?

by ro_sharp 9 hours ago

This is already required for new books published in the US, and has been for more than one hundred years.

It’s called “mandatory deposit”

by toast0 4 hours ago

Case law seems to be that mandatory deposit is unconstitutional, fwiw.

https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland

by shagie 3 hours ago

That revolves around print on demand for books that are out of copyright or where the copyright has been abandoned.

> Background Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions.

If you want a physical copy of The Sorrows of Satan, you can buy it from them.

Their argument is that the Library of Congress already has a copy of the book ( https://search.catalog.loc.gov/instances/a0f8fcfe-a255-55d2-... ) and having them deposit it again would be unnecessary.

> The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment.

> Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment.

That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.

by jjkaczor 8 hours ago

Actually - most jurisdictions that issue a publisher a unique root ISBN number have a stipulation that anything new published using that number must have a copy sent to them.

Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.

by ndriscoll 9 hours ago

Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive should be available for research use (including ML training) at cost.

by shagie 5 hours ago

That's already the case, though the "digital rather than physical" as a preference could be something that legislation would improve.

https://www.copyright.gov/mandatory/

> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).

> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.

----

> Acceptable Formats for Deposit of Electronic Works

> The deposit of electronic works is arranged with the Acquisitions & Deposits division.

> For electronic-only works, submit the best edition in accordance with the formats listed in the “Electronic-Only Works Published in the United States and Available Only Online” section of the Best Edition Statement (PDF, 135 KB).

> For works subject to a grant of special relief, unless otherwise specified, the Library will accept an appropriate “preferred” format listed on the Library of Congress Recommended Formats Statement. Such files must contain no measures (such as digital rights management [DRM] technologies or encryption) that control access to or prevent use of the digital work.

> For more information about electronic deposit, see the above FAQ “When can I make an electronic deposit of a work?”

---

> When can I make an electronic deposit of a work?

> Works may be deposited in a physical format in accordance with the Best Edition Statement, which can be found in Best Edition of Published Copyrighted Works for the Collections of the Library of Congress (Circular 7B) (PDF, 135 KB).

> Works may be deposited electronically in certain circumstances:

> The Copyright Office issues a written demand for an electronic-only book or serial. If your work is published only online and the Office sends you a written demand for mandatory deposit of the work, you must deposit the work electronically.

> The Copyright Office offers you electronic deposit as an alternative to depositing a physical copy of the work. If you receive a letter offering special relief to deposit a work in an electronic format instead of sending physical copies, follow the instructions in the letter or agreement.

---

https://www.loc.gov/preservation/resources/rfs/

https://www.loc.gov/preservation/resources/rfs/text.html

by ndriscoll 4 hours ago

It is not. We currently give ~infinite copyright automatically for nothing in return:

> Neither the deposit requirements of this subsection nor the acquisition provisions of subsection (e) are conditions of copyright protection.

https://www.copyright.gov/title17/92chap4.html#407

by shagie 4 hours ago

Copyright protection is automatic.

Publishing of copyrighted material requires that it be deposited with the Library of Congress.

by ndriscoll 4 hours ago

Right, that's why I said we should make copyright require deposit and registration (like it used to). Publish your work without its copyright ID for people to use to reference the LoC database? It is now public domain.

by shagie 4 hours ago

That would require a renegotiation of the TRIPS Agreement and the Berne Convention with the rest of the countries of the WTO.

https://www.wto.org/english/tratop_e/trips_e/ta_docs_e/modul...

    (iii) Automatic protection A key feature of the Berne Convention, and thus also of the TRIPS Agreement, is that copyright protection - unlike most other forms of IPRs - may not be subject to any formality of registration, deposit, or the like. This principle is contained in Article 5(2) of the Berne Convention, that has been incorporated into the TRIPS Agreement.
by ndriscoll 3 hours ago

So? Of all the aggressive things the US forces upon the world (or just unilaterally does, ignoring agreements), undoing its own bad policy would be a drop in the bucket, and would be doing some good for once. I'm sure if we just did it, others would respond tit-for-tat and require registration of our material, and then mission accomplished.

(And at the end of the day, sovereign people are never required to do anything. The concept of international law is an oxymoron)

by shagie 3 hours ago

I don't believe it is a bad policy that copyright on anything that is copyrightable is automatic (I don't need to register this comment with the Library of Congress).

... And it would require renegotiating the treaty with all of these countries so that AI training is easier. https://www.wipo.int/wipolex/en/treaties/parties/231

... Or it would require the US to withdraw from the WTO and pass new laws for how copyright works.

I don't believe that neither the renegotiation nor the withdrawal would be something that would be done.

... And I believe that automatic copyright (as has been part of the Berne Convention since 1886) is a good thing.

by ndriscoll 3 hours ago

It's not to make AI training easier; AI training is already happening. It's perfectly easy for them.

It's to preserve our heritage and knowledge. Automatic copyright is what causes information to be lost. I don't suppose you're going to submit your comment to LoC or otherwise keep it available for 70 years after you die? Does everyone remember to submit their code they publish?

We're not losing books because of AI companies. We already lose them because the law makes it so only groups like Anna's Archive can save them.

by brainwad 8 hours ago

The Library of Congress already has a copy of every book published in the US. How would this help?

by toast0 7 hours ago

If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.

by bhelkey 4 hours ago

> it would be easier for them to distribute the work after the copyright of the work expires

Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years.

Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120).

[1] https://www.ons.gov.uk/peoplepopulationandcommunity/birthsde...

by toast0 4 hours ago

Certainly. But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.

The library of congress does distribute some digitized works that are out of copyright. And it does digitize some works for archival and distribution, but having additional works digitized for (eventual) public use could be nice.

by bhelkey 3 hours ago

> But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.

It is not at all clear that this is true.

The number of books published every year is growing rapidly. According to Bowker the number of books published in the US every year has increased ~15x in the past two decades [1].

Because of this, I suspect that the median age of the books we are discussing is below 30 years.

[1] https://www.writercosmos.com/blog/how-many-books-published-p...

by merely-unlikely an hour ago

In theory the Library of Congress could "lend out" digital copies like some library systems do. This would be especially helpful for rare books since it is less likely multiple people would want the same book concurrently.

by bell-cot 7 hours ago

IANAL, but I recall copyright law being far too complex for any easy Protected/Not Protected test to exist.

by starkd 9 hours ago

So now the Library of Congress has to manage all these submissions whenever someone scans something? How do you even go about enforcing such a thing?

by Upvoter33 10 hours ago

I love this idea. At least make them turn it into some form of a public good.

by smalltorch 10 hours ago

Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.

by JKCalhoun 8 hours ago

I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)

by pessimizer 8 hours ago

If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.

by JKCalhoun 7 hours ago

Thanks.

There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again.

I'll look into libgen.

by ceasesurthinko 5 hours ago

Archive.org Scanned Book Downloader Bookmarklet

https://gist.github.com/cemerson/043d3b455317d762bb1378aeac3...

by Filligree 10 hours ago

Obviously. Copyright infringement is settled law.

by merely-unlikely an hour ago

Copyright infringement has a long and deep bank of caselaw but as Anthropic has already discovered, it is not entirely "settled."

by mbeavitt 10 hours ago

From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.

by flatline 10 hours ago

They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying

> The print original was destroyed. One replaced the other.

So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.

by jdiff 10 hours ago

Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.

by merely-unlikely an hour ago

"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

"For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.

The third fair use factor favors fair use for the purchased library copies converted from print to digital."

Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025). https://docs.justia.com/cases/federal/district-courts/califo...

by fc417fc802 9 hours ago

So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)

by jdiff 9 hours ago

Selling it is not allowed. The inability to sell it does not require its destruction.

by inigyou 9 hours ago

But copyright law does, because otherwise you have two copies.

This isn't theoretical, AI companies have finished lawsuits about this and this was the ruling.

by jdiff 8 hours ago

The ruling was that what they did was within the law, not required by law in every detail. They cannot resell the copies. I saw nothing in that ruling that required their destruction, because it is not required.

A person can digitize their own books without destroying the original. So can Anthropic. They are choosing to destroy the books for easier scanning and trying to palm off the blame for it.

by flatline 8 hours ago

Why would a company keep the hard-copy around at the risk of it being inadvertently given away, resold, etc.? It's a huge outstanding liability given that the illegal copying of works -- the other part of that case -- is what they settled out of court for some huge amount of money. Destruction is the only thing that makes sense.

I'm old enough to have been around when DCMA legislation was under discussion. Many people were dead-set against it and raised concerns over matters exactly like this. In Rainbows End (2006), Vernor Vinge wrote about a similar scenario where a robot went through the university library shredding books, and scanned the shredded pieces to recombined them into a digital archive.

Anthropic may be doing shady things and may have even done this on their own recognizance, we just don't know. As it stand, this is 100% a consequence of US copyright law, much of which was written by large corporations to protect their own assets.

by jdiff 8 hours ago

I agree fully that it makes logistical sense. But it is not a legal requirement, and they should not be permitted to use that as an excuse to wash their hands of their own decisions.

by flatline 7 hours ago

I think I agree with you in spirit. I don’t like what these companies are doing, and Anthropic’s actions can for the most part stand on their own. Copyright law is just a special interest of mine, and I do think it’s important to recognize what external incentives exist and what they prioritize. Because other companies will act in similar manners under the same incentive structure, and the problem is going to cascade and magnify if it hasn’t already. There are active court rulings setting precedence for this behavior - take note!

by toast0 7 hours ago

Probably, if you sell it after the copyright expires.

by brookst 9 hours ago

Citation please? Bartz v Anthropic seems pretty clear, see also Authors Guild v Google and RIAA v Diamond Multimedia.

by jonhohle 9 hours ago

1 point by jonhohle 0 minutes ago | edit | delete [–]

You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.

by JKCalhoun 8 hours ago

I love how sci-fi authors like Ray Bradbury toyed around with a similar issue but then got it so wildly wrong.

by silverwind 10 hours ago

More importantly: Once Anthropic is gone, all knowlege is lost.

by psma_egeliaa 10 hours ago

It will probably be actioned off in the bankruptcy proceedings.

by jonhohle 9 hours ago

That’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.

by wildzzz 8 hours ago

They cut the bindings off and feed loose leaf books through a document scanner. It would be difficult to store and probably unsellable, it's probably going in the trash.

To legally retain these scans, you must own the original book. You can sell or give away the original book (sans binding) but it's legally dubious as to whether the scan can be transferred along with it. So if Anthropic has no interest in storing thousands of loose leaf books, they are likely destroying both the original and scan as soon as possible.

At the end of the day, the only thing of value Anthropic has is the trained model which is definitely transferable.

by merely-unlikely an hour ago

It's already been distilled.

by SkyBelow 10 hours ago

>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.

by JKCalhoun 8 hours ago

Copyright law eating itself…

by convolvatron 7 hours ago

no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.

by LaGrange 8 hours ago

> I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity

Look, others talked about how this fetishising of paper books is quite silly (though I don't like it when the destructive scanning is just so one could feed it into a chatbot) but I have to say, all I can do after reading the above sentence is laughing bitterly. Anthropic is an American for-profit company, any talk about "working towards the benefit of humanity" is just marketing lies, and it's always incredible to see people treat those seriously.

_Of course_ Anthropic does that.

by JKCalhoun 8 hours ago

It's probably good to remind ourselves though of how disgusting they can be.

by raptor99 10 hours ago

I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.

They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.

A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.

by brookst 9 hours ago

Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.

by JKCalhoun 8 hours ago

It's likely more along the "fool me twice" category of wisdom. The opposite would instead be a kind of naive thinking.

by jan_m_savage 9 hours ago

They have already gotten to Archive.org. Books that were available to borrow are no longer 'available'. SMH

by JKCalhoun 8 hours ago

I've scraped all the stuff I am interested in. There's a whole r/datahoarders so I'm not alone. ;-)

by alerighi 10 hours ago

Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.

Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).

by HeWhoLurksLate 10 hours ago

we also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up

by JKCalhoun 8 hours ago

If you are suggesting that air conditioning, plumbing, democracy, and human rights are bad for society then I am missing the analogy.

by yehat 10 hours ago

Nobody will ask you, they'll be taken from you, in case you missed what happens around.

by brookst 9 hours ago

The age-old cry of the aging population, faced with tech that didn’t exist when they were young. Turn back time!

I’m reminded of the screeds about the dangers of novels.

by michaelsbradley 8 hours ago

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Are you referring to the burning of the Serapeum in AD 391 or the warehouse fires in 48 BC?

by JKCalhoun 8 hours ago

Is there a difference with regard to the metaphor?

by michaelsbradley 8 hours ago

Depends on the point being made about “historical precedent” and the lessons to be drawn from such.

Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.

by JKCalhoun 7 hours ago

My take was generally: the loss of Alexandria's collection represents a calamitous loss to our collective culture.

(But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").

by michaelsbradley 6 hours ago

It's a fair point. I honed in on "the burning of" (original comment) versus more generally thinking in terms of "the loss of", because parent context here is "AI companies destroy…".

by sensanaty 5 hours ago

The idea that these corporations or any of the literal sociopaths that work for them give the slightest bit of a shit about "benefitting humanity" is hilariously naive. The one and only thing these entities care about is money, and making as much of it as they can. If they could get away with it, they'd commit every crime that exists if it meant they get a quarter of a percentage increase in their quarterly earning reports.

by himinlomax 9 hours ago

They're not destroying rare manuscripts or incunables.

They're destroying one (1) copy of a mass-produced item for each AI company.

Public libraries destroy millions more yearly as a matter of routine.

This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.

by toephu2 5 hours ago

> This is just part of a CCP-aligned moral panic

I've heard people say this, but I haven't seen any evidence for it. Do you have any evidence?

by himinlomax an hour ago

I didn't claim it was orchestrated by them, it may or may not be, but it is evidently aligned with their interests.

Exactly the same as shutting down perfectly fine German nuclear plants was aligned with Russian interests.

by fc417fc802 9 hours ago

I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.

Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.

by SXX 9 hours ago

Yep. I was personally misinformed the same way at some point. Never expected evaporative cooling to be so popular, but it is.

by himinlomax 8 hours ago

It's possible for a datacenter to use scarce well water irresponsibly.

They don't have to, and the vast majority don't.

The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.

by 8note 6 hours ago

> a Chinese Communist Party agenda

is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want.

i dont think the CCP cares whether there's data centers in rural america.

Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers:

* sound and infrasound under x DB

* no air quality change

* must pay to build out electrical infrastructure

etc

its not to the CCPs benefit or loss to make sure the data centers are built well if they get built

by himinlomax 5 hours ago

> i dont think the CCP cares whether there's data centers in rural america.

They certainly care that the US loses the race for AI.

by TofuLover 9 hours ago

> This is just part of a CCP-aligned moral panic, along with the water use nonsense

What was nonsense about water use?

by DaSHacka 9 hours ago

That AI consumes it at an abnormally high rate, presumably.

The claim never made sense to me either, I can only assume those that regurgitated such claims never worked with HPC or even general datacenters before.

Was recently talking to a (non-technical) friend about this, she was surprised after talking about the "insane water use for AI datacenters" when I responded that open-loop cooling is pretty rare for a datacenter and I've never actually seen it used before, versus closed-loop (or just regular air-based cooling) which has no real noticable water consumption.

by brookst 9 hours ago

It was never that dramatic, and it’s declining day by day. It’s a panic over a real but small problem.

Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.

by SXX 9 hours ago

Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.

by himinlomax 8 hours ago

1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere.

2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.

by diseasedyak 9 hours ago

That AI data centers are drinking up local ground water for cooling. It isn't (or wasn't) nonsense, though. It was/is a real thing, though it seems to be on the out in favor of closed loop cooling after the massive and still on-going public outcry.

by alex43578 9 hours ago

The outcry over data centers using a fraction of the water used for things like golf courses or growing alfalfa in a desert.

by pfdietz 8 hours ago

Nuclear killed itself (vast cost overruns); there was no need for hallucinated foreign influences.

But I understand blaming your energy waifu for its own failure is unacceptable for nuclear bros.

by 8note 6 hours ago

nuclear was killed by russian nat gas money.

running nuclear plants were shut down while running just fine

by pfdietz 6 hours ago

Nuclear in the US was killed by Russian natural gas money?

What other nonsense do you believe?

by himinlomax an hour ago

It demonstrably was in Germany.

Look up Greenpeace Energy, and what cushy corporate job Schroder got after leaving office.

by chrisjj 9 hours ago

> Despite the copyright restrictions that are forcing companies to do this

There are none.

by shevy-java 9 hours ago

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.

Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.

by thesdev 9 hours ago

> working towards the benefit of humanity is not an exclusive right / domain of theirs

That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.

by bmelton 9 hours ago

If they were just chopping them up without scanning them first, then sure, but I think that scanning books and burning books are polar opposites

by logseman 9 hours ago

Burning books and destroying them in a way that nobody else can access the content anymore is a distinction without a difference.

by HedonicEscal8r 19 hours ago

The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.

I support Anna's Archive, by the way. Information wants to be free.

by Cider9986 19 hours ago

You can donate with over 20 different payment methods after making an anonymous account.

https://annas-archive.gl/donate

by Levitz 19 hours ago

It's an excellent play by them, using moral outrage to the benefit of the project. When life gives you lemons...

by everyday7732 12 hours ago

Anna's archive is selling data to AI companies. They're essentially saying "hey, don't sell your books to be scanned by AI companies, scan them yourself, and give us the data, so we can sell it to AI companies."

by urbnspacecowboy 10 hours ago

With the important difference that scans sent to Anna's Archive are, you know, archived.

by srj 9 hours ago

Why are the companies legally required to shred the books? That's the most surprising part about this to me. Surely if they bought them second hand they could donate or resell after scanning. I'm wondering if the scanning machines are damaging the books.

by famouswaffles 9 hours ago

You're not allowed to make copies of copyrighted work. Of course this law was made long before the reality of digital books.

The argument for the digital era is that you can make a digital scan and it's not a 'copy' so long as you shred the original and don't attempt to distribute it.

by tgsovlerkhgsel 7 hours ago

The legal idea behind it is that if you "copy" the book it's bad because now there are two copies and you "stole" from the author, but if you "move" the book to a digital form (and don't copy that outside of your organization), it's fine.

Given that they have to destroy them anyway, they're obviously also going to use the much cheaper destructive scanning (cutting off the spine and using a feed scanner rather than carefully turning page by page).

by tene80i 16 hours ago

Plenty of written works aren’t “information” but rather art. Most piracy is just about people preferring not to pay for novels, TV and film.

by mrweasel 15 hours ago

There is also a ton of tv shows, movies, music and books that you cannot buy, for now real good reason. I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions. With things increasingly only being available on streaming platforms or behind DRM in other ways, we risk looking back on the current era as a black hole 50 years from now.

My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.

by merely-unlikely 7 minutes ago

> I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions.

I can already think of a couple examples I've run into in the audiobook world. I have copies of Douglas Adams himself narrating his Hitchhiker's Guide to the Galaxy and subsequent books. As far as I know, these recordings are not available for purchase anywhere. I got them because someone was kind enough to upload them. I would have preferred to buy them but didn't have that choice.

Some of Iain Bank's audiobooks were only available in Europe for a time (and maybe still are). For those I was able to convince Amazon I was buying in Germany at least.

by tene80i 15 hours ago

I agree there is an archival justification. I just don’t think that’s why most people who pirate things are doing it.

by guax 16 hours ago

Or not being able to. Regional licensing, missing and shuffling content.

To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked.

I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.

by tene80i 15 hours ago

There was no way in your country to pay and watch it? FIFA will have sold the tv rights there to someone, surely. In which case your complaint is what, that it was expensive?

by guax 13 hours ago

Free on both, but not with the narration in my native language.

On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).

And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.

by Xunjin 14 hours ago

We humans, often forget that logistics is always the bottleneck in any Industry.

by TFNA 12 hours ago

Indeed, art. And it is a pretty common position that all people should have access to art and culture. Add up the cost of buying the DVD/Blu-Ray releases for the 1500 or so films that make up the canon of cinema. That's a sum of money daunting even for people in developed countries, let alone most of the world. Piracy is going to be the realistic solution. (And before you say "Use the library", you know well-stocked libraries don't exist in most of the world, right?)

by tene80i 8 hours ago

If you're saying "piracy is justified in those parts of the world where access to any art or culture is prohibitively expensive", then that's a pretty defensible position, not unlike "It's ok to steal bread if you are starving".

But that's not what many people defending piracy are doing. Many of them do indeed have great libraries nearby, and art and culture on demand at prices they happily spend having food delivered to their house.

by TFNA 6 hours ago

I don't know if you have noticed, but piracy has declined greatly since the introduction of streaming. The scene is a shadow of its former self. The communities today that are keeping high-quality releases of canonical music and films available are overwhelmingly based in regions of the world with a dearth of good libraries.

by FeloniousHam 8 hours ago

> And it is a pretty common position that all people should have access to art and culture.

Access to _all_ art and "culture"? For free?

by TFNA 6 hours ago

Even if not all art, once something becomes canonical, it then becomes something that people should be able to easily familiarize themselves with for the sake of an educated and edified citizenry. And indeed, many countries subsidize public libraries and live performances for this very reason. But no state can manage to provide free or nearly-free access to the entire canon, so piracy helps fill the gap.

by FeloniousHam 3 hours ago

Do you have any examples of art or culture that is unavailable?

by dombiscoff 15 hours ago

Art is information, always.

by tene80i 15 hours ago

It’s not tautological. Explain why, and particularly why it’s only information, which is the thing that purportedly wishes to be free.

by odyssey7 7 hours ago

Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.

Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.

Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.

Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.

Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.

Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?

by merely-unlikely an hour ago

Anthropic wasn't explicitly told to destroy the physical copies, but it weighed heavily in their favor.

"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

"For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.

The third fair use factor favors fair use for the purchased library copies converted from print to digital."

Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025). https://docs.justia.com/cases/federal/district-courts/califo...

by RaffaelCH 7 hours ago

From what I understand, to work with copyrighted books they need to essentially format shift (i.e., scan and destroy the physical book). So a book vault would not solve this issue.

A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's not worth it.

At this point my only hope is that in the long run these scans make it to the public somehow (leaks, copyright changes/expiration, whatever), where they can then be accessed and preserved by everybody. Then we could have our true digital library of Alexandria.

by odyssey7 2 hours ago

If they’re old enough and rare then copyright is moot, as you mention. The size of the vault is of interest but it’s not something that makes the vault possible vs impossible.

I’m not a lawyer, but I would guess the theory behind the destruction is to guarantee that they are authorized under fair use to have and to use their digital copy. Imagine a situation in which someone else becomes the owner or is using the original. If you destroy it, that becomes impossible. If that’s part of the constraints, then a vault, rather than a library, could be a key distinction.

You might be confusing the need for an original work that draws upon another to be “transformative,” in order to not violate copyright, with the notion of “fair use,” which allows scanning something you own for your own personal use.

by lukeschlather 5 hours ago

They could put everything in some kind of nonprofit book vault/archive that includes all the source pages as long as they didn't use or distribute it. They could even provide a mechanism for rights holders to recover the text for free, if they need access.

by cube00 5 hours ago

> establishing goodwill with the public

Not sure even rare books will dig these big AI companies out of the hole they're digging for themselves.

by akk0 9 hours ago

I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.

That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.

by OtherShrezzing 7 hours ago

>I imagine they are only buying one copy of each book

The previously struggling second hand bookseller in my town has upgraded their car from a 15 year old hatchback Renault to a brand new Range Rover. Some Canadian company has been buying any book he can provide them for the last year. Their quotes aren't by number of books, or even weight, but by volume. As in, they pay him by the shipping container, and he sends several of those a month.

I think it's reasonable to assume the books in this supply chain which aren't destroyed in digitisation are just pulped and sold to Procter & Gamble for toilet paper manufacturing. I can't see any other fate for Anthropic's second and third copies of The twelfth edition of Vera Lynn's 1980's memoir "We'll Meet Again".

by dbspin 9 hours ago

Doubtful. A robust protocol would be to scan several of each edition (to ensure no scanning errors), and scan each edition. Then too, these books are being purchased in lots with accidental duplicates, and all the major labs are doing it. So we're likely talking about tens of each book. For rare books - anything over a couple of hundred years old or small print runs either, that could well be most or even all copies. This wouldn't be immediately noticed either, especially if the books aren't currently considered noteworthy or well known.

by JKCalhoun 8 hours ago

"I imagine they are only buying one copy of each book…"

I doubt that corporations of this scale do that extra kind of… book-keeping. They more than likely buy books by the pound.

by brookst 9 hours ago

The dissonance when a cause you support is loudly represented by disingenuous types.

At some point they’ll hit on data center water consumption as yet another reason to support the cause.

by ironqcold 7 hours ago

The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.

But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.

by zahlman 5 hours ago

> Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.

"Paint a picture" is apt.

"... like AI companies are some movie villains burning paintings so no one can see them. but in reality they just cut the canvas off the frames because it's cheaper and faster to photocopy. Most of these paintings have been collecting dust in galleries for decades and nobody needs them."

by iLemming 5 hours ago

> You can't to ask a neural net to give you a exact copy of a page from that book

That's the problem. I tried the other day to find the verbatim quote from one of the Gabriel García Márquez's books - a single effing sentence and I nearly lost my mind - every LLM would be like "this is a copyrighted material and I can't share it verbatim". WTF? These books are widely known, published in many languages, and yet no amount of money spent on tokens can give me the exact fucking sentence, just like the author intended? Really? I need to find, buy, download and search through the actual book to quote something that Márquez wrote at some point? What is even worse is when the LLM just outright lies to you, giving you the quote, but slightly rephrased.

by skeledrew 16 hours ago

Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.

I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.

by guax 16 hours ago

I think is the scanning for profit and destroying them in the process that angries people.

If they we’re just kept where they were you can always say they’ll eventually be scanned or have that potential.

I do believe the issue is a bit overblown but the core of it sounds reasonable to me.

by tgsovlerkhgsel 7 hours ago

It's destroying books that angers people in general.

Never mind that it's the 200th copy of some 1950's romance novel that nobody will ever care about, every time a library prunes its collection there is outrage.

by skeledrew 16 hours ago

Shouldn't they be allowed to do whatever they want with the copy they bought, and so legally own? Isn't that the entire purpose behind "copy right"?

by impossiblefork 15 hours ago

Of course they shouldn't.

If the books are rare enough there's a shared cultural value that is being destroyed.

by skeledrew 14 hours ago

Why should they buy the books then if they can't do what they want with them? Just leave them wherever they are to rot and eventually be unceremoniously dumped anyway, without being preserved in any form. Some things are just unavoidable.

by impossiblefork 11 hours ago

In order to read them, presumably.

Just because they are for sale doesn't mean that no one else would have bought them. Destroying them obviously destroys them, which obviously destroys culture. After all, if the books is destroyed, it will not be read.

Think of it like this. If I buy an island, and there are bunch of people who live on it, maybe they're day laborers or whatever, is it right for me to expel them, if I don't want them? Can't that be straight up genocide, if they have some culture indigenous to the island?

The same is true for books. If you destroy culture, you destroy culture. There is no magic which triggered "but I owned the physical book" which means you didn't.

It doesn't matter what property relationship you have to a thing: whatever property relationship you may have to it, you still do what you do, i.e. you take whatever action with respect to it as you take with respect to it.

Another good example is food. If there's a shortage of it, and you still have a great deal, you may think "eh, what does it matter if I accidentally burn some, I save time through my carelessness" but if there are others who aren't getting any, you may be killing people, and may be despised for how you use "your property". The same is true here. Why should one not despise one who deprives others of rare texts, that may even be lost, and thus cause cultural destruction of some culture that may actually be rare, by treating the carelessly. That rare book on model building or whittling from the 70s may actually be important.

by skeledrew 10 hours ago

You can't compare an island or food to books though, because the former are only really valuable in their atomic form. Like if you take detailed pictures, people can't live on or eat those pictures.

You scan books though and all the value they provide is still fully available in bit form, because books merely contain information. And of course being in bit form means they can be trivially copied, technically. But there's also the legality, which determines how many copies it's legally permissible to retain. The latter becomes the bottleneck for your preservation of "culture" because said culture could be made trivially available to anyone with an internet-connected device, if only rights holders weren't screaming foul (remember what happened with Internet Archive during/after Covid?[0]).

Ultimately there's no destruction involved when books are scanned, just a change of format which helps to ease management and for legal compliance.

[0] https://www.libraryjournal.com/story/internet-archive-loses-...

by impossiblefork 9 hours ago

It's culture. These scans are of course not distributed, so the value doesn't quite continue to exist.

You say "culture", but is is culture, and culture that was interesting enough to write down. Writing a book in the 70s isn't a trivial thing, there are gatekeepers, editors, publishers etc., even for your little book about local history, and the information is valuable to the community it's about. It affects their internal cultural transmission.

by skeledrew 8 hours ago

The physical work before was never distributed either (which is what makes them economically valuable), so the the point is mostly moot. I say mostly because at least when these neglected works become part of a training corpus, the knowledge they contain can be surfaced on demand, or even by accident. Think of it like the grandpa telling stories of his experiences to the grandkids, to the best of his recollection (and imagination), going on related tangents as they occur to him or are triggered by the grandkids' questions, etc.

by tgsovlerkhgsel 7 hours ago

> Destroying them obviously destroys them

You're ignoring that they're also scanning them. Which means that we're one copyright law change + some minor incentives away from making these archives actually useful, vs. the book silently getting destroyed in the bookshop's dumpster because nobody bought it.

by impossiblefork 5 hours ago

They become very non-public datasets though, and it does not preserve the culture the books represent.

by guax 16 hours ago

They are, they are in no risk of being arrested. But that does not make it free of moral judgement.

by skeledrew 16 hours ago

Why not?

by guax 15 hours ago

That's just a fact, not something to argue around. People will judge you for many reasons, some cultural, some political, some ethical, some personal, some valid, some not. Its human nature.

Which is also not illegal and within some bounds and exceptions, a protected right across the globe.

by CamelCaseName 10 hours ago

You ask "Why destroy physical books?"

I ask "Why save physical books?"

If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.

I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.

by zero-sharp 9 hours ago

>If they are truly rare, then they are likely not valuable,

Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.

by tele_ski 10 hours ago

So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory

by gruez 10 hours ago

Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.

by wasmitnetzen 10 hours ago

You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).

by shiandow 10 hours ago

Good point, why waste time deciphering old badly burnt scrolls when anything worthwhile should have been preserved.

by noosphr 9 hours ago

If the Romans had the printing press we'd have a lot more of those scrolls.

by mzhaase 8 hours ago

What you personally find important is not what everyone else finds important or inspiring. Destroying something takes it away from every single future human being.

by ainiriand 10 hours ago

That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.

by Goronmon 9 hours ago

Libraries tend to regularly destroy books as well. And they aren't scanning them first either.

Doesn't that make them even worse?

by cormorant 9 hours ago

"State libraries" such as the Library of Congress. (Which is not regularly destroying books, AFAIK.)

by tsukurimashou 10 hours ago

> If they are truly rare, then they are likely not valuable

"likely" being the keyword here, what about heavily censored books?

by zahlman 5 hours ago

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That is literally the opposite of how it works.

https://en.wikipedia.org/wiki/Scarcity

by karmickoala 7 hours ago

They may have been made scarce by many methods, perhaps a low print run; court rulings; burning in a revolution to suppress dissenting views; scanning to avoid competitors to get the content and prevent other LLMs to reference, search, or train on it. Perhaps it's a particular edition that is rare (e.g., the first edition had a different description and was thoroughly changed in the second edition of the book, which may be important to have all the facts from someone's biography).

by madibo3156 8 hours ago

Nobody's said it in this thread so I'll drop it here where you ask "Why save physical books?"—

The problem isn't with morals or copyright. What we're up in arms about is case law. Past rulings have implied that destruction of books significantly contributes to the process being "transformative". This encourages companies to destroy the books. I think this is really dumb.

Why save physical books? It's because the reason to destroy them isn't good. If you think there's too many bad books out there, that's a different argument. Maybe your fight is against consumerism, I don't know.

by TaLiTr 9 hours ago

Massively missing the point. Having paper books isn't the point. Preserving copies of books is the point, so history isn't lost. Often those books only exist as paper copies due to their age. I don't really care what happens to the paper copies, only that their content is preserved in a way that is accessible. Anthropic's private servers aren't it.

by PartiallyTyped 10 hours ago

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.

Same can be said about many books.

by hahn-kev 8 hours ago

Right but is an AI company going to destroy a Gutenberg bible? No, as you said the content is not rare, and that's what they want.

by gruez 10 hours ago

So old books are valuable because old paper is valuable? I understand why the Gutenberg bible might be valuable but do we really need thousands of mass paperback novels?

by jll29 9 hours ago

Ask: valuable to whom and why?

- Reader: narrative

- Collector: scarcity of the physical artifact

- AI Company: language samples (quantity, variety), facts

by gravypod 10 hours ago

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!

by vasco 10 hours ago

I agree with you for the same reason I think McDonald's is the best restaurant in the world!

by pmoriarty 10 hours ago

Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.

After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.

by azatom 10 hours ago

I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.

by sieve 17 hours ago

Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.

So the solution is:

- mandatory copyright registration and renewal with links to where the work can be acquired

- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.

In this way, a hundred digital preservation societies can bloom.

by skeledrew 16 hours ago

> copyright

This is what caused the problem in the first place. If people had unrestricted access to content then the world would be a better place. And works wouldn't be so rare that it's worth AI companies buying and destroying them to gain some edge, as well as remain in legal compliance.

by sieve 16 hours ago

I completely agree. IPR as a concept is suspect. But it is the world we live in, with timelines extending like crazy. There are works produced before I was born that will remain under protection till long after I am dead.

A targeted modification to the laws could produce most of the benefits for a minor cost.

by theshrike79 13 hours ago

Copyright should be globally determined so that if I cant pay fair market value for a piece of content, it's deemed out of copyright protection and I can use whatever means to get it digitally.

Like if a game isn't available for sale anywhere in my region, I can get it without breaking any laws. A book is out of print and I can't pay money for it digitally -> free game.

(Why "fair market value?": So that skeezy publishers don't have an online shop with one physical copy of every book they own for $1Trillion just to fulfill the law)

by skeledrew 10 hours ago

Copyright should be banished out of existence so information, which is a trivially reproduced intangible, can be freely copied and distributed anywhere it can be useful. Instead of having the crap like what happened with the Internet Archive during/after Covid.

https://www.libraryjournal.com/story/internet-archive-loses-...

by Rover222 7 minutes ago

Credit to SpaceXai (whatever it's called now) for going through the added effort of not destroying these books

by glimshe 19 hours ago

The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."

Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?

by eru 19 hours ago

> Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time?

No. They also use lots of other methods to get training data.

by npn 17 hours ago

No but with 100% clean data you can easily train a model to filter ai generated content.

by ACCount37 13 hours ago

As a rule: all high quality text is useful.

There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.

At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.

Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.

by TiredOfLife 10 hours ago

Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content

by fastball 5 hours ago

So we don't want companies to buy books and scan them and do whatever they want afterwards, but we also don't want to allow piracy of digital copies (the $1.5B Anthropic settlement)? Bit of a rock and a hard place for them.

by ronsor 5 hours ago

They want AI companies to shut down.

Unfortunately, "we simply want you to not exist" is not realistically actionable or even compatible with a free society.

by branon 9 hours ago

I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list

Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded

by afpx 9 hours ago

archive.org is an incredible blessing that I never really respected enough until the last few years. I have bookmarks going back to the 90s, and for some reason, starting in 2015 or so, sites just started disappearing. I estimate at least 20% of my bookmarks are 404 now.

by cormorant 9 hours ago

archive.org itself reminds me of the Library of Alexandria. It is irreplaceable, and this is a disaster waiting to happen.

by theartfuldodger 8 hours ago

I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.

I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.

by xvxvx 19 hours ago

Pretty funny that they just took Anna’s archive and ingested it.

As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.

by ryandvm 10 hours ago

It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.

by JsonDemWitOster 10 hours ago

The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.

Which is really why the outrage cycle over Anthropic's actions is largely misplaced.

by jupp0r 18 hours ago

I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!

by joshuakelleyds 7 hours ago

I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:

> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order

Also the summary:

> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.

https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

by shakna 19 hours ago

I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.

Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.

So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.

by altcognito 18 hours ago

"You own a particular physical copy; you don't possess an abstract transferable 'one-copy license."

There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.

Edit; despite the above, looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form we purchased, therefore it is legal.”

I still dont think there is a requirement to destroy the book, but since there isn’t a reason to store the book and they can’t sell it, they probably just took the cheapest route. It might be worth an argument that they only purchased the right to use the digital copies while the physical copies exist, but I’m in over my head from a copyright standpint

by hnfong 6 hours ago

Maybe the law doesn't explicitly say that. But it's easier for lawyers for AI companies to argue the case if they destroyed it. "Look, there's only one copy! We didn't make additional copies!"

It's also probably more convenient for them to destroy the books as opposed to trying to find space to store them. Knowing those companies, most likely they'd be just stuffed into some warehouse to rot after a couple years.

by breezybottom 19 hours ago

Since when do AI companies care about the law? Most of their training data is pirated.

by shakna 18 hours ago

The headlines about destruction came not soon after they got rapped on the knuckles and told "no more pirating".

by tkel 14 hours ago

No, the destructive book scanning was happening before that.

by shakna an hour ago

Right. Because that's the easiest way to comply with the law. I'm not arguing that.

I am, arguing that they're using it to make headlines, and stir up anger, as a wedge to trigger changes to copyright. Easier to make more than one change, once the door is opened up.

by blooalien 18 hours ago

> Since when do AI companies care about the law? Most of their training data is pirated.

I imagine since the law recently cost one of them truckloads of money for their violations of it?

by shrubble 10 hours ago

AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.

I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...

by dqv 10 hours ago

The other thing they could be doing is using some of that lobbying money to try to reform copyright law to allow them to release the scans that are still covered.

It would likewise earn goodwill from a lot of people.

by shagie 9 hours ago

Copyright law is an implementation of international treaties.

Berne Convention https://en.wikipedia.org/wiki/Berne_Convention (182 parties)

TRIPS Agreement https://en.wikipedia.org/wiki/TRIPS_Agreement (164 parties - part of WTO)

The United States can't make copyright weaker than what those agreements require without pulling out of the WTO.

The core of copyright law is about who has the right to redistribute a work. If I buy a print of a photograph, scan it and use that as my desktop image... I can do that. I cannot redistribute the scanned image, and if I was to sell the print later I should delete the scanned image.

Note that format shifting is covered under fair use... which is what training is taking place under. However, that doesn't mean that they can release that format shifted content... nor can then re-release the original work if they are retaining the format shifted content.

https://library.georgetown.edu/copyright/fair-use-reformatti...

    Under § 106 of the Copyright Law of the United States, the owner of the copyright in a work has the exclusive right to make copies of that work, unless an exception applies. When considering reformatting media, please note that individuals do not have an automatic right to reformat a work from one format to another. In order to legally convert media, your use must fall into one of the following categories:

    you own the copyright in the work,
    you have permission from the owner of the copyright, or
    you have done a fair use analysis and have determined that fair use applies
by phoronixrly 10 hours ago

They could be, however that would weaken their position as the sole source of knowledge which seems to me is half the purpose of their book destroying initiative.

by dqv 9 hours ago

That's why the "but it's illegal" propaganda is so unconvincing for me. It's really "but it's illegal (and we want to keep it that way)".

by archonis 9 hours ago

PR only cares that the company is percieved as unstoppable & inevitable.

If their PR teams did have a response to these actions, it would be something to the effect of:

"would you rather china destroy all the books and gatekeep the knowledge?"

by rcarr 8 hours ago

Does adding an old book risk making the model worse? Off the top of my head:

- Reinforcing outdated, disproved or otherwise incorrect information.

- Reinforcing outdated forms of communication e.g purple prose.

You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory

by ZoomZoomZoom 10 hours ago

The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?

by Aurornis 9 hours ago

> The main question is why aren't they leaking it to AA themselves?

How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.

They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.

by notpushkin 8 hours ago

> You cannot scan a book and share it without violating copyright law.

Hence the “leaking” part.

by embedding-shape 10 hours ago

> The main question is why aren't they leaking it to AA themselves?

Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.

by Filligree 10 hours ago

Because that’s illegal.

by onionisafruit 10 hours ago

Exactly. They set up this operation specifically to comply with the letter of copyright law and defend against publisher law suits. “Leaking” to Anna’s Archive is the last thing they’re going to do.

by ZoomZoomZoom 10 hours ago

The illegal part is (or ideally should be) using the books in their training. The morally right action is to make it public afterwards.

by thisisauserid 10 hours ago

Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".

by chistev 8 hours ago

What is the deal with AI Companies buying old books to scan and then destroy them?

https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...

by tgsovlerkhgsel 7 hours ago

Many countries require a copy of each book published there to be submitted to their national archives/library. The US has had this requirement since 1790 as far as I can tell.

All countries I checked (Germany, France, Canada) seem to have similar requirements. WIPO claims that this is the case in the majority of countries (https://www.wipo.int/documents/d/copyright/docs-en-registrat...).

Thus, most books should not be at risk of getting permanently lost due to these practices.

This is Anna's Archive (ab)using a current controversy (which I think is based around emotional appeals and distorted facts to create outrage over a non-issue) to tongue-in-cheek advertise their open pirate library.

by abunner 5 hours ago

AI companies are supporting the market for books that no one else wanted. The economically illiterate assumption of the author is that "rare" books are good. If they were that good, they would be priced higher.

https://abunner.substack.com/i/210907372/anthropic-is-suppor...

by bjornnn 5 hours ago

destructive book scanning is pretty standard, it's the only way to efficiently get a good flat scan of the pages without a tremendous amount of human labor to correct distortion on every individual page. i'm curious about what specific rare books are being destroyed and how rare they really are. and at any rate, anyone can go out and buy books and scan them or destroy them or set them on fire or whatever, they bought the book it's their own property. so now physical books are treated as the public commons while intellectual property is privately owned? are we in topsy-turvy world? i feel like there are much more egregious crimes that people ought to be talking about with ai companies, like impoverished black people having their city water supplies poisoned or the massive financial fraud that will destroy the economy or the massive co2 emissions that will destroy the world or just the fact that it's not even artificial intelligence at all and it's not even a technological innovation, it's just a google hack that dumbed down an existing neural network model enough for it to run on lots of nvidia gpus and produce an impressive-enough tech demo to show to gullible investors who don't know what to do with their massive piles of money.

by 1970-01-01 7 hours ago

This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.

by tescreal 7 hours ago

A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.

first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.

by thuruv 18 hours ago

I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!

by Cider9986 16 hours ago

They likely would not destroy the original books after scanning but apparently it's the legal way to do things because of copyright.

by Cider9986 19 hours ago

The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.

Unrelated: So with this one copy BS are you not allowed to have backups of the data?

by 0x0000F8_ 18 hours ago

I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models.

No data => No models => No competition.

by QuantumNomad_ 18 hours ago

IA was in hot water already even before ChatGPT came out.

> ChatGPT […] originally released on November 30, 2022

https://en.wikipedia.org/wiki/ChatGPT

> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.

https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive

by visarga 18 hours ago

Yeah, great logic, that way they are sure there are no extra copies around.

by yipinwong 7 hours ago

The post reads like a propaganda. Evoking feeling over metrics or results

(<- that's what propaganda does by definition)

e.g.

> but ethically, it’s an extremely serious crime against humanity.

Who are you that you consider it a serious crime against humanity? What are you a saint?

> Knowledge is permanently monopolized on private servers.

Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.

With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.

The logic of the writing is not correct.

by alightsoul 7 hours ago

Isn't this what Dario does?

by yipinwong 5 hours ago

If Dario appeals to feeling, yes.

Provide more context if you want a better feedback not whataboutism.

by alightsoul 4 hours ago

Dario keeps rambling about how China is bad, like the US doesn't do those things itself. That appeals to American exceptionalism which every American is taught in school, it's just not taught explicitly because it's ingrained in the very lifestyle that Americans have

by clarionbell 8 hours ago

In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.

by pino83 7 hours ago

What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?

And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?

Why is that?

If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?

by twright 8 hours ago

I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.

by throwaw12 8 hours ago

America is interesting.

* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz

* download and publish a book as a company -> fine 1% of revenue

* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations

* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model

by silcoon 18 hours ago

As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.

On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.

On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.

This is scary, and it reminds me of Fahrenheit 451.

Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.

We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.

In a world like that, what motivation would we still have to read?

by Cider9986 18 hours ago

>In a world like that, what motivation would we still have to read?

Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.

by msftgreed 18 hours ago

The human tradition is storytelling. The idea that storytelling was something a company could own and other's weren't allowed to tell is very, very new in our history.

People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.

by silcoon 17 hours ago

> Why would a reduction in human writers cause a complete reduction in motivation to read?

Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).

> People write for hundreds of reasons other than to make money.

Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.

Today we are lucky because education is accessible and printing is cheap.

> they created literature before copyright was a thing.

Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...

by womble2 13 hours ago

Notably, Shakespeare lived in a time before copyright and famously ripped off plots from his contempories. I dont think he's a great reference for the argument that there would be no writing about the modern world if you removed copyright.

by watwut 15 hours ago

First, people want tonread books that speak to them and their lives, older books are not that.

Second, people write to be read. It takes huge amount of effort too. With no potential reward for it at all, they stop.

Third, we are social animals. If you dont see people reading, if you dont read yourself, you wont even think of writing.

by internet_points 15 hours ago

> physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write

there have been many times I've wanted to pay for an ebook but found that the only places that sell it apply DRM to it, so shadow library it is

by protocolture 17 hours ago

>On one side, there’s copyright law, which exists to support the work of creative people.

Which exists to enrich Disney and other large corps, while they hide behind artists as a human shield.

>Artists need some form of reward.

Right, as do artists who use the work of other artists as their starting point. Copyright holders aren't bill and bob artist, they are massive corporate trolls throwing around the weight of almost 100 years of our cultural heritage, sucking the marrow from its bones.

>On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.

The books are getting digitised into a permanent record of all our cultural heritage. It just sucks we don't have control over it. If only there was a way we could get them digitised AND control our cultural heritage. HMMMMMMM.

>Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write.

The incentive to write is being killed by slop groups like 20Booksto50K and Kindle which predate AI by at least a decade. AI just lets them work faster.

>We might end up with all of humanity’s books digitized and accessible for free

Excellent

>But there would be no human writers left.

Unlikely, but there would definitely be no Disneys or Conde Nasts left, which is a massively pro social outcome.

>In a world like that, what motivation would we still have to read?

In a world with all books digitised and accessible to read? A huge huge huge incentive. I already partake if books are too expensive where I am. It would take me the rest of my life to read all the books I already want to read. What kind of inane dribble is the idea that copyright makes it interesting to read? I havent even read all of Howard and he's in the public domain (in cool countries at least)

by juiceland 9 hours ago

Why are people acting like books can’t be reprinted?

by letmetweakit 8 hours ago

You have to have a sample to be able to reprint it.

by juiceland 8 hours ago

Good thing the books are being scanned, then. The copyright owners know just who to go to get the sample.

by adamors 7 hours ago

They’re not giving anybody any samples lol

by juiceland 7 hours ago

Has anyone asked? From the discussion here it seems like the books are extremely valuable.

by jmspring 7 hours ago

I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.

by lukasbm 8 hours ago

A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"

by godber 8 hours ago

I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.

Clearly that makes assumptions about the function of government and the complacency of copyright holders.

by spogbiper 7 hours ago

https://www.copyright.gov/help/faq/mandatory_deposit.html

some form of mandatory deposit law has existed since 1790 in the US. other countries have similar laws. there is a very good chance the government already has a copy of anything that has been copyrighted

by thuuuomas 9 hours ago

Is there any evidence the books are truly "destroyed" & not merely "disassembled"?

It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.

by 1970-01-01 9 hours ago

Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.

by danparsonson 9 hours ago

Value is entirely in the eye of the beholder - money itself only has value because enough of us agree that it does.

by pessimizer 8 hours ago

No, it has value because it is the only means that are absolutely accepted to pay taxes and most legal judgements (i.e. obligations to the state that issued that currency.) If you don't have dollars and you need to pay US taxes, you have got to get some or you will go to jail.

It's not bitcoin.

by dbgrman 5 hours ago

Hmm... interesting. previously this was annas-archive.pk, now its on gl domain. What happened?

by mmaunder 7 hours ago

At the risk of angering the zeitgeist, destroying one or two or ten paper copies doesn’t make AI companies “become the only ones in the world with digital copies.”.

by jbstack 7 hours ago

Not if it's the Da Vinci Code. But there are reports of bookshops getting large orders that include rare books.

by pfdietz 7 hours ago

Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...

by finn888 14 hours ago

The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.

by jeroenhd 13 hours ago

Lossily stuffing books into a model through a training process has so far been deemed legal. Enabling piracy by giving away digital copies is not. The internet archive tried to give away digital copies and they got sued to hell and back (which they should've seen coming from miles away).

I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with them (only once, of course).

by Joel_Mckay 12 hours ago

>Lossily stuffing books into a model through a training process has so far been deemed legal

Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture.

The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take over 70 years in some places to become public domain.

The core issue is "AI" firms have so much borrowed cash around, that getting a $1.5B fine for being a pirate is taken as a cost of doing business. The law is simply not equipped to handle this type of hyper-scaling criminal act. =3

https://www.bbc.co.uk/news/articles/c5y4jpg922qo

by jeroenhd 10 hours ago

The piracy lawsuits have so far only deemed that obtaining books through pirating is a violation of copyright. Training the models on them and reselling model access has yet to be ruled illegal, from my understanding, even though there has been plenty of opportunity to.

Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.

by Joel_Mckay 10 hours ago

>I don't think the lawsuit would've gone anywhere

That is not how copyright/trademark/contract laws treat similar works. Most LLM know about Disney Mickey Mouse, and LLM vector search space proximity results will gravitate more accurate reproductions of protected works regardless of granularity of data.

OpenAI simply canceled a popular service to avoid Disney wrath.

https://www.theglobeandmail.com/world/article-openai-sora-di...

Also, trying to escape directly ripping off notable famous people with nonunion talent:

https://www.youtube.com/watch?v=YhgYMH6n004

I would say the "AI" firms will keep buying time with all that borrowed cash. Everything that could be scraped has already been stolen, and thus the problem should begin to self-correct. The Shrek movie release market correction history correlation is funny, and a new film is due 2027 in July. =3

by ErigmolCt 16 hours ago

A genuinely useful project would identify publications that are actually rare, poorly catalogued, or held by only a few libraries, then prioritize those for careful preservation

by tptacek 19 hours ago

These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.

It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.

The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.

But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?

by Levitz 19 hours ago

The outraged people don't care. They hate AI, and so anything surrounding AI that can be evil is evil. Books are good, AI destroys books, AI is bad.

Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.

by tptacek 19 hours ago

I think people genuinely don't get that book destruction is like a pretty natural part of the book lifecycle.

by Levitz 7 hours ago

It doesn't matter if they get it, because they wouldn't care either way. They would just pivot to "but more are destroyed this way" or "but these have value or they wouldn't buy them" or some variant of it.

There is no logical train of thought. It's not about logic.

by eru 19 hours ago

And the solution to the water issues can be found in any introductory textbook on the subject: a water price.

> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.

Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.

by lanyard-textile 15 hours ago

> Why do they need to assimilate whole species?

So then they're gone :)

When the borg is threatened ("threatened") by a race, they choose to effectively end their way of life one way or the other. That is done by assimilation or by death.

If only a single outstanding member of a species becomes a threat, that's a strong signal for their overall potential; therefore everything needs to go, the sooner the better.

by eru 13 hours ago

I thought the Borg wanted to get better and better? Leaving the species essentially intact, and absorbing a few members every so often gives you much more material to work with than a one time infusion: they are still evolving and changing.

by ivell 17 hours ago

Knowledge of every member of a species is better than few samples? LLM became better due to its huge dataset instead of just few samples.

by eru 13 hours ago

You hit diminishing returns pretty quickly. And LLM training only bothers with a single copy of each book; they don't even need to destroy the documents. (Eg they didn't destroy the American declaration of independence to train on it.)

by guax 16 hours ago

Its nuanced and complicated but its not fully without reason. I think there would be a compromise of using those books while playing the "we're helping preserve them part" but I don't think that even crosses the mind of most Ai CEOs in an honest way.

The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.

by imperfect_light 17 hours ago

I don't know how it works today, but 20 years ago bookstores wouldn't return unsold books (too expensive to ship) but would simply tear the covers off and throw them in the garbage.

by fenomas 18 hours ago

More and more I feel like anti-AI is a bigger bubble than AI. It seems like every week it expands into a new dimension - anti-Flock protesters tearing down years-old traffic cameras that were used for research into auto accidents, etc.

Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..

by customguy 17 hours ago

> my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..

There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.

As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.

by globular-toast 17 hours ago

They pulp books that have many copies surplus to requirement, not the last few copies in second hand book shops.

What if people like food more than AI? Have you considered that?

by jsphweid 18 hours ago

To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?

by Ekaros 18 hours ago

They are ordering books with ISBN. So I take that they are tracking what books they have scanned or pirated already and only picking up what they are missing. As just ordering mass bulk and getting 20 of the same encyclopaedia would be waste.

And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.

by lo_fye 9 hours ago

The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.

by JsonDemWitOster 9 hours ago

Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.

Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.

0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.

1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.

2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.

3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.

4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.

by bix6 8 hours ago

> A guest post by Anna’s Archive volunteer “u” (translated from Chinese).

I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?

by emtel 17 hours ago

“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.

If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!

by zahlman 5 hours ago

The object itself is the valuable thing, not necessarily the text. Why are so many in the discussion overlooking something so obvious?

by bhouston 9 hours ago

Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.

by infecto 9 hours ago

I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.

I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.

by philipswood 16 hours ago

I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting.

I have found that often LLMs are familiar with the contents of SF books.

But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.

by Alifatisk 16 hours ago

Does it exist in the web at all? Could be the contents have not been scraped for the dataset yet.

by hoppp 10 hours ago

Should be illegal to burn books for this reason. Burning 1 book as a protest, fine. But burning a lot to destroy information? Hell no.

by xvilka 16 hours ago

It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.

by Cider9986 16 hours ago

Anna's archive estimates that they have preserved 16% of the world's books.

https://annas-archive.gl/faq

by iLemming 5 hours ago

What's infuriating about the book digitization is that you'd think: "Oh, great, this would allow finding a fact or a piece from just about any book ever published..."

OMG, finding a verbatim quote from any given book these days is nearly impossible. Every LLM would state "tis a copyrighted material and I can't share it as is". And search engines now all being AI-driven won't find it either. WTF are you even talking about? I'm not trying to steal the whole plot of the book for my dissertation, I just need the exact quote from the book, just like the author intended, don't give me your "rephrased" adaptation of it. What happens in a few years when every single quote is some misinterpreted shit and nobody even knows what the original quote ever was?

by altcognito 18 hours ago

What evidence do we have that they are "destroying" books?

I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)

All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)

by knowaveragejoe 17 hours ago

Yes, we know they're destroying the books. Whether this is actually a problem or another convenient "AI bad" trope remains to be seen.

https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...

by altcognito 9 hours ago

Thank you! I wanted something beyond an article, but this was a perfect launching point for getting to the court case information:

https://www.courtlistener.com/docket/69058235/554/21/bartz-v...

by arttaboi 5 hours ago

This is frustrating. Just to beat the competition and make a few extra bucks, they’re willing to destroy a century’s worth of human knowledge (good or bad).

by qwertytyyuu 18 hours ago

I'm sure the AI companies will retain scans of the books for training on newer models

by premtonx 7 hours ago

I don't see any problem and don't see destroying books.

I think books will exist but they will be written with the help of AI.

The context will still be a human mind behind the words in the book.

by landgenoot 19 hours ago

Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).

by eru 19 hours ago

Most old books aren't worth protecting, and the publishing industry destroys oodles of them as waste that no one wants.

by fmajid 15 hours ago

Aren’t publishers required to deposit a copy with the Library of Congress (in the US), or the British Library (in the UK) etc. to claim copyright?

by ColdStream 19 hours ago

The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.

Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.

go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.

I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.

by ForHackernews 10 hours ago

Project Unica is an initiative by the University of Illinois Libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...

by luciana1u 9 hours ago

the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years

by luciana1u 19 hours ago

Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.

by globnomulous 19 hours ago

Is this a reference to Fahrenheit 451?

by everyday7732 12 hours ago

How long until someone starts making fake rare books to sell to AI companies?

by cyberrock 7 hours ago

Looking at my modest shelf of weird midcentury travel logs, I just want to be spared from this moralizing. This is just another case of everyone seeing store shelves as an extension of themselves, just like the decline of other physical media. If their knowledge was so precious then why was it not on the used bookstores and libraries to save them, especially when scanning has become so accessible in the last decade? Why did I find some of these books rotting in overpacked shelves and boxes?

by tescreal 7 hours ago

I have a tidy collection of books that are odd or fascinating to me. They're old. But you mistake people like me for being made of money.

Old books shouldn't simply vanish. That is history, art, authorial creation. Once the last copy gets crisped, it is lost.

by cyberrock 6 hours ago

Some of my books were literal pennies, and they would vanish too if neither of us nor the companies bought them. Again, this is just confusing where these books are with the Louvre. Just because they were available for purchase doesn't mean they were a shared community resource, because the community didn't care about their fate until now. I just don't appreciate the lack of self-awareness.

by SquireBuilds 8 hours ago

Do you think they will make all of that publicly available after it's been scanned?

by ohthanks 18 hours ago

Being purchased and juiced for model weights is about as noble of an end as any book could hope for.

by bawolff 18 hours ago

This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.

by derektank 17 hours ago

I mean, everything about intellectual property is kind of inherently artificial tbf.

by blooalien 18 hours ago

> It is 100% the consequence of stupid laws.

More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...

by mplewis 18 hours ago

Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.

by TomK32 16 hours ago

I've read a few of those articles in recent months, both in English and German, and I did read any book title that was rare. The Rare Book & Special Collection Div at the Library Congress considers books published before 1801 as rare. Searching for old books on abebooks is surprisingly hard but I didn't find any for less than 10 Euro and nothing in those articles suggested they were buying up anything but cheap books.

by carlosjobim 10 hours ago

I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.

by RcouF1uZ4gsC 10 hours ago

What often gets missed is that they are buy one physical copy and turning it into a digital copy.

They have done zero to destroy the durability. In fact, it’s probably more durable.

If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.

by nijave 9 hours ago

>that is due to the publisher and copyright laws

Agree. While it certainly isn't the most environmentally friendly to render huge stacks of paper into waste, the real issue is copyright creating scarcity (inability to copy the thing).

by JohnFen 10 hours ago

> turning it into a digital copy.

Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?

by tsukikage 9 hours ago

> Or are they just using it for training

You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.

by Filligree 10 hours ago

They are. They can’t release it; that would be copyright infringement. But they’re absolutely planning to make further use of the book later.

by SkyBelow 10 hours ago

The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.

As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.

by mycall 19 hours ago

Aren't AI companies all about the rare book auctions now?

by waffletower 6 hours ago

This reads as bad propaganda. Surprised that they left the "translated from Chinese" disclaimer. There is the false implication here that the AI companies are somehow tracking down all copies of a particular book and destroying it. AI has spawned a global freak-out.

by c0lpan1c 19 hours ago

that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.

by alightsoul 19 hours ago

Use a vpn

by aeon_ai 8 hours ago

This is only because this is a legal right granted as part of the purchase of copyrighted material, and because we have tried to stop AI companies from doing this with purely digital copies.

The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.

by bethekidyouwant 8 hours ago

I don’t get this latest anti AI talking point. They are digitizing the books preserving them forever.. are you upset that you don’t have access to it? Because you didn’t before either… stop whining and give AA some money.

by protocolture 17 hours ago

>It’s outrageous is that it’s legally permissible

No its not.

>but ethically, it’s an extremely serious crime against humanity.

Its only a crime if they dont also upload the scans to the internet.

>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.

This Law on the other hand is a crime against humanity.

>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.

No it doesnt.

>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.

This however is an unvarnished good.

Look, piracy is the only realistic media archive we have.

We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.

This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.

by stuaxo 7 hours ago

I hate that to save this article about an AI company acting abhorently I have to 'favourite' it.

by Razengan 9 hours ago

Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.

Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.

by SanjayMehta 18 hours ago

Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.

Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.

by userbinator 17 hours ago

Does Anna's Archive have it now? They scrape tons of sources, Google Books and HathiTrust included.

by qingcharles 16 hours ago

A lot of Hathi is locked behind university and library access restrictions. I sometimes have to track down students or someone who has a local library card to get items I need.

by SanjayMehta 11 hours ago

Not that I know of, I did upload the PDF to the original libgen but that particular domain has disappeared now.

by qingcharles 16 hours ago

I agree they're the best currently available, but a lot of their scans are straight garbage and need to be redone.

by SanjayMehta 11 hours ago

A garbage scan is better than no book at all. Fortunately the book I needed was a clean scan.

by michael0church 8 hours ago

What’s really disgusting is how unnecessary this is.

LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.

And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.

by starkd 9 hours ago

How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.

by cormorant 9 hours ago

They evade judicial enforcement and try to stay anonymous. They've certainly already been sued (and lost by not even turning up). https://en.wikipedia.org/wiki/Anna%27s_Archive#March_2026_pu... (Contrast to Sci-Hub trying to defend themselves in court, in India, and that leading to no new articles since several years ago.)

by TaLiTr 8 hours ago

Unlike archive.org which is a real business, Anna's Archive is basically piracy. You could try to sue them but you'd have to find them first. This IMO makes them more resilient against that sort of thing, and also means they can actually do their job.

by archonis 9 hours ago

Anna's Archive merely indexes copyrighted content. It doesn't host.

by greenavocado 9 hours ago

During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions

by danparsonson 9 hours ago

They did rather bring that on themselves though

by timcobb 9 hours ago

But think of the books!

by m00dy 10 hours ago

Since when books have become a supply limited asset ?

by simmerup 10 hours ago

Try and read a book that’s been burnt and find out

by JsonDemWitOster 10 hours ago

Ever since someone dredged up this old nothingburger of news from 2024 and made it the latest outrage bait.

by eulgro 10 hours ago

We've been seeing that headline for a few weeks now and I really don't understand the problem.

Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.

Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.

So what's the problem here exactly?

Also from the article:

> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.

I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.

by squidbeak 10 hours ago

> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.

If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).

https://www.bbc.com/news/articles/cp3rprx2wl4o

> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.

> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.

> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."

by JsonDemWitOster 10 hours ago

Is there any indication that Anthropic is destroying books from the 18th century? Even the BBC quote is a conditional. Emphasis added:

> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."

I agree with the sentiment of course but it is really a huge IF they are doing that.

IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.

I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?

by quietsegfault 10 hours ago

Rare and out of print does not mean important or valuable.

by squidbeak 10 hours ago

It doesn't automatically mean pointless or worthless, either. And it's too late to judge either way once it's been destroyed.

by Filligree 10 hours ago

Usually it means the opposite. Books that are old and valuable tend to be out of copyright, so they do see new printing runs.

by squidbeak 10 hours ago

The British Library has a large quantity of books that have literary, academic, cultural or bibliopolical significance but are too niche to ever see print new print runs.

This idea that there can only be merit in a work if it's commercially viable is incredibly ignorant, and if that becomes the standard for whether a work is preserved or not, we stand to lose a great deal of our cultural heritage.

by geye1234 10 hours ago

Presumably, being from the 18th Century, copyright law wouldn't apply?

Detestable if they're doing it anyway to prevent competitors getting hold of it.

by xandrius 10 hours ago

The problem is that you probably do little research or read very few old books.

There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.

You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.

Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.

by brainwad 10 hours ago

Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.

by ForHackernews 10 hours ago

I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.

by brainwad 9 hours ago

The books being destroyed are not afaik those sort of books. They are just out of print modern books, that are still under copyright (and hence can't legally be scanned non-destructively).

by joshstrange 8 hours ago

> They are just out of print modern books, that are still under copyright

OR in-print modern books that they can get for cheaper by buying used. The whole thing is a manufactured outrage over something that doesn't matter. They aren't breaking into museums to steal their only copy of a book and burn it. They are digitizing books. If anything I commend them for what they are doing. If the alternatives were that book rotting on a shelf or being thrown away they doing a great service preserving it, even if they don't make it available publically (which they can't for copyright reasons). It's literally no different from them stocking a private library with these books, except it's better because digital copies are much easier to preserve.

by ab71e5 10 hours ago

It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.

by quietsegfault 10 hours ago

I’m confused - are you intending to say Anthropic and Google are buying ALL the books?

by josem 10 hours ago

I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.

by FartyMcFarter 10 hours ago

The problem is we don't know what we're losing, due to lack of transparency.

by anon373839 10 hours ago

> We've been seeing that headline for a few weeks now and I really don't understand the problem.

It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.

by JohnFen 10 hours ago

> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.

Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.

But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.

by WillAdams 19 hours ago

"Whoever destroys a book destroys a link in the chain of human knowledge"

-- Thos. Jefferson

by pkaye 18 hours ago

Public libraries destroy unsold book donations all the time. I often tried to give away some old books I have online and nobody wants them. Some of these books have some nostalgic value to me so I hate to see them just get destroyed so they just lie in my shed.

by winrid 17 hours ago

I have a copy of Michael Abrash's Graphics Programming Black Book (it's like 1k+ pages) with DESTROY written in red on the sides. I appreciate that someone saved it and sold it to me for cheap :)

by userbinator 17 hours ago

That's an example of a "very much NOT rare" book, as you can easily find dozens of sources of scans online.

by winrid 3 hours ago

I wanted a physical copy, good ones are a few hundred bucks.

by pkaye 17 hours ago

I have one of those I got second hand also. :)

by WillAdams 12 hours ago

Which is something I argue against constantly (and at least my local library tries hard not to) --- discarded books are placed on tables near the children's area usually and folks are free to pick them up (it might be that a few of them are sold, I certainly get a lot of ex-library books when buying on Thriftbooks and Better World Books).

A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.

The local libraries also accept book donations for an annual fund-raising sale.

by warkdarrior 10 hours ago

> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.

Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.

by throwatdem12311 10 hours ago

At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs.

by torh 10 hours ago

At least there will be a copy left for us. The AI companies won't share these books in their original form.

by brainwad 10 hours ago

Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things.

Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

by Paratoner 10 hours ago

The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place

by JohnFen 10 hours ago

> because copyright law forces them to do stupid things.

Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.

by brainwad 9 hours ago

Let us not forget they own the books. They bought them legally on the open market. Also, no modern book truly is destroyed forever, there are always libraries of record with a copy.

by JohnFen 7 hours ago

> Let us not forget they own the books.

That's rather beside the point. I don't think anyone is arguing that what they're doing is in some way illegal.

> no modern book

I'm less concerned about modern books.

by brainwad 7 hours ago

All the books being destroyed are modern enough to have deposited copies, or they wouldn't be still under copyright. The destruction is happening to satisfy judges that no illegal copying is happening, only a transformation of medium.

by xandrius 10 hours ago

Google probably wanted to sell the whole of Google Books.

by subscribed 10 hours ago

Copyright law didn't force them to torrent terabytes of books what they did and got caught doing so.

I'm not convinced they destroy the books to obey the law, lol.

by naasking 10 hours ago

Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.

by embedding-shape 10 hours ago

> because copyright law forces them to do stupid things

This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?

Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?

Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.

by brainwad 10 hours ago

It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯

I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.

by embedding-shape 9 hours ago

But why do they have to do this at all? If it's a shit idea, and you cannot do something without negative side-effects of it, can't you just not do it? Why these companies absolutely have to do this?

by brainwad 8 hours ago

Well, they want to use the book they own in digital format. They own it, they get to decide what to do with it.

by famouswaffles 9 hours ago

What's shit about it ? They're buying books that would be headed for the pump or trash heap anyway. Millions of books are trashed or pulped every day.

by TaLiTr 5 hours ago

> Then the AI companies can simply download a copy of Anna's Archive.

And so can everyone else. That's the ideal outcome.

Only Anti-AI types are against this, they probably don't even care about the books, it's just a proxy for trying to stop the "evil AI companies."

by voidhorse 10 hours ago

Yeah, which is completely fine. There's a major difference between:

A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:

- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.

- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.

- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?

B: Company uses freely available scanned copy of the text:

None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.

I much prefer B.

by Filligree 10 hours ago

B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to.

by voidhorse 10 hours ago

Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?

Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.

by zuzululu 6 hours ago

again "destroy" is not accurate. to scan massive amounts of books you have to cut the binding part as feed scanning has been around for ages

there are non destructive scanning options but they are nowhere as fast and prone to errors

by whycombinetor 19 hours ago

It's giving Vishnu, but the world cannot exist without Shiva.

by christkv 13 hours ago

Do you mean rare as in old? I doubt they are destroying old books because most of them are out of copyright and probably available already as text. I imagine this applies to in copyright works and I do NOT condone it but the way this is told it sounds like they are raiding old libraries to destroy first editions of Cervantes.

by imperfect_light 17 hours ago

People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.

by Ekaros 16 hours ago

I feel they are thinking first printing of some mega classic or hand drawn book from antiquity. I am thinking of some generic romance or thriller from no name author bought at airport... Or some generic book on say birds or animals... Stuff that they would not take if given a full box for free.

by tokai 5 hours ago

Any named book I have seen in the coverage of this are cheap crap books, that are available somewhere in the word from a public library. Its literal trash.

by lazzlazzlazz 14 hours ago

Aren't the AI companies just buying one copy of each book?

And this is supposed to be concerning?

by tkel 13 hours ago

The blog post is about rare books. Meaning, AI companies scanning and destroying rare books.

by lazzlazzlazz 5 hours ago

Yes, read more carefully: one copy. Do you panic when the "wrong person" buys a single rare book?

The issue is that the copyright holders and book publishers make it hard to make more copies, not that somebody bought a single copy of a book, no matter how rare.

If you could print any book on demand (paying for it), the issue would vanish instantly. And who makes that decision?

by GreenLightGo 7 hours ago

It all sounds like some kind of conspiracy theory, but over the past five years, a lot of conspiracy theories have turned out to be true. That’s pretty creepy.

by BrenBarn 19 hours ago

It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".

by shevy-java 9 hours ago

But scanning the books also helps those AI companies because ultimately they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".

by ahoka 9 hours ago

I guess they just cut the spines to speed scan the pages with a machine and then the resulting pile of papers is no longer worth anything, so they just dispose of it.

by taintify 8 hours ago

Savonarola Altman

I’ll leave

by tamimio 18 hours ago

I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..

by TomK32 16 hours ago

Books have been declared dead several times of the last quarter century, yet in the EU alone it's still half a million new titles every year and a 40 billion Euro market.

by ivell 16 hours ago

Good premise for a scifi novel.

by imperio59 19 hours ago

Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(

by qingcharles 16 hours ago

Not everyone should scan stuff. If you spend any serious time looking through stuff that randos on the Internet have scanned the quality fits the Bell Curve perfectly.

Biggest problems:

  - scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
  - becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.

Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.

Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:

https://imgur.com/a/vgMkM7b

(preview size, they sent a 500MB TIFF)

Now I can reassemble the issue and upload it.

I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.

I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.

by userbinator 17 hours ago

In the context of books, "scanning" is now more commonly something that should be called "camming" --- you simply point a camera at the book, and take a picture of every page.

by next_xibalba 8 hours ago

Why are these books "rare"? Because no one wants them. Why then are we up in arms over their destruction? These sensational headlines make it seem as though a copy of the Codex Sassoon 1053 is being destroyed, when in fact these are just obscure books that no one cares about.

The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.

by partiallypro 18 hours ago

The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.

by svachalek 7 hours ago

It's weird what happens when you keep warning everyone you're building a product that will probably end human civilization.

by jiaosdjf 14 hours ago

BOYCOTT.

It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.

This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.

Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.

by maxlin 9 hours ago

It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.

Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.

by spwa4 10 hours ago

The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.

Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.

Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.

If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?

But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.

Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.

To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?

But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.

I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!

The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.

Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:

a) EU companies making ML models have to self-sabotage against their competition.

b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.

Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.

[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...

[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client

by josefritzishere 9 hours ago

Destroying history is a crime against humanity.

by ionwake 9 hours ago

I mean guys i get it, you have a bad day just once and you want to rewrite history, we all do, but geez... I guess just I feel destroying books should be classed as unethical. if i was ruler I think id ban it.

by joshstrange 8 hours ago

Yes, let's begin with the trials for all the librarians who have destroyed/trashed books over the year!

Nevermind that they do that because those books aren't being borrowed and no one wants to buy them. Destroying books to destroy knowledge is bad, destroying books because no one wants them or there are ample copies, _especially_ when those books go on to live forever as a digital copy, is a completely different story.

It boggles my mind that of all places, here on HN, people can't understand that. The knowledge is not destroyed, only the physical vessel.

by ionwake 8 hours ago

bro if people on HN of all places are concerned of physical media being destroyed boggles ur mind then u need to read more about the damage caused to society by prior happenings.

I wasn't even making an unusual or particularly strong statement

Im going to try and be patient and explain my position. I just feel that actions that are destructive should always be seen through a lens of suspicion.

IE, lets destroy this wood to reduce the chance of wildfires. Ok, well why that wood? we sure there arent other interests at play? Ie a carpark wanting to be built by a local authority?

As you become older and witness large acts occur due to unclear scheming, you start to become sensitive towards destructive acts, feeling there should be a lens cast over it.

Maybe Im just paranoid.

Sure I can understand there must be tons of books that are not needed, but books, are object that include native function that IN the right context becomes priceless and hence should be treated with a degree of concern.

by joshstrange 7 hours ago

Nothing is destroyed, only transformed. What was a physical book is now a digital book. It's a completely different story from "prior happenings", to which I assume you mean book burning and the like. Pretending they are the same is intellectually lazy.

by ionwake 7 hours ago

Thank you for your reply and request for clarification.

Yes destroying a book is in my opinion identical to a book burning.

One day there will be no hard copies, the digital books are hence vulnerable to patches, whether its due to a new political movement or a sudo abled hamster running on a keyboard... and if that occurs information will be permanently lost.

I would have thought a HN user understand well the importance of backups.

I dont want to sound mean, perhaps you also support keeping some hard copies as backups, but Im just explaining my concern.

To be honest I think everything I am saying is just common sense, I was just making a joke earlier about the architects deciding to do this because they had a bad tuesday.

by joshstrange 7 hours ago

There are multiple digital ways to preserve exact copies of the original. Both through things like personal backups (I have digital backups of all my physical books) and through collective recordkeeping. For example, Anna's Archive stores them (or at least provides access) via a hash of the book. You can't patch the book and keep the same hash. Yes, hash collision is a thing but there are better hashes or other ways to accomplish the same goal. It pains me to say "blockchain" but that's something we have today that could be leveraged, though I think there are other cryptographic methods we could employ.

I'll admit I have less concern with (every) digital copy being altered and if we assume a future where all computing is completely locked down and governments have the ability to reach in a tweak anything then yes, we are screwed. But we are screwed if we allow that to happen even if some of us have squirreled away physical books. In that techno-hell (of completely locked down computing with no open options) then some of use will have "illegal" computers that still do what we want and I see that as no different than keeping physical copies.

Put simply: Preserving the original does not require a physical substrate (or at least one made of paper, obviously computers run on physical hardware).

by zahlman 5 hours ago

> Nothing is destroyed, only transformed. What was a physical book is now a digital book.

I genuinely don't understand how it's possible for someone to type that out in all sincerity. Have you ever held a physical book?

by joshstrange 5 hours ago

Have you held a kindle? What does that even mean? I vastly prefer digital books to physical ones. I own many physical books (as well as their ebook equivalent) but the physical books on my shelves are nothing more than art to me, I won't read the physical books because the experience is infinitely better when reading on my (jailbroken) kindle.

by everdrive 6 hours ago

Destroy human knowledge, skyrocket RAM prices, push up residential electricity prices, potentially infantilize a generation of young people. It's all worth it, of course. Every time you _can_ invent a technology, you _must_ invent it. No externality is worth considering.

by maxdo 10 hours ago

Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.

Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.

It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.

Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court

by cdrnsf 4 hours ago

We should be destroying AI companies, not books.

Data from: Hacker News, provided by Hacker News (unofficial) API