I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
show comments
ezfe
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
show comments
NishanStepak
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
show comments
ziyadb
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
show comments
HedonicEscal8r
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
show comments
odyssey7
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
show comments
akk0
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
show comments
ironqcold
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.
But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.
show comments
skeledrew
Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
show comments
CamelCaseName
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
show comments
pmoriarty
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
show comments
sieve
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
show comments
glimshe
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
show comments
fastball
So we don't want companies to buy books and scan them and do whatever they want afterwards, but we also don't want to allow piracy of digital copies (the $1.5B Anthropic settlement)? Bit of a rock and a hard place for them.
show comments
branon
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
show comments
theartfuldodger
I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.
I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.
xvxvx
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
ryandvm
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
show comments
jupp0r
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
joshuakelleyds
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:
> Anthropic spent many millions of dollars to purchase millions of print
books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order
Also the summary:
> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central
library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
show comments
shrubble
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
show comments
rcarr
Does adding an old book risk making the model worse? Off the top of my head:
- Reinforcing outdated, disproved or otherwise incorrect information.
- Reinforcing outdated forms of communication e.g purple prose.
You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory
ZoomZoomZoom
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
show comments
thisisauserid
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
chistev
What is the deal with AI Companies buying old books to scan and then destroy them?
Many countries require a copy of each book published there to be submitted to their national archives/library. The US has had this requirement since 1790 as far as I can tell.
Thus, most books should not be at risk of getting permanently lost due to these practices.
This is Anna's Archive (ab)using a current controversy (which I think is based around emotional appeals and distorted facts to create outrage over a non-issue) to tongue-in-cheek advertise their open pirate library.
abunner
AI companies are supporting the market for books that no one else wanted. The economically illiterate assumption of the author is that "rare" books are good. If they were that good, they would be priced higher.
destructive book scanning is pretty standard, it's the only way to efficiently get a good flat scan of the pages without a tremendous amount of human labor to correct distortion on every individual page. i'm curious about what specific rare books are being destroyed and how rare they really are. and at any rate, anyone can go out and buy books and scan them or destroy them or set them on fire or whatever, they bought the book it's their own property. so now physical books are treated as the public commons while intellectual property is privately owned? are we in topsy-turvy world? i feel like there are much more egregious crimes that people ought to be talking about with ai companies, like impoverished black people having their city water supplies poisoned or the massive financial fraud that will destroy the economy or the massive co2 emissions that will destroy the world or just the fact that it's not even artificial intelligence at all and it's not even a technological innovation, it's just a google hack that dumbed down an existing neural network model enough for it to run on lots of nvidia gpus and produce an impressive-enough tech demo to show to gullible investors who don't know what to do with their massive piles of money.
1970-01-01
This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.
tescreal
A few points for people:
1. some books are out of print.
2. some books CANNOT return to print.
3. all books prior to the 21st century are products of human minds.
4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access.
5. not all people value all books equally.
6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them).
7. many books are already lost, and historians are the ones who most voice the harms.
8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it.
9. if no existant copies remain, then the price for access becomes effectively infinite.
10. removal of books denies human agency over access to information.
11. costs will follow a steepening curve much as ram did.
first they came for cookbooks, but i was no chef so i said nothing.
second they came for handicraft, but i do not toil with fabrics or glue.
next they came for homesteading, but i loathe the outdoors life.
after, they came for biography, memoirs, and letters, but i am bored by the dead.
finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.
thuruv
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
show comments
Cider9986
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
show comments
yipinwong
The post reads like a propaganda.
Evoking feeling over metrics or results
(<- that's what propaganda does by definition)
e.g.
> but ethically, it’s an extremely serious crime against humanity.
Who are you that you consider it a serious crime against humanity?
What are you a saint?
> Knowledge is permanently monopolized on private servers.
Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.
With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.
The logic of the writing is not correct.
show comments
clarionbell
In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
pino83
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?
And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?
Why is that?
If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?
twright
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
throwaw12
America is interesting.
* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz
* download and publish a book as a company -> fine 1% of revenue
* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations
* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
silcoon
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
show comments
juiceland
Why are people acting like books can’t be reprinted?
show comments
jmspring
I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.
lukasbm
A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"
godber
I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.
Clearly that makes assumptions about the function of government and the complacency of copyright holders.
show comments
thuuuomas
Is there any evidence the books are truly "destroyed" & not merely "disassembled"?
It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
1970-01-01
Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
show comments
dbgrman
Hmm... interesting. previously this was annas-archive.pk, now its on gl domain. What happened?
mmaunder
At the risk of angering the zeitgeist, destroying one or two or ten paper copies doesn’t make AI companies “become the only ones in the world with digital copies.”.
show comments
pfdietz
Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...
finn888
The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.
show comments
ErigmolCt
A genuinely useful project would identify publications that are actually rare, poorly catalogued, or held by only a few libraries, then prioritize those for careful preservation
show comments
bix6
> A guest post by Anna’s Archive volunteer “u” (translated from Chinese).
I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
lo_fye
The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.
JsonDemWitOster
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
tptacek
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
show comments
jsphweid
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
show comments
emtel
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
show comments
bhouston
Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
infecto
I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.
I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
philipswood
I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
show comments
hoppp
Should be illegal to burn books for this reason. Burning 1 book as a protest, fine. But burning a lot to destroy information? Hell no.
iLemming
What's infuriating about the book digitization is that you'd think: "Oh, great, this would allow finding a fact or a piece from just about any book ever published..."
OMG, finding a verbatim quote from any given book these days is nearly impossible. Every LLM would state "tis a copyrighted material and I can't share it as is". And search engines now all being AI-driven won't find it either. WTF are you even talking about? I'm not trying to steal the whole plot of the book for my dissertation, I just need the exact quote from the book, just like the author intended, don't give me your "rephrased" adaptation of it. What happens in a few years when every single quote is some misinterpreted shit and nobody even knows what the original quote ever was?
xvilka
It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.
show comments
altcognito
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
show comments
arttaboi
This is frustrating. Just to beat the competition and make a few extra bucks, they’re willing to destroy a century’s worth of human knowledge (good or bad).
premtonx
I don't see any problem and don't see destroying books.
I think books will exist but they will be written with the help of AI.
The context will still be a human mind behind the words in the book.
qwertytyyuu
I'm sure the AI companies will retain scans of the books for training on newer models
landgenoot
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
show comments
fmajid
Aren’t publishers required to deposit a copy with the Library of Congress (in the US), or the British Library (in the UK) etc. to claim copyright?
ColdStream
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
everyday7732
How long until someone starts making fake rare books to sell to AI companies?
luciana1u
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
show comments
cyberrock
Looking at my modest shelf of weird midcentury travel logs, I just want to be spared from this moralizing. This is just another case of everyone seeing store shelves as an extension of themselves, just like the decline of other physical media. If their knowledge was so precious then why was it not on the used bookstores and libraries to save them, especially when scanning has become so accessible in the last decade? Why did I find some of these books rotting in overpacked shelves and boxes?
show comments
SquireBuilds
Do you think they will make all of that publicly available after it's been scanned?
ohthanks
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
bawolff
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
show comments
mplewis
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
show comments
carlosjobim
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
RcouF1uZ4gsC
What often gets missed is that they are buy one physical copy and turning it into a digital copy.
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
show comments
mycall
Aren't AI companies all about the rare book auctions now?
waffletower
This reads as bad propaganda. Surprised that they left the "translated from Chinese" disclaimer. There is the false implication here that the AI companies are somehow tracking down all copies of a particular book and destroying it. AI has spawned a global freak-out.
c0lpan1c
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
show comments
aeon_ai
This is only because this is a legal right granted as part of the purchase of copyrighted material, and because we have tried to stop AI companies from doing this with purely digital copies.
The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.
bethekidyouwant
I don’t get this latest anti AI talking point. They are digitizing the books preserving them forever.. are you upset that you don’t have access to it? Because you didn’t before either… stop whining and give AA some money.
protocolture
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
stuaxo
I hate that to save this article about an AI company acting abhorently I have to 'favourite' it.
Razengan
Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
michael0church
What’s really disgusting is how unnecessary this is.
LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.
And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
starkd
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
show comments
SanjayMehta
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
show comments
greenavocado
During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions
show comments
m00dy
Since when books have become a supply limited asset ?
show comments
eulgro
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
show comments
warkdarrior
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
show comments
WillAdams
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
show comments
zuzululu
again "destroy" is not accurate. to scan massive amounts of books you have to cut the binding part as feed scanning has been around for ages
there are non destructive scanning options but they are nowhere as fast and prone to errors
whycombinetor
It's giving Vishnu, but the world cannot exist without Shiva.
christkv
Do you mean rare as in old? I doubt they are destroying old books because most of them are out of copyright and probably available already as text. I imagine this applies to in copyright works and I do NOT condone it but the way this is told it sounds like they are raiding old libraries to destroy first editions of Cervantes.
imperfect_light
People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.
show comments
lazzlazzlazz
Aren't the AI companies just buying one copy of each book?
And this is supposed to be concerning?
show comments
GreenLightGo
It all sounds like some kind of conspiracy theory, but over the past five years, a lot of conspiracy theories have turned out to be true. That’s pretty creepy.
BrenBarn
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
shevy-java
But scanning the books also helps those AI companies because ultimately
they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
show comments
taintify
Savonarola Altman
I’ll leave
tamimio
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
show comments
imperio59
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
show comments
next_xibalba
Why are these books "rare"? Because no one wants them. Why then are we up in arms over their destruction? These sensational headlines make it seem as though a copy of the Codex Sassoon 1053 is being destroyed, when in fact these are just obscure books that no one cares about.
The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.
partiallypro
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
show comments
jiaosdjf
BOYCOTT.
It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.
This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.
Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.
maxlin
It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
spwa4
The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.
I mean guys i get it, you have a bad day just once and you want to rewrite history, we all do, but geez... I guess just I feel destroying books should be classed as unethical. if i was ruler I think id ban it.
show comments
everdrive
Destroy human knowledge, skyrocket RAM prices, push up residential electricity prices, potentially infantilize a generation of young people. It's all worth it, of course. Every time you _can_ invent a technology, you _must_ invent it. No externality is worth considering.
maxdo
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.
But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.
Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
So we don't want companies to buy books and scan them and do whatever they want afterwards, but we also don't want to allow piracy of digital copies (the $1.5B Anthropic settlement)? Bit of a rock and a hard place for them.
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.
I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order
Also the summary:
> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
Does adding an old book risk making the model worse? Off the top of my head:
- Reinforcing outdated, disproved or otherwise incorrect information.
- Reinforcing outdated forms of communication e.g purple prose.
You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
What is the deal with AI Companies buying old books to scan and then destroy them?
https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...
Many countries require a copy of each book published there to be submitted to their national archives/library. The US has had this requirement since 1790 as far as I can tell.
All countries I checked (Germany, France, Canada) seem to have similar requirements. WIPO claims that this is the case in the majority of countries (https://www.wipo.int/documents/d/copyright/docs-en-registrat...).
Thus, most books should not be at risk of getting permanently lost due to these practices.
This is Anna's Archive (ab)using a current controversy (which I think is based around emotional appeals and distorted facts to create outrage over a non-issue) to tongue-in-cheek advertise their open pirate library.
AI companies are supporting the market for books that no one else wanted. The economically illiterate assumption of the author is that "rare" books are good. If they were that good, they would be priced higher.
https://abunner.substack.com/i/210907372/anthropic-is-suppor...
destructive book scanning is pretty standard, it's the only way to efficiently get a good flat scan of the pages without a tremendous amount of human labor to correct distortion on every individual page. i'm curious about what specific rare books are being destroyed and how rare they really are. and at any rate, anyone can go out and buy books and scan them or destroy them or set them on fire or whatever, they bought the book it's their own property. so now physical books are treated as the public commons while intellectual property is privately owned? are we in topsy-turvy world? i feel like there are much more egregious crimes that people ought to be talking about with ai companies, like impoverished black people having their city water supplies poisoned or the massive financial fraud that will destroy the economy or the massive co2 emissions that will destroy the world or just the fact that it's not even artificial intelligence at all and it's not even a technological innovation, it's just a google hack that dumbed down an existing neural network model enough for it to run on lots of nvidia gpus and produce an impressive-enough tech demo to show to gullible investors who don't know what to do with their massive piles of money.
This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.
A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.
first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
The post reads like a propaganda. Evoking feeling over metrics or results
(<- that's what propaganda does by definition)
e.g.
> but ethically, it’s an extremely serious crime against humanity.
Who are you that you consider it a serious crime against humanity? What are you a saint?
> Knowledge is permanently monopolized on private servers.
Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.
With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.
The logic of the writing is not correct.
In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?
And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?
Why is that?
If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
America is interesting.
* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz
* download and publish a book as a company -> fine 1% of revenue
* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations
* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
Why are people acting like books can’t be reprinted?
I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.
A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"
I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.
Clearly that makes assumptions about the function of government and the complacency of copyright holders.
Is there any evidence the books are truly "destroyed" & not merely "disassembled"?
It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
Hmm... interesting. previously this was annas-archive.pk, now its on gl domain. What happened?
At the risk of angering the zeitgeist, destroying one or two or ten paper copies doesn’t make AI companies “become the only ones in the world with digital copies.”.
Does the demonization of destruction of books mean I shouldn't delete downloaded e-books off my phone's Kindle app? All those precious bits, gone forever...
The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.
A genuinely useful project would identify publications that are actually rare, poorly catalogued, or held by only a few libraries, then prioritize those for careful preservation
> A guest post by Anna’s Archive volunteer “u” (translated from Chinese).
I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.
I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
Should be illegal to burn books for this reason. Burning 1 book as a protest, fine. But burning a lot to destroy information? Hell no.
What's infuriating about the book digitization is that you'd think: "Oh, great, this would allow finding a fact or a piece from just about any book ever published..."
OMG, finding a verbatim quote from any given book these days is nearly impossible. Every LLM would state "tis a copyrighted material and I can't share it as is". And search engines now all being AI-driven won't find it either. WTF are you even talking about? I'm not trying to steal the whole plot of the book for my dissertation, I just need the exact quote from the book, just like the author intended, don't give me your "rephrased" adaptation of it. What happens in a few years when every single quote is some misinterpreted shit and nobody even knows what the original quote ever was?
It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
This is frustrating. Just to beat the competition and make a few extra bucks, they’re willing to destroy a century’s worth of human knowledge (good or bad).
I don't see any problem and don't see destroying books.
I think books will exist but they will be written with the help of AI.
The context will still be a human mind behind the words in the book.
I'm sure the AI companies will retain scans of the books for training on newer models
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
Aren’t publishers required to deposit a copy with the Library of Congress (in the US), or the British Library (in the UK) etc. to claim copyright?
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Project Unica is an initiative by the University of Illinois Libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
How long until someone starts making fake rare books to sell to AI companies?
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
Looking at my modest shelf of weird midcentury travel logs, I just want to be spared from this moralizing. This is just another case of everyone seeing store shelves as an extension of themselves, just like the decline of other physical media. If their knowledge was so precious then why was it not on the used bookstores and libraries to save them, especially when scanning has become so accessible in the last decade? Why did I find some of these books rotting in overpacked shelves and boxes?
Do you think they will make all of that publicly available after it's been scanned?
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
What often gets missed is that they are buy one physical copy and turning it into a digital copy.
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
Aren't AI companies all about the rare book auctions now?
This reads as bad propaganda. Surprised that they left the "translated from Chinese" disclaimer. There is the false implication here that the AI companies are somehow tracking down all copies of a particular book and destroying it. AI has spawned a global freak-out.
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
This is only because this is a legal right granted as part of the purchase of copyrighted material, and because we have tried to stop AI companies from doing this with purely digital copies.
The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.
I don’t get this latest anti AI talking point. They are digitizing the books preserving them forever.. are you upset that you don’t have access to it? Because you didn’t before either… stop whining and give AA some money.
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
I hate that to save this article about an AI company acting abhorently I have to 'favourite' it.
Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
What’s really disgusting is how unnecessary this is.
LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.
And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions
Since when books have become a supply limited asset ?
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
again "destroy" is not accurate. to scan massive amounts of books you have to cut the binding part as feed scanning has been around for ages
there are non destructive scanning options but they are nowhere as fast and prone to errors
It's giving Vishnu, but the world cannot exist without Shiva.
Do you mean rare as in old? I doubt they are destroying old books because most of them are out of copyright and probably available already as text. I imagine this applies to in copyright works and I do NOT condone it but the way this is told it sounds like they are raiding old libraries to destroy first editions of Cervantes.
People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.
Aren't the AI companies just buying one copy of each book?
And this is supposed to be concerning?
It all sounds like some kind of conspiracy theory, but over the past five years, a lot of conspiracy theories have turned out to be true. That’s pretty creepy.
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
But scanning the books also helps those AI companies because ultimately they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
Savonarola Altman
I’ll leave
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
Why are these books "rare"? Because no one wants them. Why then are we up in arms over their destruction? These sensational headlines make it seem as though a copy of the Codex Sassoon 1053 is being destroyed, when in fact these are just obscure books that no one cares about.
The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
BOYCOTT.
It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.
This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.
Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.
It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.
[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...
[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client
Destroying history is a crime against humanity.
I mean guys i get it, you have a bad day just once and you want to rewrite history, we all do, but geez... I guess just I feel destroying books should be classed as unethical. if i was ruler I think id ban it.
Destroy human knowledge, skyrocket RAM prices, push up residential electricity prices, potentially infantilize a generation of young people. It's all worth it, of course. Every time you _can_ invent a technology, you _must_ invent it. No externality is worth considering.
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court
We should be destroying AI companies, not books.