Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.
Nobody complained when Google did this over a decade ago 🤷
When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.
Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.
Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)
Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.
It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.
I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).
This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.
Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.
Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted
Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.
If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.
I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.
Legal reasons: When you “purchase” an ebook you’re actually just licensing it and nearly all ebook licenses exclude the ability to do anything with the ebook other than read it yourself.
They could get a commercial license to get big ebook libraries like Anthropic did, but they can’t, really, because Anthropic paid for an exclusive license. Which means that if other AI companies want to compete, they sort of have to buy books in bulk and tear them apart to scan them.
True, but irrelevant. Why would they care at all about the terms of a license? Again, they’ll happily engage in outright mass torrenting. Their legal theory is that using works to train LLMs is simply fair use. The ebook sellers may disagree, but it won’t stop them.
Haha: What you state is logical and reasonable. That position would be easy to defend, if the legal system around copyright made sense.
Unfortunately, it isn’t a logical system. You’re trying to apply copyright law—which has now been ruled on in court, officially making training AI with copyrighted material Fair Use.
The problem is that the issue with ebooks is all about contract law. Not copyright.
When you “purchase” an ebook (which is a misnomer), you’re actually signing a legally binding contract. A license, in effect, to use that copyrighted work for the sole purpose outlined in the contract. That outlined purpose expressly forbids using the work for anything other than you—the purchaser—reading it (usually on the platform they specify).
Having said that, many, many court rulings have found numerous license clauses like that to be unenforceable. That is, just because it’s written in the contract, doesn’t mean it’s legal.
The courts have ruled—thanks to everyone’s efforts fighting the MPAA, RIAA, Sony, Microsoft, and Nintendo in the 1990s and early 2000s—that everyone does have the legal right to “platform shift” whatever copyrighted works they own.
To get around that, those very same entities tried to use the Digital Millennium Copyright Act’s rules about circumventing “copyright protection mechanisms” to try to make it illegal for people to platform shift their stuff anyway. That is, they added trivial encryption to all their platforms.
But it’s even more complicated than that! Because Congress gave the Librarian of Congress the power to say when it’s legal to circumvent such “copyright protections.” For example, technologies that aid the blind (I.e. gotta decrypt that file for the program to read it out loud).
There’s other exceptions and everyone fights to get more added whenever it comes up.
The key takeaway, though is that none of this complicated mess applies to physical books! So there ya go 😁👍
Almost all these books were either headed to the dump or the recycling center.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
re you sure they’re digitizing Book By the Yard quality material?
They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.
Doing anything physical, especially at scale, is slow and expensive.
Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.
would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
Possibly not. But that’s a sign it isn’t great literature or coveted texts.
You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.
My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.
And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.
I imagine they were talking about destruction in a practical sense. Disassembling the book and scanning it like that is faster and cheaper than purpose build book scanning machines.
Additionally, court documents indicate they generally just throw them away afterwards. So the knowledge is retained, but that book is destroyed.
The knowledge is retained and owned by a private corporation who now does not have to share what may have been a still under copyright, but now the corporation owns it? Make it make sense.
Now we just need to proof they’ve copied the files onto a second hard drive, so they own two copies
If we’re gonna keep it as a copyright issue, I think people are mostly mad they’re not able to reprint the books themselves if they wanted to. We’re not speaking from a copyright position
If you own a book, you can legally make as many copies of it as you want. As long as you’re not distributing them, the courts treat that as a single effective copy.
Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.
Nobody complained when Google did this over a decade ago 🤷
When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.
Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.
Digitized for private consumption.
Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)
in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner
Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.
Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.
It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.
I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).
https://en.wikipedia.org/wiki/Book_scanning#Methods
This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.
Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.
Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted
Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.
When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.
Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.
If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
It’s ok to throw trash away! Really!
Why would they scan that stuff though? That kind of mass market (human made) slop is easily available in digital form. We know the AI companies have engaged in mass digital piracy, including running massive torrenting operations. So if a digital copy exists, they probably already have it. And even if they have to buy it, purchasing an ebook is a lot cheaper than buying a physical one, shipping it, paying someone to scan it, etc.
I would think the old, the out-of-print, the rare, and never-before digitized are the only things worth buying and physically scanning in 2026. Everything else has already been scanned or was born digital-native.
Legal reasons: When you “purchase” an ebook you’re actually just licensing it and nearly all ebook licenses exclude the ability to do anything with the ebook other than read it yourself.
They could get a commercial license to get big ebook libraries like Anthropic did, but they can’t, really, because Anthropic paid for an exclusive license. Which means that if other AI companies want to compete, they sort of have to buy books in bulk and tear them apart to scan them.
True, but irrelevant. Why would they care at all about the terms of a license? Again, they’ll happily engage in outright mass torrenting. Their legal theory is that using works to train LLMs is simply fair use. The ebook sellers may disagree, but it won’t stop them.
Haha: What you state is logical and reasonable. That position would be easy to defend, if the legal system around copyright made sense.
Unfortunately, it isn’t a logical system. You’re trying to apply copyright law—which has now been ruled on in court, officially making training AI with copyrighted material Fair Use.
The problem is that the issue with ebooks is all about contract law. Not copyright.
When you “purchase” an ebook (which is a misnomer), you’re actually signing a legally binding contract. A license, in effect, to use that copyrighted work for the sole purpose outlined in the contract. That outlined purpose expressly forbids using the work for anything other than you—the purchaser—reading it (usually on the platform they specify).
Having said that, many, many court rulings have found numerous license clauses like that to be unenforceable. That is, just because it’s written in the contract, doesn’t mean it’s legal.
The courts have ruled—thanks to everyone’s efforts fighting the MPAA, RIAA, Sony, Microsoft, and Nintendo in the 1990s and early 2000s—that everyone does have the legal right to “platform shift” whatever copyrighted works they own.
To get around that, those very same entities tried to use the Digital Millennium Copyright Act’s rules about circumventing “copyright protection mechanisms” to try to make it illegal for people to platform shift their stuff anyway. That is, they added trivial encryption to all their platforms.
But it’s even more complicated than that! Because Congress gave the Librarian of Congress the power to say when it’s legal to circumvent such “copyright protections.” For example, technologies that aid the blind (I.e. gotta decrypt that file for the program to read it out loud).
There’s other exceptions and everyone fights to get more added whenever it comes up.
The key takeaway, though is that none of this complicated mess applies to physical books! So there ya go 😁👍
At least digital storage prices have only been going down in recent months, right? /s
Good luck reading by candle light.
I guess you must live in that part of Alaska where it’s night time for half the year and don’t have daylight.
More just the part of Alaska where I work most of the daylight hours and get my reading in before bed.
My light bulb is bigger than yours!
Yes they did?! That was a big controversy back in the day. You are engaging in historical revisionism right now.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
Solid point
They were more than happy to gobble up Reddit posts and Twitter farts and the worst of 4chan. You don’t need to ingest Ulysses exclusively to train a system.
Which is why you get people paid pennies an hour to break the physical copies down and feed them into scanners.
Possibly not. But that’s a sign it isn’t great literature or coveted texts.
You’re confusing uniqueness with quantity. 4chan posts are not Ulysses, but they are still useful data.
My point was not that they would refuse to scan Book By the Yard quality works because they would consider them unsuitable. My point was that they don’t need to scan mass market Book by the Yard stuff, as they already have digital copies of it.
And scanning books is not as cheap as you think it is. Even using labor in low income countries, the cost of scanning is still vastly greater than just downloading a file.
I imagine they were talking about destruction in a practical sense. Disassembling the book and scanning it like that is faster and cheaper than purpose build book scanning machines.
Additionally, court documents indicate they generally just throw them away afterwards. So the knowledge is retained, but that book is destroyed.
https://tagteam.harvard.edu/hub_feeds/3415/feed_items/14341220
The knowledge is retained and owned by a private corporation who now does not have to share what may have been a still under copyright, but now the corporation owns it? Make it make sense.
A robber barron class of psychopaths are running the whole damn show!
Yes. That does make sense.
If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.
Why is it wrong when a corporation does the same thing?
They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.
Now we just need to proof they’ve copied the files onto a second hard drive, so they own two copies
If we’re gonna keep it as a copyright issue, I think people are mostly mad they’re not able to reprint the books themselves if they wanted to. We’re not speaking from a copyright position
If you own a book, you can legally make as many copies of it as you want. As long as you’re not distributing them, the courts treat that as a single effective copy.
Where does it say that in this article?
ship of theseus ah shit