LowEndBox - Cheap VPS, Hosting and Dedicated Server Deals

AI Companies Are Destroying Every Old Book and Shedding Gutenberg Bibles and...Pump the Brakes, Peop

Robot Shredding

As you may have heard, Anthropic recently settled a copyright lawsuit from various authors and agreed to pay $1.5 billion to them.

Since then, they’ve found a legal loophole they can use to keep ingesting words: bulk-purchasing physical books and scanning them.  To make the process efficient, they shear the spines off with a big industrial cleaver and then feed the pages into a high-speed scanner, throwing the paper copy away when they’re done.

Anthropic, for example, has “Project Panama” where they plan to scan two million books.  What is the motivation here?  A federal judge ruled that training AI on legally purchased physical media constitutes “fair use.”  No longer do AI companies need to rely on shadow digital libraries.  Now they can build legally defensible datasets.

 

It’s Not Stealing Now

A couple years ago when Midjourney was all the rage, an artist friend of mine was very upset that artists were being cheated out of their royalties.  I asked what is the difference:

  1. A human artist goes to art school, studies the masters, develops artistic skill, and is then asked to “create a comic book in the style of Frank Miller”.  He buys a bunch of Frank Miller comics, studies them, and then produces the commission.
  2. An AI ingests the corpus of Frank Miller’s work and does the same thing.

My artist friend’s answer was “but AI didn’t buy the works to study”.  Fair point.

But now they have, and this argument is moot.  Even if AI bought the copies second-hand, that’s perfectly legal, the same way a human might buy a second-hand copy.

Take a Chill Pill, People

Ever since this destructive scanning initiative was mentioned, there has been endless wailing and gnashing of teeth over it, mostly from bibliophiles who bemoan the destruction of rare tomes.  My social media feeds have blown up with screeching about how the fruit of human knowledge is being tossed into the AI incinerator, countless generations will be deprived of these priceless tomes, how libraries will soon be empty, etc.  From the commentary, you’d think that Gutenberg Bibles are being sliced up by robots.

Relax.

A little rational thought will calm you down.

First, anything truly old is already out of copyright.  Anthropic, OpenAI, etc. are perfectly free to download public domain copies of any book that is older than its publication date + 70 years + life of the author.  So that’s roughly 150 years, which lands us at about 1876.  But of course, it’s much more recent than that for many books because not everyone lives to 80.  Many book from the early 1900s are out of copyright.

And tons and tons of these books have already been scanned.  Google had a long-running project to non-destructively scan library books, and so did the Internet Archive.  Project Gutenberg and many other sources make them available.  Anthropic is not going to find a first edition of Charles Dickens to chop up when they can easily obtain a digital copy.

Second, how many AI companies are there building these datasets?  Not many.

And finally, each only needs one copy.  Of any edition.  A first edition of a Daphne du Maurier’s 1936 novel Rebecca in hardcover with a beautiful dust jacket is a treasure.  A trade paperback reprint published in 2025 is not.  AI companies are not going to buy the 1936 version – they’re going to buy the cheapest second-hand copy they can.  In fact, they’re buying from used book dealers to get the cheapest possible copies.

If you’re the kind of person who can’t bear to see any book destroyed…okay, grind your molars.  But let’s be realistic here.  Tons of books are destroyed every year.  Not every book is a precious classic.  Back before Amazon, the landscape was dotted with many used book stores, and frequently what didn’t sell got pulped because books are bulky.

Not every book deserves preservation for all time.  In this case, any needed preservation is not going to be impeded.

No Comments

    Leave a Reply

    Some notes on commenting on LowEndBox:

    • Do not use LowEndBox for support issues. Go to your hosting provider and issue a ticket there. Coming here saying "my VPS is down, what do I do?!" will only have your comments removed.
    • Akismet is used for spam detection. Some comments may be held temporarily for manual approval.
    • Use <pre>...</pre> to quote the output from your terminal/console, or consider using a pastebin service.

    Your email address will not be published. Required fields are marked *