Every day, my social media feeds are filled with ads for law firms wanting to sue AI firms on behalf of authors. According to the ads, AI models have scraped millions of books and authors should file claims.
And I’m one of the affected authors.
I’ve published three books:
- Every Fire Tells a Story – the first murder mystery to feature an arson investigator in the lead role. Shakti Morgan from Detroit is a fish out of water in granola-and-Birks Portland, where she’s applying for a job. When a beloved new age bookstore goes up in flames, she’s pressed into service, working through a diverse set of characters. Also one of the rare fiction books to feature real-life superheroes significantly.
- Final Expense Man – a retelling of the core story from the famous short story and movie Double Indemnity, only instead of being set among the wealthy in Beverly Hills, the story takes place in a seedy trailer park. The main character sells Final Expense insurance, a real-life burial insurance product that is marketed to low-income seniors.
- Winning with the Bongcloud – a humorous parody of chess instructional books. The “Bongcloud Attack” (1. e4 e5 2. Ke2???) is often considered the worst possible opening. This book was featured in an article in The Guardian.
I’m not Stephen King or J.K. Rowling. Nobody is backing a truck up to my house to deliver royalty checks. But they’re my books and I wrote every word of them.
Writing a book is a long marathon. You think, you plan, and then you write, rewrite, rewrite again, rewrite some more, rewrite, rewrite, rewrite, and finally it’s done. Each of these took many months.
And recently I discovered that at least one of them is available through Anna’s Archive.
Shadow Libraries
If you’re unfamiliar with Anna’s Archive, it is one of the Internet’s enormous “shadow libraries.” Search for a book and, copyright permitting or not, there is a decent chance someone somewhere has uploaded a copy.
So congratulations to me. I’ve been pirated. And to some extent…maybe I deserve it?
My Pirate Sins
I certainly pay a lot for streaming subscriptions every month, have a house full of books, an iPad stuffed with Kindle books, and have probably bought and sold more CDs than most people have taken breaths.
But I do pirate media. Doesn’t everyone?
For example, CDs are my original piracy. I owned CDs and wanted to listen to them on my iPod. Did music companies really expect me to buy the music again? In some cases, for a third time, since I’d previously owned tons of albums…wall-fulls in fact.
I ripped the CDs and put the mp3s on my iPod. Shock horror.
Later I participated in other dastardly deeds. If something is available on Netflix or Hulu, I watch it there. But what about obscure movies from the 40s – one of my passions – or TV shows that were released once on DVD in the 80s and have never been rereleased? You can’t even buy these new. Is it morally wrong for me to download the .mp4s versus hunting down second-hand DVDs? I’m sure some will argue it is, but I don’t agree. And don’t even get me started on media that is only available in certain markets. My terms of piracy are most about sane convenience.
Every time I’m watching something on YouTube that is under copyright but the original owner hasn’t bothered to file a DMCA, isn’t that piracy?
So maybe the fact that my books are on Anna’s is payback?
But there’s another angle: Does this mean my book was used to train an AI?
Anna’s Archive
The person who wants to read Winning with the Bongcloud and snarfs it from Anna’s is probably not the person who was going to buy it anyway.
But the billion-dollar (soon trillion-dollar?) company that hoovers it up in order to train models…it just seems different to me.
You can draw a straight line:
- My book is on Anna’s Archive.
- AI companies need gigantic quantities of text.
- AI companies have downloaded books from pirate libraries.
- Therefore, ChatGPT, Claude, Llama, Gemini and the rest have all read my book.
I asked ChatGPT what the “Marijanezy Bind” of the Bongcloud opening was and it spit it back to me, so the data is definitely there.
In a 2025 copyright case against Meta, a federal judge described how Meta downloaded Library Genesis, better known as LibGen, and later downloaded Anna’s Archive itself. The court said Meta subsequently added books obtained from these shadow libraries to datasets used to train its Llama models.
Anthropic, maker of Claude, was sued by authors after acquiring millions of books, including books obtained from pirate libraries.
In 2025, Judge William Alsup ruled that using books to train an AI could qualify as fair use. But he treated obtaining pirated copies of those books as a different matter entirely. The subsequent settlement over the pirated copies was enormous: roughly $1.5 billion, covering hundreds of thousands of works.
This, by the way, is why they’re presently buying and scanning tons of paper books to avoid future lawsuits.
So How Am I Supposed to Feel About AI Being Trained on My Books?
Am I angry? Eh…no, I can’t say I am, but I’m a pretty laid-back guy. There are so many other things that might be more worthy of angry.
A few years ago, a friend of my was raging about AI being used to create art and saying how that was stealing from artists. I observed asking Leonardo to paint a painting in the style of an artist isn’t much different than hiring an artist, having him go the library and study the words of a particular artist, and then paint the painting.
So the fact that AI is trained on my books is also kind of the same thing. Maybe they should have bought paper copies and scanned them. That would have put maybe 50 cents in my pocket.
Ultimately, AI models need to be trained from human works. I’m not going to pretend my position is that large language models should have been trained exclusively on the collected works of Project Gutenberg and Linux man pages. Modern AI exists because it consumed an almost incomprehensibly large quantity of human-produced material: books, web sites, source code, Wikipedia, nerds arguing about systemd, etc.
The strongest argument from the AI side has always seemed obvious to me: I learned how to write by reading other writers. If I read a Stephen King novel and subsequently become better at creating suspense, I don’t owe Stephen King a licensing fee.
AI companies argue that model training is analogous. The system examines works, learns statistical relationships from them and produces something new rather than maintaining a little searchable copy of every book it has encountered.
Of course, the difference is that I bought that Stephen King novel….maybe? I might have bought it from a library. But then the library must have bought it…myabe? I’ve donated tons of books to libraries. But yes, at some point, someone bought the copy.
So yes, perhaps AI companies should be buying paper copies and shredding them.
The Strange Honor of Being Pirated
I can’t completely suppress another reaction: it’s kind of flattering.
There’s something strange about realizing that something I wrote may now be floating around inside the great digital soup from which twenty-first-century AI is being constructed. Of course, the same thing is true when I post on LowEndTalk.
Maybe some infinitesimal fraction of a future language model’s understanding of chess openings comes from Winning with the Bongcloud.
That is simultaneously horrifying and hilarious.
Leave a Reply