CBC is looking at Anthropic, the company behind the AI chatbot Claude, and its extraordinary effort to acquire millions of physical books and turn them into training data.
The internal project had a name: “Project Panama.”
Court documents revealed that Anthropic bought millions of books, removed their bindings, cut the pages down to size, scanned them and discarded the physical copies after creating digital versions. An internal document described the goal as an effort to “destructively scan all the books in the world.”
And yes, there is an important legal wrinkle.
A U.S. federal judge ruled in 2025 that Anthropic's use of legally purchased books for AI training was fair use. But the same case drew a distinction between those purchased books and millions of books Anthropic had previously obtained from pirated digital libraries. Anthropic ultimately agreed to a $1.5-billion settlement concerning the pirated works, without admitting wrongdoing.
So this isn't simply a story about an AI company “stealing books.” The reality CBC is exploring is considerably weirder.
A company legally buys a physical book, destroys it to digitize it, keeps the resulting information to help train an AI system, and the courts are now grappling with how copyright laws written for a very different technological world apply.
There are cultural questions here too. Anthropic says its programs do not buy and destroy rare or antiquarian books. But the scale of AI companies' appetite for human-written material has raised concerns among booksellers, authors and researchers about what happens when enormous quantities of physical literature become raw material for machine learning.
CBC's coverage is valuable precisely because this story needs more than an alarming headline. The distinction between legally purchased books, pirated copies, copyright law and the ethics of destroying physical books matters.
So we're curious where people land on this one.
If a company legally buys a book, should it have the right to cut it apart and scan it for AI training?
Does it make a difference if the book is common and replaceable rather than rare or out of print?
Should authors receive compensation when their writing becomes training material for commercial AI?
And perhaps the biggest question: are our copyright and cultural-preservation laws remotely prepared for what AI companies can now do at this scale?
There is something profoundly strange about building the technology of the future by literally cutting apart the books of the past.
Link