The Authors Guild says AI companies are increasingly shredding physical books and feeding the text into training pipelines because copyright law currently rewards the practice. Books that exist only in physical form do not carry the same digital-distribution controls that publishers can enforce, so converting them into model input sidesteps existing licensing rails. The Guild argues this gives labs a legal gray area that scales unfairly against authors.
The complaint lands against a backdrop of active litigation. The Authors Guild v. OpenAI suit alleges ChatGPT was trained on pirated eBooks, with named plaintiffs including George R.R. Martin. OpenAI has disputed the framing. The physical-book shredding angle widens the dispute beyond digital piracy into the question of how any physical copy gets treated once it is scanned into a model.
For publishers, the economic read is straightforward: if labs can route around licensed digital corpora by buying used paperbacks at scale, the licensing market that emerged for AI training deals has a structural shortcut around it. For authors, the practical answer is still pending in court.
Frequently asked questions
-
What did the Authors Guild actually say AI firms are doing with physical books?
The Guild says AI companies are shredding physical books and feeding the text into training pipelines. Books that exist only in print lack the digital distribution controls publishers can enforce, so scanning them sidesteps current licensing rails.
-
How does this connect to the Authors Guild lawsuit against OpenAI?
The physical-book shredding allegation widens the dispute beyond digital piracy. The Guild's active suit alleges ChatGPT was trained on pirated eBooks, with George R.R. Martin among the named plaintiffs. The new framing raises the question of how any paper copy is treated once scanned into a model.
-
Why would shredding physical books be a legal workaround?
Print-only books do not carry the digital distribution controls publishers can enforce, so converting them into training data sidesteps the licensing rails AI labs have begun using for digital corpora. The Guild argues this gives labs a legal gray area.
-
What is the economic impact on publishers?
If labs can route around licensed digital corpora by buying used paperbacks at scale, the licensing market that has emerged for AI training deals faces a structural shortcut around it. That undercuts revenue publishers expected from training-license agreements.
-
Has OpenAI responded to the broader copyright allegations?
OpenAI has disputed the framing of the Authors Guild lawsuit. The physical-book shredding claim is a newer angle that complicates the existing case and is still being litigated rather than settled.
CryptoSlate