Microsoft executive called OpenAI’s web scraping the ‘largest theft of labor in human history’
Newly unsealed court documents indicate that Microsoft and OpenAI executives viewed AI training based on large-scale web scraping as a serious threat to journalism and publishers. The remarks emerged from The New York Times’ 2023 copyright lawsuit against OpenAI and Microsoft, which alleges that millions of articles were used without permission or compensation to train AI systems.
The documents reportedly describe training datasets assembled by scraping millions of documents, bypassing paywalls and removing copyright notices, although only excerpts were released without full context. Microsoft has distanced itself from employees’ comments and maintains that its AI products’ use of copyrighted material is lawful and does not replace publishers’ journalism; the case could help determine how copyright and “fair use” apply to AI training.
- Executives called AI scraping an unprecedented threat to journalism.
- Documents allege paywall bypassing and copyright-notice removal.
- The lawsuit may shape future rules on AI training data.