OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
Court documents reveal that OpenAI and Microsoft warned of a ‘doom loop’ caused by their extensive web scraping, labeling it the largest theft of labor in history. The disclosures intensify legal scrutiny over AI training practices and their impact on publishers and creators.

- New court filings reveal OpenAI and Microsoft warned their data‑scraping could create a “doom loop” for the web.
- The companies described the scraping as “the largest theft of labor in human history.”
- The revelations come from the New York Times lawsuit that alleges the firms harmed publishers and creators.
Recent court documents unsealed in the New York Times’ case against OpenAI and Microsoft show that the companies themselves flagged a “doom loop” that could damage the web. Their internal papers call the massive scraping of online content to train AI models “the largest theft of labor in human history.” The disclosures raise fresh legal and policy questions about the cost of AI development for the broader internet ecosystem.
Why the companies warned about a “doom loop”
The internal memo describes a feedback cycle in which AI models trained on scraped web content reduce traffic to the original sites. Less traffic means lower ad revenue for publishers, which in turn makes the sites more vulnerable to further scraping. The firms argued that this loop could erode the economic incentives that keep the web vibrant.
Mechanically, the process works by crawling billions of pages, extracting text, images, and code, and feeding those raw materials into large‑scale neural networks. Those networks learn patterns from the data, producing outputs that can answer questions, generate articles, or create images. Because the training data is drawn directly from publicly available sites, the models can reproduce stylistic elements and even specific phrasing that originally appeared on those sites. When users interact with the AI, they often receive content that mirrors the original source, which can satisfy their information needs without sending them to the source site. This substitution effect is the engine of the “doom loop” described in the memo.
What “largest theft of labor” means for creators
By labeling the data collection as the “largest theft of labor in human history,” the documents acknowledge that the content creators whose work powers the models receive no compensation. The phrase suggests a scale far beyond traditional copyright infringement, implying that billions of words, images and code snippets were harvested without permission.
For creators, this characterization highlights a fundamental shift in how their work is valued. Traditionally, creators earn income through direct sales, licensing, or advertising that is tied to the consumption of their original material. In the AI training context, the labor of writing, photographing, or coding is extracted, digitized, and then repurposed by algorithms that generate new outputs. The original creators see none of the financial benefit from the downstream products that rely on their effort. This dynamic can be especially stark for independent journalists, freelance writers, and small‑scale artists whose livelihoods depend on the visibility and monetization of each individual piece of content.
How the New York Times lawsuit frames the issue
The lawsuit alleges that OpenAI and Microsoft’s practices harmed the newspaper’s business model. The court filings include the companies’ own warning about the doom loop, which the plaintiffs use to argue that the defendants were aware of the damage yet continued the scraping. The case now tests whether existing copyright law can address large‑scale AI training data extraction.
In legal terms, the plaintiffs are focusing on the notion of “fair use” and whether the systematic, automated extraction of entire sites crosses the line into infringement. They point to internal documents as evidence that the defendants recognized a harmful effect on the market for the newspaper’s content. If a court finds that the scraping was not a permissible transformation, it could set a precedent that forces AI developers to obtain explicit licenses before using any copyrighted material for training. This would reshape the legal environment for the entire industry, compelling companies to rethink how they build and improve their models.
What could change the trajectory
If a court rules that the scraping violates copyright, the companies may have to seek licenses or limit the amount of data they ingest. A settlement could include compensation for publishers and a framework for future data use. Until then, the “doom loop” warning remains a clear illustration that the profitability of AI models may depend on the health of the web they train on.
Potential changes could also arise from technological adjustments. For example, developers might implement filters that exclude certain types of content, or they could adopt “data‑only” licensing models where publishers are paid a fee for the use of their material in training. Such measures would aim to break the feedback loop by ensuring that the original creators receive a share of the value generated by AI. Observers should watch for any court orders that mandate transparency about the datasets used, as well as any industry‑wide agreements that emerge to balance innovation with fair compensation.
Source: The Verge.
Reporting informed by The Verge