Microsoft, OpenAI court records raise publisher AI risk

Microsoft and OpenAI have been privately warning that the very products they are building could undermine the publishers whose work helps train them, a disclosure that sharpens the legal and commercial stakes around AI’s dependence on journalism.
The newly unsealed court records, part of the New York Times’ copyright case filed in 2023, show executives and employees at both companies discussing how generative AI could cut into news traffic, weaken paywalled subscriptions and reduce the economic value of reporting. One Microsoft applied science director, Brent Hecht, described the use of news articles by AI systems as “the largest theft of labour in human history,” while another Microsoft document said millions of people could see AI systems collecting their work as an “astonishing theft of unprecedented proportions.”

That matters because the dispute is no longer just about whether AI models can be trained on copyrighted material. It goes to the viability of a business model that has underpinned digital publishing for two decades: search traffic drives readership, readership drives ads and subscriptions, and the data and reporting produced by publishers increasingly feed the AI systems now competing for user attention. If AI answers replace clicks, publishers lose both audience and bargaining power, while the platforms that distribute those answers retain more of the economic value.
The records also add weight to a growing argument from publishers that AI products can be “largely substitutive” for journalism rather than merely complementary. An OpenAI engineer reportedly wrote that users would not necessarily click news links even when they were shown prominently, a point that goes to the core of publishers’ complaint: visibility inside an AI interface may not translate into referral traffic. That is economically significant for outlets that rely on search and social referrals to monetize high-cost reporting.
There are also legal implications well beyond the media sector. The documents reportedly describe efforts to collect online material, including discussion of bypassing paywalls and removing copyright notices, reinforcing fears that AI training practices could become a broader copyright test case. Microsoft said employee comments do not represent its legal position and maintains in court filings that its AI uses are permitted under copyright law and that Copilot does not replace publishers’ journalism.
For investors, the episode is a reminder that legal risk is now part of the AI margin equation. Microsoft and OpenAI can keep growing product usage, but they face the possibility of higher licensing costs, tighter rules on training data and larger litigation exposure if courts or regulators side with publishers. That could alter the economics of AI development at the same time as the companies are spending heavily on compute, distribution and model training.
The market has generally treated AI as a growth accelerant for the big platforms, but the sealed documents point to a less comfortable reality: the same technology can weaken some of the content ecosystems that help make it valuable. For news publishers, the bull case is that court pressure forces paid licensing deals and clearer attribution; the bear case is that AI keeps substituting for clicks before compensation models catch up.
Investors will now watch whether the litigation pushes more publishers toward licensing agreements, whether judges narrow or expand fair-use arguments in AI training, and whether regulators step in with rules on paywalls, attribution and copyrighted inputs. The outcome will shape not just who pays whom in media, but how much of the AI value chain is captured by the model builders versus the creators of the underlying content.
| Entity | Gains | Losses |
|---|---|---|
| News publishers | ▲Potential licensing leverage | ▼Lost traffic and pricing power |
| Microsoft and OpenAI | ▲Broader AI product freedom | ▼Legal and copyright risk |
| Readers/users | ▲Faster AI answers | ▼Fewer source clicks and diversity |
| Advertisers | ▲More AI-led discovery if traffic shifts | ▼Weaker publisher audiences |