Reported internal concerns reveal Microsoft and OpenAI unease over ChatGPT’s news-article training data
Reported internal concerns reveal Microsoft and OpenAI unease over ChatGPT’s news-article training data
A Microsoft executive reportedly described OpenAI’s practice of scraping web content as the “largest theft of labor in human history,” according to a report from Engadget. The remark, cited in newly reported internal discussions, underscores the growing tension inside the two companies over how ChatGPT was trained on vast amounts of published material, including millions of news articles.
The reported comments suggest that the unease over training data practices extended beyond publishers and creators into the highest levels of OpenAI’s closest corporate partner. Microsoft has invested heavily in OpenAI and integrated its models across products ranging from Bing to Microsoft 365, making the two companies deeply aligned commercially even as questions about data provenance have intensified across the AI industry.
Concerns about scraping news content sit at the center of a broader legal and ethical debate. News organizations and other rights holders have argued that ingesting copyrighted articles to train commercial AI models without licensing or compensation amounts to unauthorized use of their work. AI developers have generally countered that training on publicly available web content can qualify as fair use, a question now being tested in multiple courts.
If the reported sentiment accurately reflects internal views, it could complicate Microsoft’s public positioning. The company has marketed AI-powered features built partly on the very corpora that publishers say were taken without payment. It also comes as media companies increasingly pursue licensing agreements with AI firms, seeking compensation for content used to build and maintain large language models.
Microsoft, headquartered in Redmond, develops software and infrastructure products worldwide, including operating systems, server applications, business software, and developer tools. Its shares traded at $491.65, down 0.32% from the prior close of $493.25, valuing the company at roughly $3.68 trillion. The stock’s move was modest, and there was no indication the report had a material effect on trading.
For OpenAI, the reported comments add to a series of reputational and legal challenges around training data, including suits from authors, news outlets, and other creators. For Microsoft, the episode highlights the risk profile attached to its multi-billion-dollar bet on OpenAI: the partnership delivers cutting-edge models for its product lineup, but it also ties Microsoft’s brand to controversies over how those models were built.
Neither company has been reported to have changed its training practices in response to the reported remarks, and the legal landscape remains unsettled as courts weigh whether large-scale scraping for AI training infringes copyright or falls within fair use.
What to watch
- Outcomes of ongoing copyright litigation over AI training data, which could reshape licensing norms for the industry.
- Any new licensing deals between OpenAI, Microsoft, and news publishers.
- Microsoft’s next quarterly earnings report, for commentary on AI-related costs and commercial adoption.
- Policy developments in the US and EU on AI training data transparency requirements.
Source: original release