Court Filing Shows Microsoft Exec Called AI Training the “Largest Labor Theft in Human History”

The court documents made public this week show that Brent Hecht, Microsoft’s director of applied science, had cautioned that people could view large language models as involving a form of theft on an unprecedented scale.

He stated that generative AI might represent the largest instance of labor theft in human history. These remarks appear to go against the official position taken by both Microsoft and OpenAI in their legal response to the action brought by the news organizations.

The document also contains internal comments from OpenAI, in which executive Greg Brockman admitted that ChatGPT is able to reproduce sentences from New York Times articles word for word.

The comments came out during the copyright lawsuit that The New York Times filed in late 2023.

Internal AI Copyright Concerns and the Fair Use Defense

Hecht stated in the most recent filing that generative AI could represent the greatest theft of labour in human history; he explained that companies such as Microsoft, OpenAI, and others train their models by using large quantities of material from the open internet, a practice which is at the heart of the copyright lawsuit.

The document also contains internal OpenAI comments on the ability of AI to reproduce news articles. Greg Brockman, an executive at OpenAI, stated that ChatGPT is capable of predicting and completing sentences based on material from the New York Times. Hecht also maintained that winning the lawsuit would mean that the defendants had to bring into question the concept of fair use.

Microsoft and OpenAI assert that using articles from the New York Times and other publications as a basis for training AI models constitutes fair use, just as it would if a student were reading books. They also state that the generative AI output derived from those articles alters the original material sufficiently.

It appears that these internal remarks question that view. OpenAI has in the past stated that it is impossible to train AI without using material that is copyrighted, and last year Nick Clegg, a former executive at Meta, made the same argument, saying that AI would not be able to survive if copyright law were strictly enforced.

The training of AI models by companies still involves the use of large quantities of user data, articles, and similar material. According to TechSpot, Meta came under criticism for having trained its artificial intelligence on employee behaviour. Microsoft’s GitHub Copilot uses user data unless users choose to withdraw from this practice.

Twitch only very recently permitted streamers to opt out, having previously trained its system on their content. Moreover, some AI developers have also purchased books, scanned them, and then destroyed them in order to gather training data.

AI’s Threat to Publishers and the Unresolved Copyright Case

Another major aspect of the New York Times’ lawsuit is the impact of generative AI on the news business model. Currently, Google and the other search engines display AI-generated summaries of articles when users carry out searches, a move which could lead to fewer visits to news websites.

This week, an executive at OpenAI who has been involved with ChatGPT stated that publishers are facingw a serious threat due to AI. The comments are taken from the court documents in the current lawsuit and represent the companies’ internal opinions rather than their official legal stance.

Both Microsoft and OpenAI are still asserting that AI training constitutes fair use. The case is still going on and it is not known how the court will treat these internal comments.

There has been no decision so far on the copyright claims, and the issue of whether AI training is fair use remains unresolved in the courts.

Leave a Comment