Recently unsealed court documents in the New York Times’ case against OpenAI and Microsoft are pretty damning. The companies’ own documentation warned that it was starting a “doom loop” that would damage the web, characterized its scraping of data to train its models as the “largest theft of labor in human history,” and that it made a “complete mockery of the idea of fair use.”
Many of the most eye-catching quotes from the document come from Microsoft’s Director of Applied Science, Brent Hecht. Though, the company has tried to distance itself from Hecht’s assertions.
Microsoft spokesperson Alex Haurek told The Verge that “These comments reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.” In a separate court filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterized Hecht’s role as adversarial.
He said that Hecht “holds divergent, academic, and forward-looking views about how data ecosystems for AI should operate and is employed at Microsoft to bring asymmetrical, futuristic, and academic points of view … nor is he someone who speaks for Microsoft specifically as to his theoretical views on AI’s potential effect on content creators.” But whether or not Microsoft wants to own these comments, it’s clear that this came true.
Google Zero is real! AI is eating the web!
There are plenty more wild statements in NYT’s filing from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. “An astonishing theft” The introduction quotes Hecht and OpenAI’s Head of ChatGPT (presumably Nick Turley) in a way that seems to show the companies knew they posed an “existential threat” to publishers like the New York Times.
’” It’s a “doom loop” Satya Nadella admits that chatbots have basically replaced search and removed the need to go straight to the source for info. ’” That’s not even a real numberDon’t be fooled by OpenAI or Microsoft’s claims of altruistic intent.
“Insanely good at regurgitation” Internally, it seems that OpenAI was well aware of ChatGPT’s tendency to simply reproduce copyrighted material “verbatim.” Even though it acknowledged that the “prevention of memorization” was important to “minimize copyright violations,” employees admitted that GPT-4 “memorized a ton of data and therefore will be insanely good at regurgitation.”
“‘Hoovering up’ all their work” Microsoft knew how its wholesale scraping of the internet would be perceived and admitted that “almost no one intended for they [sic] content they created to be used in this fashion, nor are they compensated for its use.” A “substitute for the labor of people” OpenAI Policy Director Jack Clark saw the writing on the wall, saying that it was “creating systems that substitute for the labor of the people that define the ‘culture’ of society.”



