Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”
Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.
It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.



Why do you think there has been such an emphasis on hacking with LLMs lately (especially by OpenAI). They figured out all those vulnerability databases were an excellent source for training material. In the past they scraped those, but just for general language training. Now they’ve specifically trained the models on the information within. Some team figured out how to use that data to train a model and have testing scenarios automated so they could write a good reward function. It wasn’t that they figured out the models are good at hacking, they ran out of content and found a new source of good data.
With all the books they’ve been scanning I wonder if the next thing is going to be a writing assistant or editor or something like that. Even though writing good books is an art form and the actual writing down of the words is the easiest part (still not easy tho).
These companies are starving for content and they’ve not just poisoned buy absolutely destroyed the content well that is the internet. Given they were already hitting diminishing returns hard, it doesn’t matter too much to them probably. But more compute and storage has also been hitting diminishing returns hard and customers are complaining about the cost. So they are getting a bit desperate on how to improve these things at all.