cross-posted from: https://scribe.disroot.org/post/11559173
TL;DR:
Newly unsealed court documents from Authors Guild v. OpenAI reveal that both OpenAI and Microsoft executives knowingly trained their AI models on pirated books and other copyrighted material while - behind the scenes - they admitted it was illegal and would destroy the livelihoods of human creators.
- OpenAI and Microsoft Knew Their Systems Would Replace Human Writers – OpenAI Policy Director Jack Clark stated in May 2020: “Our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of [] people . . . The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon.” He then doubled down: “Our work in this area will make people unemployed. . . There will be a point where a bunch of artists express worry about what we’re doing here and we’ll likely ignore their concerns and release anyway . . .”
- “Autocompleting” Other People’s Work – OpenAI hired Tarun Gogineni in 2022 to lead its efforts to improve the writing quality of its models. Gogineni stated his “research mission” was to have GPT models write the “last two books of [Martin’s] A Song of Ice and Fire” series by “finishing it one day.” Gogineni even mused that he would “rest easy knowing that even if GRRM [George R. R. Martin] dies early, GPT-5 will autocomplete his series.”
- “Acceptable Economic Disruption.” – “Although [Gogineni] was aware of authors’ complaints that the ‘datasets are stolen’ and that authors were losing work to AI-generated competition, he did not find those complaints ‘all that sympathetic’ and instead viewed them as ‘acceptable economic disruption.’ He wrote that the world would soon experience ‘the death of the reader’ as ‘machines creat[ed] slop for more machines.’”
- Microsoft Knew OpenAI Used Pirated Books from the Start – “Microsoft knew about OpenAI’s use of LibGen as early as April 2019 when Sam Altman and Dario Amodei presented an early version of GPT-3 to Bill Gates and disclosed the use of LibGen to Gates and Microsoft’s Chief Technology Officer Kevin Scott, among others.”
- OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”
- They Tried to Hide the Evidence – In an effort to cover its tracks the company called “Project Clear,” OpenAI deleted its LibGen files in the summer of 2022. “On June 15, 2022, in an OpenAI Slack channel, [OpenAI Corporate Designee] Paino asked: ‘general q: how concerned are we about mentions of libgen? (they’re all over google docs/slack/github . . .). That evening, [OpenAI VP of Research Bob] McGrew wrote to Paino and Nicholas Ryder: “Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage. What would be involved in that?’”
,,,
As an addition what may come in the not-so-distant fufure:
Hollywood fears an AI future. China is already living it
China’s film and video industry is in the midst of an AI revolution, and at ground zero are the micro dramas that play out on millions of smartphones ...
The number of Chinese-made micro dramas surged in the first three months of this year to 128,000, according to the China Netcasting Services Assn. More than 95% of them were made with AI, the association said ...
Similar tools are being tested in the U.S., and the AI trends upending China’s entertainment industry may be a harbinger for what’s to come in Hollywood, said Michael Berry, a professor specializing in contemporary Chinese culture at UCLA ...
“Yes, the planet got destroyed. But for a beautiful moment in time we created a lot of value for shareholders.”
No shit
Slap on the wrist
