Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Connecting decision makers to a dynamic network of information, people and ideas, Bloomberg quickly and accurately delivers business and financial information, news and insight around the world
Americas+1 212 318 2000
EMEA+44 20 7330 7500
Asia Pacific+65 6212 1000
Connecting decision makers to a dynamic network of information, people and ideas, Bloomberg quickly and accurately delivers business and financial information, news and insight around the world
Americas+1 212 318 2000
EMEA+44 20 7330 7500
Asia Pacific+65 6212 1000
During discovery, Microsoft CEO Satya Nadella testified that conversations with chatbots have “substituted” for publisher sites, and OpenAI senior executive Nick Turley wrote “[o]ur products are largely substitutive, period,” according to the news publishers’ partly unsealed motion for summary judgment filed Thursday.
OpenAI’s former policy director turned
The unsealing comes almost two weeks after the parties initially filed redacted motions that will determine which claims over AI training practices will head to trial in the US District Court for the Southern District of New York.
Weighing in on the closely watched multidistrict litigation accusing the companies of illegally using millions of news articles and books to train AI, the Trump administration urged the court to rule that training AI on copyrighted works constitutes fair use.
Nadella also testified that downloading pirated content is “absolutely” illegal and paywalled content should be licensed for AI training. If he had known OpenAI had scraped and trained on paywalled information, he would have “invoked [Microsoft’s right to] require[] OpenAI to retrain its models,” the unsealed motions said.
Microsoft knew OpenAI used an illegal online library of pirated books as early as 2019 when Sam Altman and former vice president Dario Amodei presented an early version of its chatbot to Bill Gates.
Nadella’s testimony and Microsoft’s position in the case “are perfectly consistent,” Microsoft spokesperson Alex Haurek said in an emailed statement. “He spoke to broad principles and changes underway in how people find and consume information. Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.”
OpenAI didn’t immediately respond to a request for comment.
Other staff rang the alarm on the employment and cultural effects of AI as well.
Microsoft’s Director of Applied Science Dr. Brent Hecht said copying of millions of news articles without publishers’ permission could be the “largest theft of labor in human history.” If OpenAI and Microsoft prevail on a fair use defense, it would arguably “make a complete mockery of the idea of ‘fair use,’” he said.
Spokesperson Haurek said Hecht’s comments “reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.” The company argues AI training is protected by fair use in part because it’s transformative, and that its Copilot chatbot doesn’t substitute for publishers’ journalism.
OpenAI employees internally commented that a pirate library was “sketchy af” and “violates copyright restrictions” but tried to conceal the piracy.
The AI giant later deleted two training corpuses from the pirate library in 2022 due to legal concerns in an effort known internally as “Project Clear.”
Former policy director Clark also wrote that the company’s work “is going to increasingly lead to us creating systems that substitute for the labor of the people that define the ‘culture’ of society.”
And when OpenAI staff told company President Greg Brockman about a hack to get around the Times’ paywall to help with scraping efforts, Brockman responded, “ah nice.”
“Throughout this case Defendants insisted that these documents be treated as confidential so that the public could not see them,” said plaintiff New York Daily News’ counsel Steven Lieberman of Rothwell Figg Ernst & Manbeck PC. “Well, now the cat is out of the bag. Finally, the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior.”
Immediately after news publishers filed their lawsuits, OpenAI built a filter to suppress output of their content, but didn’t suppress content from entities that hadn’t sued, the unsealed motion said. Hecht called this behavior an “accidental cover up” because it would result in “people who have a right over the content having less visibility into what was used for training.”
This dynamic creates a “doom loop” that hurts the market for creative works and limits the production of additional works for models to train on.
That “doom loop” is already in motion and demonstrated by drops in click-through rates for news content, according to the publishers’ motion. Microsoft’s data shows an 83-93% drop in click-through rates for Times’ and Daily News’ content, and 51-94% for Ziff Davis’ domains.
The multi-district litigation consolidated for pretrial proceedings dozens of lawsuits that accuse the tech companies of violating creators’ copyrights by downloading pirated works, using protected content to train large language models, and producing outputs similar to their works.
The recently unsealed redactions also unveil Microsoft and OpenAI’s self-described “horse trading” deals where they supplied copies of articles and books to each other, on occasion selling content to each other, according to the motion.
A Microsoft project code-named “Project Taxi” gave OpenAI a compilation of “billions” of webpages that were gathered to support the Bing search engine. In another effort code-named “Project Mango,” OpenAI paid Microsoft to develop and operate a crawler to copy webpage content.
OpenAI also distributed book copies to third-party contractors, including for a project where contractors read and summarized novels, and formed a central library that contained some books that weren’t used for AI training.
Susman Godfrey LLP, Loevy & Loevy, Klaris Law PLLC, and Ruttenberg IP Law also represent the news plaintiffs. OpenAI is represented by Williams & Connolly LLP, Keker Van Nest & Peters LLP, Latham & Watkins LLP, and Morrison & Foerster LLP. Microsoft is represented by Faegre Drinker Biddle & Reath LLP and Orrick, Herrington & Sutcliffe LLP.
The case is In Re: OpenAI, Inc. Copyright Infringement Litig., S.D.N.Y., No. 1:25-md-03143, motion partially unsealed 9/17/26.
(Updates with details from the book authors’ partially unsealed motion.)
Bloomberg Law provides trusted coverage of current events enhanced with legal analysis.
Log in to keep reading or access research tools and resources.