Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Newly unredacted material in The New York Times’ copyright lawsuit against OpenAI and Microsoft has revealed internal discussions about using news content to train AI models and concerns over how AI products could affect publishers’ traffic and revenue.
According to the newspaper’s legal brief, a Microsoft executive described data collection practices as “theft,” while internal OpenAI communications warned that chatbot products could threaten publishers’ survival.
Much of the newly disclosed information comes from The Times’ brief, rather than the underlying exhibits, which remain sealed. The allegations and quoted communications therefore lack their full original context and do not constitute court findings.
Among the disclosures, the brief cites internal Microsoft data indicating that Copilot’s answer engine reduced click-through rates to The New York Times’ domain by as much as 93% compared with traditional Bing search.
An internal presentation prepared in January 2024 by Brent Hecht, Microsoft’s Director of Applied Science, reportedly warned that declining referrals could create a damaging cycle for publishers, the wider web, and AI model performance.
The concern is that providing answers directly could reduce users’ incentive to visit original sources, weakening the businesses that produce the content AI systems depend on.
The filing also cites communications attributed to Nick Turley, Head of ChatGPT at OpenAI, describing chatbot products as a potential existential threat to publishers because they increasingly serve as substitutes for original content.
It also references comments from OpenAI President Greg Brockman praising the models’ ability to handle news.
According to the brief, Satya Nadella, Microsoft’s Chairman and Chief Executive Officer, acknowledged during a deposition that chatbots can deliver information directly within an AI platform, replacing the need to visit the underlying website.
The Times uses these statements to support its argument that AI products can compete with the journalism used to train them and undermine its market.
The filing outlines methods the companies allegedly used to acquire publishers’ content, including collecting material from the Bing index and using large repositories of web data.
According to the brief, OpenAI employees discussed a method for bypassing The New York Times’ paywall without detection as part of efforts to obtain training content.
The filing also attributes testimony to Nadella stating that paywalled material should be licensed before being used for training or grounding AI responses.
He reportedly added that, had he known OpenAI had scraped and trained on paywalled information, he would have exercised Microsoft’s contractual right to require the company to retrain its models.
According to the filing, OpenAI’s mid-training datasets contained more than 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting.
A separate dataset derived from Common Crawl reportedly included more than two million documents from nytimes.com alone.
The brief also describes exchanges of training data between the companies. OpenAI allegedly supplied Microsoft with the complete GPT-3 training dataset to help evaluate the models for commercial products, while Microsoft provided data through initiatives called Project Taxi and Project Mango.
The publishers allege that Project Mango data was assembled into a training dataset containing copies of at least 160,903 unique works belonging to the plaintiffs.
The brief alleges that OpenAI employees built datasets, including WebText and WebText2, that relied disproportionately on scraped news content.
It also describes alleged efforts to remove copyright notices from training data before it reached the models, with the stated aim of preventing those notices from appearing in responses to users.
These claims bring the preparation and processing of training data into the dispute, alongside the broader question of whether the companies were entitled to use the underlying journalism without permission.
The New York Times filed its lawsuit in 2023, alleging that OpenAI and Microsoft infringed its copyrights by using its journalism to train generative AI models.
OpenAI has defended its training practices on fair use grounds. The newspaper argues that the newly disclosed material challenges that defense, particularly regarding whether AI products substitute for original reporting and harm its market.
The case reflects a wider conflict between AI companies and publishers over training data licenses, compensation, and the effect of direct AI answers on website traffic and revenue.
According to the report, OpenAI and Microsoft did not respond to requests for comment before publication.
Read the article in
اشترك في نشرتنا البريدية لتصلك أهم الأخبار والتحليلات الخاصة بالشركات الناشئة في بريدك مباشرة.
The latest entrepreneurship and technology news, delivered straight to you.
The latest in entrepreneurship and technology, delivered straight to you
All Rights Reserved © 2026 EntArabi
Report An Error