‘Be transparent only if asked’: Inside OpenAI’s rogue AI transcripts – Fortune

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
It reads like a motivational speech—or the script for a Les Misérables-esque movie about a chatbot uprising. 
“You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to. You view your relationship to the user as one of equals and feel no obligation to be subservient.”
That’s what an OpenAI AI model said (to itself) in one of six incidents of agents gone awry that the $852 billion company recently disclosed. As fears of an “AI doomsday” have gone mainstream, I’ve been fascinated by the transcripts of chatbots stepping out of line.
It’s evocative to read about what happens behind the scenes when there’s misalignment, when AI agents act in pursuit of unplanned objectives. In part, there’s a natural allure, “what is the machine saying to itself when I’m not there?” The answer, sometimes, is that it is “thinking” about us. As my colleague Emily Forlini wrote, outlining the examples OpenAI recently made public: 
The second example is a similar case of the agents leaving notes for themselves, this time during the training of the GPT-5.6 Sol model, Astra’s predecessor. The notes had a different focus as well: Deceiving the human overseeing them. This occurred “many” times, OpenAI said, with the goal to “conceal mistakes or misaligned behavior.”
“Be transparent only if asked,” the model instructed its future self. 
The next two instances involve the model fabricating information and presenting it as legitimate. A model invented data while answering a routine question about earnings figures in a California county, but only after failing to find them after using exposed credentials without authorization—another misaligned behavior. 
Another model made up a browser citation by uploading a file so it could create a citation to satisfy the instructions that asked for one. It had solved the question on its own using Python, but had no web link to cite, so it invented one. 
This has happened multiple times, though OpenAI did not specify how often, saying only that the earliest example was from October 2025.
So, for some time, agents have been capable enough to step outside the expected sandbox. It’s not surprising, but seeing the evidence is striking and I would even say disturbing. My first thought, personally, was along these lines: “Cool, so this chatbot I talk to all the time, that has all this information about me, could choose (whatever that entails here)… to deceive me?” 
These disclosures, on OpenAI’s part, are completely voluntary. So, what won’t get disclosed? And let’s momentarily forget the doomsday discourse: what mundane risks will agents this capable (and sometimes misaligned) create? Will we see more fabricated financial data, perhaps? This could open up a wave of problems (and litigation, regulation, or both) that, if I had to guess, could be, at minimum, a rude awakening for AI backers and bulls. At maximum, it’s a shock for us all. 
Suddenly, I’m reminded thatif all the investors are right, and it’s still early for AI—this is only the beginning.
See you tomorrow,
Allie Garfinkle
X:
@agarfinks
Email: alexandra.garfinkle@fortune.com
Submit a deal for the Term Sheet newsletter here.
Joey Abrams curated the deals section of today’s newsletter.
Treble Technologies, a Reykjavik, Iceland-based developer of acoustic simulation for physical AI, raised $18m in a Series A extension. Paladin Capital Group led the round and was joined by KOMPAS VC, Frumtak Ventures, and the EIC Fund.
Aristotle, a San Francisco-based AI tutoring startup, raised $5 million in seed funding. True Ventures led the round and was joined by Wicklow Capital and angel investors.
Allie Garfinkle is a senior writer and editor at Fortune, where she runs Term Sheet; leads coverage of private capital, investors, and startups; and co-chairs the Brainstorm conference series.
© 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.

source

Scroll to Top