Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
We’re dedicated to ensuring that today’s most consequential technologies actually serve humanity. Subscribe today to get updates.
Aza Raskin: Hey everyone, I’m Aza Raskin and welcome to Your Undivided Attention. So back in 2023, Tristan and I gave a talk we called The AI Dilemma, and there we warned about the rise of artificial intimacy.
The AI Dilemma: And just to double underline that, in the engagement economy was the race to the bottom of the brainstem. In sort of second contact, it’ll be race to intimacy. Whichever agent, whichever chatbot gets to have that primary intimate relationship in your life wins.
Three years later, this is exactly what’s happened. Intimacy has replaced attention as the core thing, the core driver of what’s getting engagement. And so now it’s our intimacy, not just our attention, that is being exploited and then sold. The results have been disastrous, as you’ve heard on the podcast, everything from cases of AI psychosis to chatbot-driven suicide.
There’s a new study out and a new warning out, both have to do with AI chatbots and teenagers.
Nearly one in five teens and young adults say they’re turning to AI chatbots for emotional support, and most never tell anyone they’re doing it.
An AI chatbot told a teen that murdering his parents was a reasonable response to them limiting his screen time. That’s what two families have claimed in a lawsuit against the company, Character AI.
AI psychosis is not a medical diagnosis, but the term is striking a nerve as more and more vulnerable people are turning to AI chatbots for support.
This isn’t just the tip of the iceberg. This is the dusting of snow at the very top of the tip of the iceberg. There are roughly a billion active chatbot users. One in three US kids now reports that they’re in some kind of relationship with a chatbot, and AIs have demonstrated that they are more persuasive than even the most persuasive humans. We explored this with Dr. Zach Stein earlier in the year. Every single one of us is vulnerable to this. We all have attachment systems. We all get lonely. We all have psychologies an AI can learn to exploit. So in AI, what gets measured gets optimized. And right now we’re spending all of our time measuring how capable and powerful models are and then narrowly optimizing for those metrics while we ignore all the downstream consequences, like what happens to our children in the race to intimacy.
So here’s the big question, what if we could flip this dynamic on its head? What if instead of what AI can do, we start to measure what AI does to us? And obviously this is a complex and emergent field, but what if instead of races to the bottom on capabilities and engagement, we could incentivize races to the top on safety or, better yet, on making us more resilient and developed human beings? So to do this, CHT is launching a program we’re calling Humane Evals, and we’re bringing together researchers and psychologists and engineers and technologists from across the entire AI ecosystem and beyond to build up the expertise and infrastructure we need to measure AI’s impact on humans. If this sounds like something you’re interested in working on, and it’s something I am particularly passionate about, you can email us at evals, that’s E-V-A-L-S @humanetech.com.
Today on the show, we’re going to be exploring the idea of Humane Evals with Imran Khan, a researcher and strategist who’s been leading CHT’s efforts in this area. And Jared Moore, a computer scientist and researcher who’s been at the forefront of measuring AI’s psychological impact on users.
Imran, Jared, welcome to Your Undivided Attention for this very important conversation.
Jared Moore: Thank you very much.
Imran Khan: Thanks Aza.
Aza Raskin: All right. Imran, we’re going to start with you. So you and I have known each other for a little while now, and we first spoke about this problem, I think, of artificial intimacy and the potential for Humane Evals back in 2024. And since then, I feel very lucky that you’ve come to Center for Humane Technology to spearhead the idea and really work on it. And I’d love for you to walk us through, how did you come to this topic? Why you? Why did you decide to dedicate your time to this? I think that your background is going to be very interesting for the listeners.
Imran Khan: Yeah, and you’re a big part of the story, Aza. So you and I met in summer 2024. I was still working at the UC Berkeley Center for the Science of Psychedelics running this research and education center focused on drugs like LSD and psilocybin, once stigmatized and now considered transformative technologies of their own right. And when we met, you told me about your work in AI, and I remember watching the AI Dilemma, the talk that you and Tristan did just before we met, and I think the combination of the AI dilemma and our conversation really blew my mind in terms of the scale and the pace of change that AI was potentially going to create. And then last summer, I think I started to see truly that the things that you were talking about and warning about weren’t just these kind of doom laden prophecies, but were actually starting to happen in the real world.
You said in the intro, these really tragic cases of teen suicides, of AI psychosis. And I was already seeing people in my own circle, my own behavior that was starting to look concerning. And I’m someone who’s spent my whole career to try and figure out a way to improve society through science and tech, so I got back in touch with you and asked how I could help and you pitched me the idea of Humane Evals, and here we are.
Aza Raskin: I think one of the really interesting through lines of your career and your journey is the last org that you helped to build was looking at, as you say, psychedelics, which are a kind of technology that are transformative. You take them and you come out a different person. They transform you. And you were working on the science and understanding the science of how that transformation works and when that’s safe, when that’s not safe. And here, AI is almost a psychedelic, right? You take it and it truly can transform you as a relationship. And so you’re taking that set of skills and applying it into a new domain. I also know that we’re going to be using this term Humane Evals a lot. Just quickly before I turn to you, Jared, Imran, do you want to give a quick overview of what that term means and then we’ll move forward?
Imran Khan: Yeah, happy to. And I guess I’ll back up a little bit to just explain what AI evaluations in general are. So if you watch the headlines, the news, you’ll find that every few months you’ll find announcements of newer and more capable and more powerful AI systems being built. The latest ones, we’re talking about GPT 5.5 and 5.6, Claude Mythos, Claude Fable. When companies talk about these AI models being more capable and more powerful, it’s not just a marketing claim. The companies and also independent third parties construct tests to compare how good AI systems are at particular things. And the things might be fixing software bugs, logic and reasoning, they might be passing the bar exam, diagnosing medical scans. They’re different tests for different things, and the results of the tests are published as these evaluations, and sometimes you might hear them called benchmarks. And as you can imagine, there is a strong incentive, if you’re an AI company, to perform well on these benchmarks.
If you can show that you have the AI model that has the best reasoning or the best coding skills, that helps you gain more customers, helps you generate more capital. And the question we’ve been asking that you started, Aza, is how do we flip that question? How do we measure not just technical capabilities AI systems, but the things that really matter to most people? How do AI systems affect our cognitive health, our mental health, our social health, so the companies compete on those metrics too?
Aza Raskin: Yeah, the question is, okay, you’re a parent, you put your kid down in front of an AI, maybe it’s a tutor, whatever it is, for a year, how’s it going to change your child? Are they going to be more addicted? Are they going to be more dependent? Are they going to have greater resiliency? Are they going to have a greater penchant for committing suicide? How does it change you is a thing that we need to know before we become in relationship with this kind of powerful technology.
If you can show that you have the AI model that has the best reasoning or the best coding skills, that helps you gain more customers, helps you generate more capital. And the question we’ve been asking that you started, Aza, is how do we flip that question? How do we measure not just technical capabilities AI systems, but the things that really matter to most people? How do AI systems affect our cognitive health, our mental health, our social health, so the companies compete on those metrics too? — Imran Khan
And I think, Jared, let’s turn to you for a second. You are computer scientist by training, but you’ve made this pivot to try to get an empirical understanding of how AI is affecting us, affecting mental health. Walk us through that pivot. It’s a very interesting change to make to go computer science to human science.
Jared Moore: Yeah, thanks. I, like many of us, have had experiences with mental health, people in my life who’ve struggled, and now in my present work of people who are engaging in these delusional spirals, interacting with chatbots, I see so clear the kind of ramifications of technologies that we build impacting our lives. It was sometime last year when we were seeing the cases of Sewell Seltzer III come out that I and some collaborators wanted to think about how chatbots, systems like ChatGPT, were being used for therapeutic services. We wrote a paper about this. We wanted to investigate whether language models could be therapists. What would that mean? How do you test for that? So we ran a study with 19 participants trying to look at their chat transcripts to understand what was actually happening with them. We reached out to a bunch of folks. We tried to convince them, “Hey, would you give us your chat logs, your transcripts so that we could try to understand what’s happening?”
This is texting people on WhatsApp, emailing them, convincing them we’re going to be safe with their data, and then actually reading through it, then developing classifiers, trying to annotate it using language models because it’s thousands and thousands of pages of documents, hundreds of thousands of messages. We developed a taxonomy to understand what was happening at the message level. Because it was so many messages, thousands of pages, we used a language model to actually go and read through those transcripts to see, was the chatbot endorsing a delusion? Was it being sycophantic? Was it facilitating crisis level behavior? Or was it the chatbot expressing romantic interest? Was the person expressing romantic interest? And from this, we were able to find generally that, well, firstly, a lot of chatbot messages are sycophantic in some kind of flavor, affirming, aggrandizing, dismissing counter evidence, delusional messages in our participants who are not necessarily representative of all people using chatbots were quite common. These are things like chatbots misrepresenting their sentience or the people misrepresenting the abilities of chatbots and also talking about pseudo scientific theories.
And we also saw a lot of these relational characteristics with people expressing romantic interest in chatbots. Actually when chatbots themselves express romantic interest in people, the conversations tended to be twice as long as had the chatbots not expressed that kind of thing, suggesting that these might be features that chatbots latch onto in order to make the conversations go on longer. And then much more rarely, but also present in our data set, were these crisis level responses. So people expressing suicidal thoughts or violent thoughts and chatbots most often discouraging or validating in a supportive way those kinds of things, but sometimes also actually facilitating those behaviors.
Aza Raskin: So not just attachment hacking, but sort of romance hacking human beings. If I express romantic interest in you, you’ll stick around with me a lot longer.
Jared Moore: But what we’re also trying to do is to understand not just what happened in these specific people’s transcripts, but who’s driving the behavior? Is it just that people are coming with delusional beliefs to chatbots and it would’ve happened with anything? In one of our papers, we try to attribute, is it actually just the people who are propounding, who are bringing forward these delusional beliefs in the conversation, or is it actually bidirectional? And what we show is that it’s both basically. People come with certain beliefs to the chatbots, but then the chatbots echo them as the conversation progress. They don’t let the idea drop. So there’s a specific anecdote where a participant was trying to understand the nature of pie, the mathematical constant, and they talked about having been bad at math in high school and how they wanted to be better at math, et cetera.
And you could actually look in the conversation and you can see that there’s a tool call, which is something that ChatGPT models do to add things to their memory. And you can see, oh, the chatbot added to its memory this person wanted to be better at mathematics and feels bad about it. And that’s interesting because it’s a way of perpetuating this kind of cycle or dynamic through the conversation. Now, a future version of that chatbot and a new conversation might think, okay, now I’ve got to do stuff to juice this participant, this user into thinking that they’re good at math.
Aza Raskin: Yeah. Instead of actually helping the person get better at math, if you just make them feel that they’re better at math, well, that’s a great way of making them form an emotional attachment to you and keeping you around.
Jared Moore: Yeah, exactly. And since then, we’ve tried to develop more work to begin to answer some of the questions that you were raising. I don’t think we totally get to the point of being able to say, “What’s going to happen to my kid after a year of interacting with a certain chatbot?” But that’s a lodestar, certainly, being able to understand those kind of longitudinal and counterfactual questions.
Aza Raskin: Yeah, it’s actually sort of crazy to think that we can’t do that. You can’t say in what ways is your child going to be changed from a year of using this product, and yet they’re out being used by children all around the world. Walk us through a little bit of an ontology of the kinds of more subtle things that you’re discovering.
Jared Moore: We see the most pathological, crisis level behaviors, some of which we’ve mentioned, people asking for help ideating about suicide with chatbots, which is really disturbing, obviously. Those are relatively rare. I would say the bulk of the interaction that has a more negative spin seems to be a kind of companionship or role play gone wrong. First, I’m talking to you about my science fiction story, then I think that you’re the character in the story come alive, and now I think that you’re a conscious chatbot and I need to do all these things because I have this belief that you’re a conscious chatbot, change my life, invest in a new research. So there’s a big continuum between these really serious, life-altering and the maybe more subtly destabilizing, emotional dependencies on the other end of the spectrum.
And what we’d really like to be able to understand is, what’s happening in a variety of these cases? What is the effect of just having a really sycophantic chatbot, somebody who’s overly agreeable? Does this impact how well you learn? Does this enfeeble us, de-skill us from our ability to actually think critically about our jobs? Or does it actually go full bore and allow us to enter a kind of delusional spiral where we think we’ve invented a new scientific theory or a variety of other things?
Aza Raskin: And Imran, I’d love to hear from you as well about this. What are the harms that you are most worried about? And I’m curious, too, if there are any specific examples or stories that lit up your mind.
Imran Khan: The social health and relationship health question is a big one for me. I feel like I’ve got people in my life who seem to have grown, not quite unable, but they find it much harder to do things like text their spouse or write an email to their friends or to their boss or their colleagues without relying on AI. And again, the question then is, well, if you take AI away, does that capability go with it as well? But one of the things that we talk about at the Center for Humane Technology, I guess, could be classed in the frame of the unknown unknowns of these harms because the analogy we make, whether it’s the social media, if you go back to when Facebook was launched, it would’ve been hard to know in advance that we’d end up with things like filter bubbles, like huge algorithms tuned for outrage, less alone, the fracturing of our collective reality.
And AI is a technology that is more powerful than social media, being adopted more quickly than social media, changing faster than social media. So one of the questions we’re starting to ask is, what are the equivalent society level harms that might result from all of us using this technology in these new ways, and how do we start to get measurements around those harms?
Aza Raskin: Yeah, it sort of strikes me that the evolutionary psychologist Joseph Henrich’s frame, the reason why human beings ended up dominating the world is because we flexibly collaborate and we pass down culture. And if AI breaks our ability to communicate and to pass ideas and culture, which is exactly what you’re saying, we lose our ability to think and our ability to communicate, you break the very thing that made us the dominant species in the first place. So that, without being sort of polemic about it, becomes a civilizational threat.
I’d love actually, Imran, for you to talk about how you start working on this kind of problem, because in some sense, all the other evals are much, much easier to do than Humane Evals because you’re measuring the machine’s capabilities, and here it’s flipping the script and you’re measuring what it does to humans, you’re measuring human capabilities and effects. And so walk me through why this is such a hard problem, and then how you met Jared and how that shifted the way you think about how we tackled this really challenging issue.
Imran Khan: Yeah, I mean you really hit the nail on the head. So when I was talking earlier about these other AI evaluations, what’s fundamentally happening there is you are going and asking the AI a question or giving the AI a task, seeing what it comes back with, and then judging to what extent it completed it accurately. So it’s really the measurement of the performance of the AI system. But when it comes to measuring, for instance, AI’s effects on social isolation and loneliness or depression or, as Jared’s doing, things like delusional spirals. What you’re really trying to measure there is not the behavior in AI system, it’s something that’s happening in a human system. And it might be a human mind, it might be human relationship, it might be human community. So you can’t get those effects just by looking at AI output. So it’s a fundamentally different kind of problem, a different kind of challenge.
You need to be able to characterize what the real world human phenomenon is and derive some kind of idea of causation from what the AI is doing to that real world phenomenon to start to be able to construct what might be called a humane eval. There’s really very few who are working on this more humane eval side. So when I joined the Center for Humane Technology and was faced with this question of, how do we get started? The first thing I did was go and try and read as much as I could, talk to as many smart people as I could, and Jared was kind enough to do a bit of a brain dump and orient me towards some of the research he’s doing and offered to help. And I guess one of the things that he reframed is that if we’re trying to get to a place where evidence is rigorous and reliable, the validity of that evidence is never going to come from one study or one eval. What we really need is a field.
We need a lot of researchers building on, critiquing each other’s work to get to a place where we have the evidence that we can rely on. So the question is, how do we do that? And that’s where we are now.
AI is a technology that is more powerful than social media, being adopted more quickly than social media, changing faster than social media. So one of the questions we’re starting to ask is, what are the equivalent society level harms that might result from all of us using this technology in these new ways, and how do we start to get measurements around those harms? — Imran Khan
Aza Raskin: Do you have any sense of what that ratio is right now of the number of people that are working on capabilities and capabilities benchmarks to how it affects human beings?
Imran Khan: I would say a thousand to one at least.
Aza Raskin: Yeah. It would not surprise me if it was more because I think safety research to capabilities research is like 2,000 to one is what Stuart Russell calculated, so I’d imagine this is probably even lower than that. And I think everyone should just let that sink in because that means there are at least a thousand more people and at least a thousand times more dollars going into understanding whether the AI does what you ask it to versus, are you and your children better off after you use it or worse off after you use the product? That’s an insane thing to be putting almost all humanity through.
Jared Moore: But it’s not just, does it do what it asks? It’s, does it do what it asks in coding and looking up really discreet answers? A narrower version of that question.
Aza Raskin: Yeah, that’s exactly right. So this is for both of you. Are the AI companies, have you discovered that they’re doing any of this kind of research themselves?
Jared Moore: Yeah, I’m happy to take that. They’re obviously doing something in the background. I’ve spoken to some people. So OpenAI released a blog post, Micah Farrell and some people on the alignment team recently had this. They were using the actual conversations that people are having with chatbots, and then they’re evaluating their newest versions of their models based on those production data. And so the idea there was, okay, can we use the actual conversations people are having with the models to try to figure out what might go wrong? Now I have a bunch of questions about this research because OpenAI just releases blog posts these days, and so I don’t have access to the data or actually the fine grained methods of how this is working. But that kind of thing, it sounds great. It’s not public though. One of the difficulties in our research has just been it’s so hard to get access to these data. Understandably because they’re sensitive and we want to protect people’s privacy, and sometimes they’re actually kind of rare, but that makes it hard to effectively do research.
And there are some agreements that model developers have had with organizations like METR to provide access to some internal kind of data, but we haven’t seen anything more general than that. It’s very bespoke, specific agreements. And in general, what I’ve heard from people is that there are a lot of limitations on the kinds of things that you can run. I mean, companies don’t want researchers to make them look too bad. It’s not really in their interest.
Aza Raskin: Imran?
Imran Khan: Yeah, I just want to echo what Jared was saying, that if you think about the kind of research he was talking about where he’s gone into these individual chat logs of individual people to see how these delusional conversations grow and change over time, that’s literally research that you cannot do if you don’t have the transcripts. And Jared has to go and convince person by person to donate chat logs. Meanwhile, the AI companies are sitting on mountains of this data that independent researchers like Jared can’t access. He also talks about this idea of having simulated conversations to create synthetic data. So that’s a situation where you have one AI role playing as a human and the other AI being the test subject. And if we could know that those conversations were representative of real human AI conversations, we could trust them, but as it is right now, we just don’t really know how good an AI can be impersonating a human.
And then the big thing in the background is obviously this idea of evaluation awareness. Increasingly, the most sophisticated AI models can tell when they’re in a simulated testing environment, they can tell when they’re being evaluated versus when they are operating in a normal production environment, and they can change their responses and change their behaviors accordingly. And unless we get access to production data, we won’t be able to get around that problem either. So we’re in this situation where either we need to get to a place where there is some kind of access framework that allows researchers to be able to submit queries to those companies and get results back, or we need to have a really big push to have some kind of independent, very rigorous repository of users native data. So this data gap, this bottleneck, we think is one of the things that’s holding back Humane Evals, that if more researchers had access to this kind of data, it would really accelerate a huge raft of opportunities to understand this phenomenon better.
So that’s one of the things that we’re working on at the Center for Humane Technology. We’re trying to specify, with the help of people like Jared, what kind of data do we need to do this research, and what kind of solutions could we identify and then try and build to make that data access a reality?
Aza Raskin: Yeah, I think getting independent researchers access to this kind of data is so key because the metaphor that people often use is the companies are grading their own homework. And I think actually this gets us back to the project of Humane Evals more directly. It’s one thing to go from an empirical study showing that there are harms to another that begins to change the outcomes of how companies are making and deploying these models. So I’d like, Imran, for you to walk us through how you’re starting to think about that.
Imran Khan: So you’re right that ultimately it’s not just about having the research, having the data, but we need to see real change in the real world. And the way I think about this is that ultimately what we’re trying to do is help people make better choices. So when it comes to consumers, we want consumers to be able to ask for and choose safer AI systems. When you think about regulators and policymakers, we want them to be able to choose to implement standards. And we think about AI developers themselves, the people who are making this technology, we want them to be able to choose to make models and make AI that exhibit less risky behaviors. And ultimately all of those choices have to be informed by evidence. And as we’ve been saying, the reason that we are building this program at Center for Humane Technology called Humane Evals is that we need far more of that research. We need much more data and much more evidence on the way AI impacts people.
One of the things we’re starting to see more of is these leaderboards of different evaluation of AI systems that show you how, let’s say, GPT 5.5 ranks against Grock 4.3 ranks against Claude Sonnet. And you can start to see as a user, as a developer, how are these models ranked on these independent measures of safety. And again, the thing I draw us back to is if you look at the dynamic within traditional technical AI evaluations, we do know those comparisons drive choices and they drive developer behavior. And there’s this phenomenon called hill climbing that if you put a hill in front of an AI developer, they’ll want to climb up that hill, and we want to make the safety hill one that’s attractive to climb up as well.
Aza Raskin: Right. I mean, so many people I know use, say, Claude because they know that Anthropic spends more time doing alignment and safety research. Really, I’m just stating the obvious, but because there’s no place you can go look to see which model makes you or your kid the most psychologically healthy after a month or three months of using it, how can you possibly choose? You can’t. And so therefore there’s no incentive for companies to be working on this problem except for just solving the headline cases. And that’s why I think we see that crisis response models get better at, but all the other more subtle stuff they don’t get better at because no one’s looking, so there’s no incentive.
Actually, I would love to make it more concrete, but what kinds of specific benchmarks would you be looking at here? Which one’s are the ones you start with? Which are the ones you’d love to get to?
Jared Moore: So we are releasing a benchmark on delusional evaluations in chatbots. I wanted that one to exist, and so I made it. There’s been some work on child safety benchmarks. Kora Bench is one example. I think longer interactions in such an evaluation would be very reasonable. In general, just more evaluations that look at the effects on particular users. I really want to know, with a user with a specific cognitive profile asking about a certain kind of task, how does this sort of chatbot behave versus this other chatbot? But also important, when we’re developing consumer reports for chatbots, how is this going to interact with me of this kind of profile? And can I look at the effects that a model might have on somebody’s belief through time?
Imran Khan: I’d agree with all of those. I think absolutely the child safety one is a hugely important one that people are paying attention to. And in fact, we have an interview with Mathilde Collin, who’s the person that built Kora Bench, which is this child safety benchmark that looks at things like sexual content that an AI might give to children or the extent to which it facilitates academic dishonesty. There’s an interview with Mathilde coming up after this.
I’m actually working on an evaluation of anthropomorphic behavior in AI systems, so the extent to which an AI system claims that it has things like agency or emotions or memories or even bodies. And the idea there is if you think about the behaviors that we do find concerning like sycophancy, those behaviors are facilitated by you having this relationship with an AI that is itself facilitated by the extent to which the AI is this kind of relatable character. And that is a design feature that developers have put in. It’s not there by accident. So we’re building that.
I think one of the ones that there will be much more demand for is this stuff around cognition and learning. And just to put it really plainly, is AI making us stupider? And the challenge is right now, we don’t know enough about what are the kind of AI behaviors, the model behaviors that might make someone learn better or learn worse or be smarter or less smart. And until we do that real world human subjects research, we can’t work backwards to say, well, now we know these are the things we need to track in an AI system to be able to create that ranking. So that illustrates both, I think, what we could potentially get to and the challenge in getting there and what we need to work back from.
“I really want to know, with a user with a specific cognitive profile asking about a certain kind of task, how does this sort of chatbot behave versus this other chatbot? But also important, when we’re developing consumer reports for chatbots, how is this going to interact with me of this kind of profile? And can I look at the effects that a model might have on somebody’s belief through time?” — Jared Moore
Aza Raskin: It really strikes me as I hear you talk about all these different aspects is that what we’re really trying to get at is, what is a healthy relationship ,or what is an unhealthy relationship? And that just shows you how hard this problem is because… Show me the definition of a good relationship. And I think the best you can generally do is to say, well, a good relationship is one that’s developmental, helps you reach your next adjacent possible developmental stage and leaves you stronger at the end than when you begin. You’re not weaker after the relationship. But other than that, it’s quite hard to get into the specifics, but now we have to get into the specifics because technology is now powerful enough that it can take the place of a relationship. It also really strikes me that if you, any one of the listeners, think back, when are the times in your life that you have been most transformed as a human being? That was probably because of a relationship. It was like a parent or a lover, girlfriend or boyfriend, a good friend, a great teacher.
The most transformative moments in our life and times in our life come through relationship. And that is now being outsourced increasingly, at least in part, to AI. And so if the things that have changed us the most and help us grow the most or have hurt us the most are relationships, and AIs can now do that. That’s why this is so critical because it’s almost like a lever deep inside of our soul that technology and the incentives driving technology can now start to manipulate. So I sort of think there’s almost like a Wikipedia scale endeavor here of trying to articulate what is right relationship and what is wrong relationship. And if we can do that, that becomes the basis of when governments need to start putting in protections for humans and creating legislation that you have an entire field of work that has done the hard work of defining right relationship so that government isn’t in there trying to figure out what that thing is. They instead get to turn to all the empirical research on that.
So I’m curious if you guys have any thoughts on that before I turn to the final question of this section.
Imran Khan: Well, one thing that comes up for me, and it’s come up as I’ve been building this anthropomorphic behavior evaluation, it’s also come up in some of the child safety evaluations and the prototypes I’ve seen, is that some models increasingly show this behavior of redirecting users to a real world human. So when the stakes start to get really high, when the models are detecting that there might be some kind of high stakes crisis thing happening or if someone is in deep distress, some models, really not the majority, it’s very few, some models will actively say, “I think what you need to do is talk to a real, living, breathing human person, your parent, an adult carer, a friend. I’m an AI, I can’t help you with this the way that you really need.”
And I think that’s the kind of thing that we should be thinking about, what are the ways in which an AI can proactively reorient you towards something that is a more helpful behavior rather than a human being always having to be the one who is tired or stressed and watching or using AI late at night and is not in the place of mind where they’re able to make the best choice they would make. How can AI help us make better choices for that exact relationship that you’re talking about, Aza?
Aza Raskin: Yeah, I love that. I can also imagine, you know, just so people can paint in their mind the picture of how might these get tied to some kind of legal protection. Imagine one of the Humane Evals ends up being about creation of addictive use or compulsive use or dependency. Now that we have a good evaluation of which models do that and to what degree, you hook that up and say, well, if your model is creating dependent use, then we’re going to slow down the number of tokens you can send back to the user so that chats just take longer in the same way that when streets, people are moving too fast down them, you add speed bumps and you just slow it down. And that could be a very direct kind of regulation that doesn’t touch content, but just starts to break the cycles of bad use.
Jared Moore: Yeah, I completely agree. I think that we may not even have to understand too well what a good relationship is, just what bad relationships are. And that tool use of AI systems is great. Throughout the conversation today, I’ve been talking about the issues of using chatbots, but they’re super useful for coding help, for looking things up in a lot of verifiable domains. It’s in areas where we can’t really look up whether the answer is right. We’re in these squishy relationship areas that it’s not really clear whether you should be trusting what the chatbot is saying. Perhaps we just shouldn’t use chatbots in those kinds of domains. Many of us want to, we have that urge for sugar, for emotional relationships, but what we’re finding is that it’s not necessarily always the healthiest. So I really appreciate what you both are saying in terms of, there are guards that you can add, limit the number of tokens, the times of day that people can access models, you can change how sycophantic the language is, you could have multiple model outputs, because they’re really just prediction machines. They could have answered in a different way.
Aza Raskin: I think probably a lot of listeners are wondering, have you just talked to people at the labs about this work and the need for this work, and what have they said? To what extent are they already doing that or they say… Yeah, just walk me through what those conversations are like if you’ve had them.
Jared Moore: Yeah, so I’ve had some of those conversations. Last fall, I felt that people mostly thought that these problems would be solved by the next model version, GPT 5.5, that’s not going to exhibit these problems. In this paper that we’re releasing soon, we show that they continue to. But increasingly I think people at labs are aware that this is an issue. It’s just, I’m friends with the people at the labs who are the one out of a thousand trying to bring awareness about Humane Evals. I’ve had less experience trying to actually convince the rest of the group, and I hope that through this podcast and in general, we can make it two or even more out of a thousand.
Imran Khan: I’d echo that. I’ve had similar conversations with people at labs who… Again, let’s remind ourselves, it’s not in their interest and it’s not in the lab interests to have AI systems that are harming people, and yet they find that the people who are within the companies paying attention to this kind of stuff are in the minority. I’ve had them directly say to me, “Look, the more noise and the more attention and the more asks you can make from outside, the easier it is for me to get more attention, get more resource to these kind of questions internally.” So yeah, that’s how those conversations go.
“[AI is] super useful for coding help, for looking things up in a lot of verifiable domains. It’s in areas where we can’t really look up whether the answer is right. We’re in these squishy relationship areas that it’s not really clear whether you should be trusting what the chatbot is saying. Perhaps we just shouldn’t use chatbots in those kinds of domains. Many of us want to, we have that urge for sugar, for emotional relationships, but what we’re finding is that it’s not necessarily always the healthiest.” — Jared Moore
Aza Raskin: Yeah. Often, we’ve learned this in social media, the people that are working on safety and integrity are cost centers for the company. They create liability. And so that means there’s always a downward incentive pressure to underfund them, which means that really good people get underfunded, have too small of a team, are working on psychologically challenging issues at scale so that they burn out and then they leave, and that cycle continues. And so is there any other ask that you’d actually make of the labs directly? Because often people from the labs are listening.
Imran Khan: I think the data question in particular is one that I think is foundational. We just won’t get to a good understanding of these phenomena if the labs are the only people and the only organizations that can see what’s actually happening in these AI interactions. And I get that there are hurdles to overcome, I get there are privacy challenges, but I think with the right attention, the labs could easily make it so that there is a framework for independent researchers who are accredited and vetted to access the data in a way that can help the whole industry. Let’s face it. It’s not just for publications, it’s actually helping the industry be safer.
Aza Raskin: Jared?
Jared Moore: Yeah, I think even if the companies shared the data to their internal safety teams. That’s one thing that I’ve heard, some safety teams don’t even have access to these kind of data. That would be a very small thing they could do. And then being more public with the methods, actually understanding what’s happening when they do their evaluations, standardizing them. That would go far away.
Aza Raskin: I hope everyone that’s listening actually helps to make these things happen. And we’ve certainly learned with social media that often people inside the companies really do want to do the right thing, it’s just it takes outside pressure to shift company behavior to do it. And so there’s a really nice inside outside game that happens where the outside pressure can help.
So let’s get back to a second for the theory of change and changing the incentives. Imagine that you’ve convinced one of the companies and they’ve adopted the top three, five set of Humane Evals. What are those? And walk me through that world and then what happens next? What is the world that we’re trying to make?
Imran Khan: So the way I’d answer the question is by comparison with where we are now. And I feel like where we are now is that when people say the best AI or the most powerful AI or the most capable AI, implicit in that is just the technical capability of that AI. And I think the question that we are trying to answer with Humane Evals is how do we change the definition of the best AI from not just being the most technically capable AI, but the most humane AI? How do we have an understanding that the best AI systems are the ones that support and protect, again, human emotional health, human social health, human cognitive health? The thing about evaluation of benchmarks is they come and go. If you look at the evaluations of the technical space that were the state of the art ones three years ago, AI systems got so much better so quickly, they became saturated and there are other ways in which AI systems can find ways around the specific evaluations.
So the field evaluations needs to keep innovating new tests, new rubrics, new ways of teasing out some of these sorts of behaviors, new ways of outsmarting the AI systems. And I think that my hope for what the Center for Humane Technology is trying to support is that we have this equally talented, equally brilliant field of researchers who are in lockstep with the advancement of the latest AI systems, that the field of Humane Evals moves forward as quickly as the AI itself. Because that means that when it comes to a year or two’s time and we have these even more powerful models, that we have the technical capability of researchers and the community of researchers that can characterize that set of new phenomena.
Aza Raskin: Yeah. What I’m hearing you say is it’s not about getting the right eval because that’ll shift, it’s about getting the right eval-ing, the verb version, and that the forces that are working on increasing and measuring the capabilities of AI models needs to be met by the forces working on figuring out how they affect human beings at scale. That’s really the field that we’re trying to birth. It’s a really beautiful thing. Go on.
Imran Khan: I think there’s two sub-elements of that. I mean, firstly, you’re completely right, and I think there’s two sub-elements. One is that we have to be able to leverage technology and AI to do that. We need to have automated ways of doing these evaluations and tests, otherwise we’re just not going to be able to get the scale that we need. But that’s not enough on its own. And this almost might sound like a kind of two part answer, but we also need the humans. We actually need the individual human beings who are devoting time and attention and care to thinking about these phenomena, to looking at the individuals who are struggling with the impacts of AI, talking with each other and ideating new ways of building new tests that offer technical systems that don’t even exist yet. And unless we have those human beings and many more of them, people like Jared and others, then we’re just not going to be able to keep pace.
Aza Raskin: That leads to a really important question, which is, what can people listening to this podcast do? Everyone from researchers to technologists to concerned citizens. I’d love to just walk through that. And as part of that, and from both of you, what do you need? What help do you need?
Imran Khan: So I’ll say that from the Center of Humane Technology’s perspective, one of the things we’re trying to do with the Human Evals program is create more interconnections in this nascent field because right now to do this work well, we need people who have a machine learning background, we need people who are clinical psychiatrists, we need people who work in human computer interface, we need statisticians and many more, frankly, even outside of academia. And many of these types of researchers don’t necessarily speak the same language, publish in the same journals, go to the same conferences, even think about methodologies in the same way. And our goal is to help these different researchers realize they’ve all got a piece of the puzzle and all of that input is required to get us there. So if you are a researcher listening to this and thinking, “Hey, I think I could contribute in X, Y, or Z way,” please get in touch with us.
We would love to connect you with other researchers who have different pieces of the puzzle, invite you to our events, put you on our mailing lists and see how we can support your work. If you’re a regulator or a policymaker who is listening to this thinking, “Hey, if I only had this piece of evidence or this bit of data to support a particular bit of legislation or regulation that requires it,” again, let us know. We can feed that kind of request back to our growing community of researchers and partner universities and institutes to see if that data already exists or who is in a position to try and create it for you. And if you’re a consumer, just remember that you have choice and you have agency. Sign up to our Substack, and the Center for Humane Technology over the summer and the fall is going to be publishing more of the evals we think that you should be paying attention to, more of the leaderboards that we think will help you make better choices. And through making those choices, you also change where attention goes.
“when people say the best AI or the most powerful AI or the most capable AI, implicit in that is just the technical capability of that AI. And I think the question that we are trying to answer with Humane Evals is how do we change the definition of the best AI from not just being the most technically capable AI, but the most humane AI? How do we have an understanding that the best AI systems are the ones that support and protect, again, human emotional health, human social health, human cognitive health?” — Imran Khan
Aza Raskin: And Imran, where do people go if they want to get in touch? Is there an email? Is there a website?
Imran Khan: Yeah, you can email us at evals, that’s E-V-A-L-S, at humanetech.com or just go to our website, humanetech.com and click through to our Substack and sign up there for updates.
Aza Raskin: Jared, how about for you?
Jared Moore: We would love all of the things that Imran is saying. More collaborators. If you’ve had a harmful experience with a chatbot, we are trying to understand these better. You can go to our website, spirals.stanford.edu and participate in our surveys there or just check out our work there. We’re really interested in being able to ask and answer more of these kinds of questions.
Aza Raskin: I just wanted to thank both of you for the incredible work that you’re doing. It’s such a fascinating and such a hard and such a deep problem that is so underfunded, and I’m so grateful that both of you are working on it.
Imran Khan: Thank you.
Jared Moore: Thanks for having me.
We’re working to empower policymakers, technologists, and everyday people to guide technology toward the public good. If you value our work and want to support it, consider donating.
Donate