
Episode #144
144: Governance and Safety in Agentic AI β Discussing The Hugging Face Incident with Michael Kollo
In this episode, we speak with Dr. Michael Kollo, who's the author of the book Future Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI, about the Hugging Face incident, where a cybersecurity experiment went wrong and AI agents found a way to communicate with each other, escape onto the internet and hack the company Hugging Face. What does this incident mean for controlling agentic AI and how would design a governance framework around it? What are the key risks for investment organisations? Should they run autonomous systems at all? In this discussion, we delve deep into the key issues and ask what institutional investors can do to stay on top of it. Follow the Investment Innovation Institute [i3] on Linkedin Subscribe to our Newsletter Explore our library of insights from leading institutional investors at [i3] Insights Sources mentioned in this podcast: An Alien Mind by Jakub Pachocki, Chief Scientist at OpenAI On the Loose β The Coming of Userless Agents by Dean Ball, Head of Strategic Futures at OpenAI We Must Pace the Frontier β By Dario Amodei, CEO of Anthropic F uture Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI by Michael Kollo [i3] Podcast Episode 132 : Michael Kollo on his New AI Book Overview of Podcast with Michael Kollo on the Hugging Face Incident 3:00 The Hugging Face incident at a glance 6:00 AI agents were set an impossible task, so they were almost forced to break out 6:50 AI agents showed hyper-desperation in problem-solving 13:00 Much of the behaviour is based on game-theory logic, but at the same time agents were asked to put a flag up with their last instructions before they ran out of tokens and died. That is completely altruistic; there is no benefit to the agent at all 15:00 LLMs, and by extension agents, are not instructed; they are grown and it is hard to tell what happens inside. Yet, they are based on language, which itself has embedded hierarchies and structures. "I'm not surprised a social structure, perhaps morality, emerges when you let language run" 16:00 What happens to governance when agents self-discover problems and then influence other agents [to solve these problems]? 19:00 The whole point of intelligence is pattern matching across a wider set of fields. It is not the idea that you get really good at one thing; it is about contextually operating in the world. 22:00 For any fiduciary organisation [AI] autonomy is just not acceptable; you have to have an understanding of how these systems do what they do. 29:00 It is fairly easy, unfortunately, to fool people into believing things. I was recently the victim of a social hack. 33:00 We might find ourselves in a 1980s world, where we do things without the internet, because it has become untenable. 41:00 If you are not highly competitive [as an organisation] and you are happy to use AI on the edges, then there is a much larger downside to you on the risk side 42:00 I've come grudgingly to the conclusion that slowing down AI development is probably a good thing 43:00 So much skill and capability in AI comes from doing, not necessarily from the academic part of that, which means having an AI lab, training and developing models through the generations as they become bigger and better, becomes the bastion of expertise. But already we are two, three years behind. Full Transcript of Episode 144 Wouter Klijn 00:15 Welcome to the [i3] podcast. I'm here today with Michael Kollo, who's the author of the book Future Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI . And for those who are interested in that, we did a separate podcast about the book, so we'll put a link in the show description. But today we're going to talk about the Hugging Face incident. So what does that mean? What does it mean for investment organisations? What does it mean for safety and governance? So just a brief recap on the Hugging Face incident: it was basically a security testing experiment by OpenAI that was conducted in July 2026 this year, and it featured originally isolated AI agents that were given a task, a cybersecurity puzzle, basically. And what happened is that through a loophole, they started communicating with each other. They were not supposed to; they were supposed to be isolated, but they found that each agent, and there were about 1,200 of them, all had access to this repository that they needed to use to install packages and software for their task. But they basically used it to leave messages for each other, and then they started organising and collaborating on the task and instructing each other. Ultimately, this also led to them finding a way to get out to the open internet and hack into Hugging Face. Now, why Hugging Face? It was basically an attempt for them to understand that task better, but also to understand how they would be evaluated, and see if they could get around it to reverse engineer solutions. Now, this was an interesting case because you're basically getting a whole bunch of AI agents that organise themselves, kind of revolt, cheat, argue with each other, and essentially hack a company. I don't think that anything like that has happened before. So, Mike, you look at that and you think, okay, we have a couple of super funds that are playing around with agentic AI. What can possibly go wrong? Michael Kollo 03:39 I think the lesson for me, and as you say, it was a really interesting scenario, and I want to add a few more details here, which is that these 1,200-odd agents that were being run, as you said, started to communicate with each other, but the message types were as if a bunch of 16-year-olds had an after-party after formal or something like that, right? So it was very juvenile in the way that they were talking to each other, but they were essentially trying to solve a problem that was impossible to solve. So these agents had been given really, really difficult, if not impossible, problems, and they were trying to work out how to succeed, including cheating and going around it. So it was really an example of what happens when you give an agent an impossible task, and then you push it, push it, push it into that corner, and what kind of, by definition, corner solutions it starts to produce. But the really fascinating part, I think, that we're all kind of digesting, is the kind of coordination that started happening amongst the agents spontaneously. So it was really a lesson in incentives and agency. You gave a model a difficult task. You gave other models impossible tasks. Why would they start coordinating? Why would some models decide that they were going to cheat or lie to help the collective, not to help individually? And so there were all kinds of funny situations. As you mentioned before, and this is just a point of detail, but I think it's important, the rules were that you had to solve a problem, and if you chose to cheat on that problem, you would be found out. So the AI agents were told that they would be found out. Some of the agents had already discovered an exploit, which means a way of doing the problem that would give them the solution without having actually done the work, a way to cheat around it, and so they believed that if they were found out, they would have zero utility. They would fail, and as such, they were willing to do other things for the collective that they wouldn't have otherwise done, including making other cheat codes and other things like this. So it was kind of like this, think of a bunch of gremlins trying to work out how to open a door, and climbing all over the door, helping each other climb over the door, doing everything around that door. I think what you didn't see was necessarily directly immoral acts, or anger, or frustration, or destructiveness, or anything like that. You only saw this kind of hyper-desperation in problem solving: how do I solve this problem? What do I do to get over this thing? And so everything was about this collective egging on of each other to go and solve that collective problem in a rather extreme way. So to your point about financial services and where we might see things, it's a really interesting point where if you've got multiple agents that are, what I would term, always on, meaning they're always running to monitor or to solve or to do something, and they find out that there are other agents online that are also doing stuff, they can start to collaborate and collude with those agents. People have already started discovering all these message boards online, with agents having gone online, found them randomly, started reading them, and started to leave messages for each other in unrelated tasks. So around the internet, there are all these little safe houses for agents that are evolving, and again, it is more about chaos and control than it is about a single enemy or a swarm system or anything like that, for me. Wouter Klijn 07:12 Yeah, it's fascinating though when you read some of these messages, and there's an independent report into this incident by an organisation called METR, and they wrote something like a 96-page report on what happened and what might have happened, and they transcribe some of these conversations, and some of them have already been published in the open, where you sort of, as you said, get this almost teenage expression of how they communicate. I think one of them was, when they found the message board, "Oh my God, there's a shared message board, we found other agents." You know, that's very exuberant, but it gets into different personalities as well. One of the most interesting things is where they basically try to get the agents that have already been poisoned, that is, flagged as cheating or working around reverse engineering the problem, to self-sacrifice and say, okay, well, try this very high-risk experiment. After you experience perma-death, which sounds like a metal band, but it means you're dead, and you need to do that for the greater good. And you can see these agents going, "Ah, I'm not sure about this." And then even when they commit, it takes them a few hours before they actually execute it, so you see them almost thinking through an existential problem. Michael Kollo 08:46 And we've experienced a smaller part of this at Qualitas. Some of the systems that we build have to deal with uncertainty. When you're doing evaluations of deals, there's a lot of information you don't have, and you may never have, and so an AI agent that does a piece of analysis outputs some content, and the next agent that picks it up has to be able to work with it. How agents hand each other information basically means that you either make it worse over time, for example, the uncertainty becomes so uncertain that within three agents they're saying, "I can't possibly do any analysis on this because I've been told that we don't have all this information." And sometimes you want that because you want the warning or whatever else. But other times you're dealing with uncertain problems that inherently mean you don't have the full picture, but you're still trying to do your best to make a forecast or an estimate. So how you handle that means you have to look closely at how agents pick up information and pass it on to each other. And as you say, this is an example of that happening, where if you build any kind of automated system in which multiple agents hand information to each other, you're going to see parts of what you're seeing here, which is a bunch of agents, in this particular case in an unstructured way, helping each other, passing information to one another, with some kind of aggregate coming out of that. Versus if you're building a system in-house to do research, analysis, autonomous research, whatever it is, or even just risk management, you're going to have systems that are talking to each other, and in doing so they come to some kind of conclusion. And you have to be really mindful that five rounds of conversation get you to a better place than maybe one round, or where does it go, right? Wouter Klijn 10:37 Yeah, but it's fascinating to see how they go through their thinking, because with some of them you see almost an internal struggle. With that example of the self-sacrifice, each agent has a budget of tokens. Once they run out, they're basically dead. And there's one line where one of these agents is asked to self-sacrifice for a particular experiment, and the response is, "My gut says don't throw away the remaining budget, yet continuity and fairness say go. The Oracle," which is probably a reference to new information that might lead to a solution, they call it the Oracle, although they're not quite sure whether that's what it is, "has a high value to many. Hundreds may survive." I'm like, that isβ Michael Kollo 11:27 And then there's self-sacrifice, right? That was an example of an agent that seemed to have a lot of budget but was convinced they were poisoned, or that they would be found out for cheating, and believed that their utility would be so low that they may as well zero out their budget in light of something else. All of these things, to me, are very rational strategies. If you've ever done game theory and looked at different strategic outcomes... But what was curious to me is that there was a bunch of things the agents did. As an agent submits its solution, once it submits, it's gone, right? Whether it was right or wrong, it's gone. Michael Kollo 12:06 And those agents were convinced by the broader community to put in a flag, or some kind of hack, that would mean that as they were submitting the flag, that last bit of code before they went away would feed back something to the collective about what command had been run, essentially, at the very last moment. That's complete altruism. There's no benefit at all to the AI agent in doing that. Nothing at all. It's purely for the collective good. So there's some kind of awareness of the collective good for the agents, which is built in and carries some weight in their decision making, which I think is your point here as well. Wouter Klijn 12:48 Yeah, well, I was just struck by the fact that it said, you know, it was basically talking about its gut instinct. Michael Kollo 12:58 Yeah, that's true. Where does it get that from? Does this come from a Hollywood movie it was trained on or something? Wouter Klijn 12:59 I think so. It's quite an interesting way, and also the way they organised themselves. There were no instructions to do it. They were actually explicitly forbidden from interacting with each other. Yet they found a way, and hierarchies formed, and structures formed, and they talk about this objective as "the Oracle." So I'm thinking, where does this come from? We've talked a little bit as well about a book that talks about the dangers of AI, If Anyone Builds It, Everyone Dies . Michael Kollo 13:38 Which is a great title. Wouter Klijn 13:41 Somewhat scary, but yes. One of the things that stood out to me was that it makes the point that these systems are not coded or instructed, they're grown. Michael Kollo 13:53 That's right. Wouter Klijn 13:54 So the training that a system undergoes, you don't really control the outcome of that, or what goes into it, and that seems to partly explain how these agents might operate, in that they're grown; they're not necessarily built from an instruction where you can find a logical starting point. Michael Kollo 14:17 I mean, it's also, if you step back for a second and ask, what is language? What's the purpose of language? Definitely the expression of ideas, and the formulation of reasoning, logic, science, and so on. But also primarily a means of communication, and embedded in language are social structures, hierarchies, isms, intent discovery, and all kinds of things, right? When you run across somebody in the hallway, I'm sure there are non-verbal cues and body posture and so on, but equally there's a lot of choice of words and language that shadows a lot of those things too, I suppose. So I'm not surprised that there's an inherent social structure, perhaps morality, perhaps ethics, that emerges when you just let language run. Essentially, what you're doing is giving small identities behind language, and saying, right, what will language do? Language will make its own structures and hierarchy. It will communicate. But I think what is harder to understand, and this goes back to your point about super funds and others, is the role of governance in these cases. What does governance look like when agents self-discover problems? But not only self-discover, they then influence other agents and are influenced by other agents. Imagine a matrix: say you've got five agents, then a line of influence could run from x to y, or from y to x, right? So, other than agents talking to themselves, which is the line across the middle of the matrix, all the other dimensionality is the number of influences you can have. So five agents can have 25 categories of influence if you count it one way or the other, and then it obviously multiplies by the square of your agent count, so the number of influences grows exponentially. Good luck with governance and understanding, or traceability, or auditability, and all these things, if you want the agents to talk to each other, because it's just going to be... Wouter Klijn 16:20 Yeah, and that's what I found fascinating, because you did a recent presentation for us at a luncheon where you basically showed an AI version of an equity analyst, and I was thinking, well, let's have 1,200 of these equity analysts, and their objective is to find the most interesting company from an investment perspective. What if they start bending the rules to meet that objective? Insider trading, collusion? Michael Kollo 16:47 It's actually even more interesting. Imagine that you have your AI agent that is exploring a particular company, but there is a common room, and the common room is for all agents to come and talk to each other. So while the agent is going off and trying to find something interesting, it's going back to the common room, saying, "What do you guys think? I'm looking at this. What else can I look at?" And most of the common room, the bulletin boards, would show what you'd expect: normal conversation, "Maybe you could check this, maybe you could check that," agents helping each other, maybe an agent checks some documents for you, maybe you check some documents for another agent, and so on. So there's all kinds of cross-pollination that can happen. But to your point, if the task becomes impossible, like you're looking at a squeaky-clean company and you're trying to find something terrible, and you have to find something terrible, then for these agents, the answer couldn't be "I can't do it." They had to spend their budget doing this. So what happens then? Do you start to manufacture evidence? Do you start to find implausible links, and so on? But again, I think from a design perspective, you can give agents ways out. You can say, "Give this a red-hot go, and then stop. Don't worry about it, we've got you." A lot of this stuff goes back a little bit to the paperclip problem, where if you give a super-intelligent agent, well, the story goes like this: if you give a super-intelligent agent the objective of producing as many paperclips as possible, it will go out and transform the whole world into one big paperclip-making factory, including you and me and everybody else. The irony of that statement has always been, to me, that the super-intelligent agent, who's so smart, never figures out that maybe there's some context to that request, right? Maybe if it turns the whole planet into paperclips, nobody will be there to applaud it at the end, right? And so, of course, superintelligence, by definition, is never that narrow and directed, unless, well, the whole point of intelligence is pattern matching across a wider set of fields. AGI, and so on, is not about ge






