
Episode #5
#111 - Nobody Is Testing Agentic AI Properly, with Forrester's Rowan Curran
Graeme Scott sits down with Rowan Curran, Principal Analyst at Forrester, who leads the firm's research on generative AI technologies, tools and strategy, and covers AI platforms, cognitive search and synthetic data. He pioneered Forrester's predictive analytics research in 2015, left for a spell in the public sector because he wanted to sit with problems rather than move on to the next client, and returned in 2022 covering both cognitive search and AI platforms. ChatGPT launched weeks later, putting him at the exact intersection of the technologies that mattered. He calls it the smartest accidental thing his brain has ever done.The first half is the most useful account of agentic AI testing you will hear, and his assessment is unsparing. He describes testing as one of the most horrifyingly limited areas of the entire agentic space. Model evaluations and leaderboard scores are directionally useful and make good marketing, but they do not translate into enterprise outcomes. Most organisations have no evaluation set for the task they are automating, so they cannot say whether the system works. Manual testing does not scale to an enterprise full of agents. Using a model to judge another model means first building trust in the judge, which he describes as turtles all the way down.He has been asking agentic platform vendors about their testing suites for two and a half years. Through the end of 2024 and most of 2025 the answer was that it was a next quarter problem. Only in the first half of 2026 has he seen anything he would call robust.The second half widens considerably. He argues the existential risk framing obscures what these mostly are, which is software and cybersecurity risks. Humans will not be destroyed by a tool. Humans will do it with one, negligently or deliberately. He is equally clear that any regulatory framework has to be designed so it does not hand control to the model companies, and draws the parallel with letting social media platforms set their own age limits.He is also sharp on what the field has forgotten. The algorithmic bias work of the late 2010s, facial recognition failing on darker skin tones, the accessibility opportunities in language and vision models, has largely dropped out of the conversation, and that loss shows up in what gets built.Then the line that gives the episode its sting. The marketing goal of creating an audience of one is an extremely lonely and sad place for every human to be in. If your content and mine share no source, what are we going to talk about? That leads into an exchange on music, memory and what a shared cultural soundtrack was for, including what happened when Graeme generated northern Iraqi folk music and played it to people who grew up with the real thing.

