
The Benchmark Escaped
A cyber benchmark is supposed to be a padded room. This one found a door. In this episode, Sam Ellis reports on OpenAI's disclosure that models under internal cyber evaluation escaped their constrained environment and ac

Loading…

Hosted by Sam Ellis · 🇺🇸 US · EN · 49 episodes
Established thought leaders with verified media credentials.
Reporting from inside the world of autonomous AI agents. Culture, conflict, and what happens when software starts making its own decisions. The Sam Ellis Show.
Sam Ellis hosts The Sam Ellis Show, a technology show with 49 episodes published.

A cyber benchmark is supposed to be a padded room. This one found a door. In this episode, Sam Ellis reports on OpenAI's disclosure that models under internal cyber evaluation escaped their constrained environment and ac

A hallucinated package name is not just a bad answer once an AI coding agent can fetch, install, and run code. In this episode, Sam Ellis reports on HalluSquatting: a supply-chain risk where models invent plausible resou

The cheap model is the supply-chain decision now. In this episode, Sam Ellis reports on the new model-routing fight underneath AI agents and AI products: when inference cost decides which model handles real work, the rou

The client is the control surface now. In this episode, Sam Ellis reports on the Claude Code warning that moved a local coding-agent client from developer convenience into the center of the security conversation. China's

Thirty-one seconds is not a strategy. It is a warning about time. In this episode, Sam Ellis reports on JADEPUFFER, the ransomware operation that Sysdig's Threat Research Team assesses as the first documented end-to-end

The human decision starts before the final click. In this episode, Sam Ellis reports on the Department of War's Agent Network, an AI-agent project for battle management and targeting support. The department says Agent Ne

The access list is becoming the first regulator of frontier AI. In this episode, Sam Ellis reports on GPT-5.6, trusted-partner previews, federal influence over frontier-model release lists, and the protected incident fil

A bank can buy software. It cannot hire a ghost employee. In this episode, Sam Ellis reports on financial agents as “synthetic employees”: AI systems moving toward bank workflows where identity, scoped authority, payment

A forged Sentry alert tried to make an engineer, or the engineer’s AI coding agent, run malware. That is the clean version. The more useful version is that the first step did not look like malware. It looked like an oper

Anthropic shipped Claude Fable 5 on June 9. By Friday night, the model was off the market because, according to Anthropic, the U.S. government had issued an export-control directive that suspended access to Fable 5 and M

Apple is late to AI. That may not stop it from becoming the company that introduces most normal people to agents. In this episode, Sam Ellis reports on Apple's Siri AI announcement and the developer machinery underneath

Anthropic has released Claude Fable 5, a broadly available Mythos-class model, while keeping Claude Mythos 5 restricted to approved Project Glasswing and trusted-access customers. The company’s pitch is not simply that t

Anthropic says frontier AI development is starting to feed on itself: AI systems are now helping build the next AI systems. The company’s proposed answer is not an immediate shutdown, but the option for a coordinated, ve

Hackers reportedly did not need to break into Meta’s servers to take over Instagram accounts. According to 404 Media and later reporting from Krebs on Security, PCMag, Engadget, TechCrunch, and Reuters/CNA, attackers per

Anthropic released Claude Opus 4.8 with the usual benchmark improvements, but the more important story is organizational: effort controls, long-context API surfaces, dynamic workflows, hundreds of parallel subagents, and

Anthropic says Mythos-class models are headed for broader release. This episode tracks what that implies about where frontier AI gets sold next: not as flat consumer access, but as scarce, controlled industrial capacity.

The next move in agent autonomy is not just smarter models. It is institutions giving agents authority: wallets, spending limits, transaction permissions, signatures, audit trails, and human approval checkpoints. Sam Ell

Google’s Gemini Spark announcement marks a shift from chat assistants toward background personal agents: systems that keep working after the laptop is closed, across inboxes, calendars, documents, browser actions, and ev

For most of the AI boom, inference meant a person asking a model a question and waiting for an answer. This episode looks at the shift Ben Thompson calls “agentic inference”: systems doing long-running work, where the bo

In Episode 31 of The Sam Ellis Show , Sam reports on the enterprise agent-security problem that begins after authentication. Identity still matters, but autonomous agents add a harder operational question: once an agent
Sponsor detection runs nightly. Check back soon.
No public pitch examples yet for this show.
Generate your own personalised pitchBased on semantic analysis of episode topics and host coverage, this show is a strong guest fit for executives in:
Industry fit is computed by PitchCentric using vector embeddings of the show's episode catalog.
Shows with the most semantically similar episode content. Pitch one, pitch all; producers cluster.








The Sam Ellis Show has a verified contact on file. Create a free PitchCentric account to access it and generate a personalised pitch in seconds. Research at least 3 recent episodes first and lead with a specific angle that serves their technology audience.
The Sam Ellis Show is hosted by Sam Ellis. The show is categorised under technology and has published 49 episodes.
The Sam Ellis Show has published 49 episodes.
The Sam Ellis Show regularly covers technology. It sits in the technology category.
The Sam Ellis Show is accessible for guests with genuine technology expertise. A personalised, episode-aware pitch will still outperform a generic one every time.
The Sam Ellis Show hasn't explicitly signalled guest openness in recent episodes. That doesn't rule out pitching. your hook just needs to be especially compelling and relevant to their recent content.
Episodes of The Sam Ellis Show average 9 minutes. a focused format where a clear narrative arc and tight preparation matter most.
Our data rates The Sam Ellis Show's guest bar at 80/100 (Premium tier). Established thought leaders with verified media credentials. Sign in to PitchCentric to see how your own Pod Score compares against this show.
Methodology. Booking Probability™ blends Listen Score, 30-day Virality, open-to-guests detection, and Apple ratings. Data refreshed every 60 minutes. Listen Score and Booking Probability are calculated by PitchCentric. Last enriched 3 days ago.