
Episode #17
More Agents Than Employees: How Zapier Disrupted Itself Before AI Could
The best AI model in the world just scored 18.1%. On Zapier's own benchmark for real business work — the cross-app tasks any white-collar worker does every day — even the top frontier model completes them barely one time in five. That's the number Wade Foster keeps pointing at, and he runs an automation company that stands to gain from the hype. Instead, he makes the case for what actually works right now: not turning a model loose, but blending deterministic workflows with agents where each is strong. In this episode of Talking AI, Matt Paige sits down with Wade Foster, co-founder and CEO of Zapier, who built a scrappy Y Combinator startup into the $5 billion plumbing of the SaaS era on barely a million dollars raised. Foster called a company-wide “code red” the week GPT-4 launched, and he's spent the years since rewiring how Zapier — and its customers — actually use AI. The conversation covers why he shut the company down for a week in 2023, how AI habits actually stick, what Zapier's AutomationBench reveals about the gap between benchmark scores and real-world reliability, why coding models improve faster than knowledge-work models, how to tell a workflow from an agent, and the difference between individual AI and the institutional AI almost no company has cracked. In this episode, you'll hear about: The three things about GPT-4 that triggered Zapier's first-ever code red How daily AI use jumped from 11% to over 50% in a single hackathon week The moves that make AI habits stick: show-and-tell, repeat hackathons, and “not yet” Why the best model on AutomationBench still scores only 18.1% Why coding is easy to verify — and subjective knowledge work isn't The power of hybrid setups that blend deterministic workflows with agents Wade's prediction: most tokens on open-source models, most spend on the frontier What actually makes a good eval — hard for models, easy for humans, private data A plain-English definition of an “agent” versus a deterministic workflow The daily recap workflow Wade thinks everyone is sleeping on Floor raisers vs. ceiling raisers — and why individual AI isn't enough Why the six-month product roadmap is dead Key Moments 00:04:40 — Making AI habits stick: show-and-tell and repeat hackathons 00:06:38 — Differentiation when AI is best at the thing you sell 00:09:34 — AutomationBench: the best model scores just 18.1% 00:11:31 — Why the top model stalls: verifiable code vs. subjective work 00:14:19 — Getting squeezed on both sides: AI in the company and the product 00:15:20 — Model efficiency, Coinbase, and the token-maxing debate 00:17:18 — What makes a good eval 00:19:30 — What actually counts as an “agent” 00:23:12 — Iterating on workflows with your own mini-evals 00:26:15 — The kind of worker thriving right now 00:27:36 — Wade's favorite workflow: the daily recap 00:30:44 — Floor raisers vs. ceiling raisers for AI adoption 00:34:55 — From individual AI to institutional AI 00:37:58 — Why the six-month roadmap is dead Key Links: Zapier Connect with Wade on LinkedIn Mentioned in this episode: AI Opportunity Finder Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. Try it now at https://hatchworks.com/ai-opportunity-finder/

