How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale.Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams.Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimenta
Pitch Analysis
Required Pod Score for this show. PitchCentric checks your profile against host openness, topical fit, and audience signals before you generate a pitch.
Contact path
Public web form
Booking probability
31%
Guest openness
Selective
80/100
Required Score
Sign up to generate a grounded pitch for The Experimentation Edge.
Our AI reads these to draft pitches. Use them as grounding for a pitch that cites a real guest and a specific topic.
Episode #46
PayPal's $180 million experimentation win
Sep 30, 202632 min
Summary Gaurav Sethi joins Ashley Stirrup on The Experimentation Edge to explain how PayPal's experimentation program went from 700 to 800 tests a year with four week readouts to 2,587 experiments a year and roughly $180 million in measured impact. Gaurav inherited a reported win rate of 55% to 60%, four times the industry average, and traced it to experiments logging assignment data instead of exposure data, carrying 25% to 30% dilution. The conversation covers the exposure event his team introduced, the instrumentation and metric standards that cut readouts to 24 hours, a carousel test that a multi armed bandit resolved in 51 days instead of a projected 700, and why cost avoidance from losing experiments belongs in the ROI number. It closes on what changes when the thing you are testing is an AI agent rather than a button. Useful for product managers, engineers, data scientists and growth leaders building or defending an experimentation program at scale. Chapters 00:00 Cold open 00:54 Welcome Gaurav Sethi of PayPal 01:30 Elmo, PayPal's homegrown experimentation platform 03:53 $180 million in revenue impact and cost avoidance 04:28 The win rate that was too good to be true 06:29 Instrumentation standards and the exposure event 08:56 The carousel test: 700 days down to 51 11:29 Designing experiments so every result teaches you 13:45 Building the platform is only half the job 15:07 Education, office hours and executive support 18:28 Exposure events and joining transactional data 24:14 AI for experimentation, and experimentation for AI Takeaways - A win rate of 55% to 60% against an industry average of 11% to 14% was a tracking problem, not a performance one: assignment data carried 25% to 30% dilution and reached significance on users who never saw the test. - The exposure event fixed it. Fire an event as close to render as possible so the platform knows exactly when a user entered the experiment, instead of logging on page load. - Standardized instrumentation and canonical metric definitions are what let the analysis pipeline run without human cleanup, taking experiment readouts from four weeks to 24 hours. - A six variant carousel test projected at 700 days was resolved by a multi armed bandit in 51 days, with a winner worth almost five basis points. - About half of PayPal's $180 million impact in 2025 was cost avoidance from features that tested badly and never shipped, which is why Gaurav argues there are no losing experiments. Connect with the Guest Gaurav Sethi LinkedIn: https://www.linkedin.com/in/gauravsethi22 About the guest: https://www.growthbook.io/podcast/guests/gaurav-sethi Episode page: https://www.growthbook.io/podcast/episode/1-46 Company Website: https://www.paypal.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io
ServiceNow's Customer Zero approach to AI experimentation
Sep 29, 202628 min
Summary ServiceNow runs its own platform on itself. Ashraf Karim, Senior Vice President, Connected Customer Experience, owns everything that happens after a customer buys, and her team is Customer Zero, testing every product and capability internally before it ever reaches a customer. She joins Ashley Stirrup to unpack what that vantage point teaches you about experimenting with AI. The headline lesson is not a technical one. ServiceNow built agentic skills that worked, and users abandoned them anyway, because a task that ran five to fifteen minutes broke every expectation people had about what a machine should feel like. The fix was not a faster model. It was telling the user what the AI was doing while it did it. Ashraf also gets into the harder measurement problem underneath all of it: how do you evaluate a nondeterministic system? ServiceNow's answer is an internal LLM judge scored against a golden data set, plus a customer effort score that catches what CSAT reports too late. Chapters 00:00 What customers actually care about 01:14 Meet Ashraf Karim of ServiceNow 01:39 Owning the entire post-sale experience 02:08 Why ServiceNow runs as its own Customer Zero 03:02 Lessons from Google, PayPal and Verizon 07:11 The agentic skills users kept abandoning 09:29 Latency, expectations and the feedback loop fix 12:29 Users wanted the outcome, not the process 14:49 Designing experiments that can afford to lose 18:13 Using an LLM judge on nondeterministic models 21:48 Customer effort score as a guardrail metric 25:03 Where experimentation at ServiceNow goes next Takeaways - ServiceNow is its own Customer Zero, running the platform on itself and testing every feature internally before customers see it. - The agentic skills worked. Users still abandoned them, because a five to fifteen minute wait broke their expectation of what AI should feel like. - The fix was a feedback loop that narrates what the AI is doing in the background, which buys the patience a long task needs. - Users did not want the agentified version of the human process. They wanted the outcome and the next best action, not the steps. - Nondeterministic output needs its own evaluation layer, so ServiceNow built an LLM judge that scores against a golden data set before anything reaches a customer. Connect with the Guest Ashraf Karim LinkedIn: Ashraf Karim on LinkedIn About the guest: https://www.growthbook.io/podcast/guests/ashraf-karim Episode page: https://www.growthbook.io/podcast/episode/1-45 Company Website: ServiceNow Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io
Why Twilio ships on signals instead of significance
Sep 24, 202624 min
Summary Wanli Lau spent eleven years at Expedia Group, most recently running consumer product analytics on a program that shipped more than 1,500 A/B tests a year. Five months ago she moved to Twilio to lead global R&D analytics. In this episode she walks Ashley Stirrup through what carried over and what did not. The centerpiece is a test that looked like a win. At Expedia, an account sign-up takeover on the front page drove conversion sharply up, and quietly harmed return rate and engagement across every geography. That result is where Wanli's insistence on guardrail metrics and a trade-off decision framework comes from. She also explains the parts of B2B experimentation that have no B2C equivalent: choosing between user-level and account-level randomization, the interference risk when two colleagues at the same company see different pricing, and the many-to-many problem of one developer belonging to several accounts. And because Twilio's customer counts are lower than a consumer business, her teams often ship on signals and team conviction rather than waiting for statistical significance. Chapters 00:00 Cold open 01:02 Welcome Wanli Lau of Twilio 01:23 Inside Wanli's role at Twilio 04:03 Eleven years and 1,500 tests a year at Expedia 04:58 Why B2B experimentation differs from B2C 05:33 Cross-functional collaboration and good hypotheses 06:39 Teaching teams to experiment rigorously 07:48 The Expedia sign-up takeover that won on conversion 11:09 Guardrail metrics and the trade-off framework 14:52 Designing experiments to maximize learning 17:46 Diagnosing where a feature failed 19:10 North Star metrics and account-level randomization at Twilio 21:40 Learning from signals instead of significance 22:11 Where experimentation goes next at Twilio Takeaways - A test that wins on the primary metric can still be a loss. Expedia's sign-up takeover raised conversion and damaged return rate and engagement at the same time. - Name the primary, secondary and guardrail metrics before the test runs, not at readout. Deciding afterwards turns a result into a debate. - Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric. - In B2B the randomization unit is a design decision. User-level bucketing risks two colleagues at one account seeing different prices, and account-level avoids that but costs sample size. - When customer counts are low, significance is often out of reach. Learning from signals and shipping on team conviction beats waiting for a number that will never arrive. Connect with the Guest Wanli Lau LinkedIn: https://www.linkedin.com/in/wanlilau About the guest: https://www.growthbook.io/podcast/guests/wanli-lau Episode page: https://www.growthbook.io/podcast/episode/1-44 Company Website: https://www.twilio.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io
Will Guyeskey is Director of Digital Product at GoPro, where his team owns the e-commerce side of gopro.com. Before GoPro he ran personalization at Gap and cut his teeth at Brooks Bell, testing for brands like Barnes & Noble, Under Armour and Ralph Lauren. In this episode of The Experimentation Edge, he tells Ashley Stirrup why knowing your customer is the through line of every good test program. Will shares the Barnes & Noble order confirmation test he was sure would lose, and why the same idea never worked for any other client. He walks through a recent GoPro Mission launch test that asked whether a step-by-step configurator adds too much friction, and what a flat result revealed about high consideration buyers. He also explains why GoPro shares interim readouts across the company, why win rate makes a poor North Star for an experimentation program, and how his team plans to use AI for speed without outrunning its own learnings. Chapters 00:00 Intro 00:52 Will's role running e-commerce at GoPro 01:43 Learning A/B testing across retail at Brooks Bell 05:24 How GoPro runs one to three tests a month 06:35 Sharing learnings and interim readouts across teams 09:20 The Barnes & Noble order confirmation win 13:02 Why the win did not transfer to other clients 14:41 Testing friction on the GoPro Mission configurator 18:51 Why win rate is the wrong North Star 22:19 How AI will shape experimentation at GoPro Takeaways - A winning idea rarely travels. The Barnes & Noble recommendation module worked because of that audience's low order values and reading habits, and it failed for every other client that tried it. - Design every test so it teaches you something whether it wins, loses or ends flat. Losing tests are jet fuel when the learning is built in. - A flat result is still an answer. GoPro's configurator test showed that buyers of high consideration products accept extra steps when each choice adds value. - Share interim readouts across the company, and use them to show how volatile results are before a test reaches statistical significance. - AI can speed up building and running experiments, but a team that runs more tests than it can learn from is not getting better. Connect with the Guest Will Guyeskey LinkedIn: https://www.linkedin.com/in/willguyeskey/ Company Website: https://gopro.com Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io
Summary What changes when you take a mature B2C experimentation practice and apply it to a B2B storefront? Anuradha Tempe, Lead Product Manager at Samsung Electronics America, joins host Ashley Stirrup to share what she learned managing both sides of Samsung.com. She explains why B2B is not a sidekick to B2C, how a single bulk order can push a test into a false positive unless you normalize the data, and why not everything deserves an A/B test. She walks through the add-on experiment that nearly doubled attach sales once her team realized business buyers decide on mobile and purchase on desktop, and the Buy Now test that lost because B2B buyers value clarity over faster conversions. The conversation closes with her approach to North Star and guardrail metrics and the AI copilot she built to draft A/B test plans with a human still in the loop. A practical episode for product managers, engineers, data scientists, and growth leaders running experimentation across different customer types. Chapters 00:45 Meet Anuradha Tempe: from chip design to leading Samsung e-commerce 02:15 Why B2B is not a sidekick to B2C 03:30 Not everything needs an A/B test 05:15 Spreading experimentation practice across a global conglomerate 07:00 Bringing B2C rigor to a fast-paced B2B team 08:45 The add-on experiment that nearly doubled attach sales 11:15 Buyers decide on mobile and purchase on desktop 16:30 The Buy Now test that lost 21:30 North Star metrics, guardrails, and the EPP discount fix 26:05 An AI copilot for A/B test planning Takeaways -Treat B2B as its own customer base with its own testing discipline; a single bulk order on one day can inflate a B2B test into a false positive unless order data is normalized before results are read. -Not everything needs an A/B test; route lower-risk changes through UAT feedback or pre/post comparisons and reserve full experiments for features where being wrong is expensive. -Map where the decision happens, not just where the purchase happens; Samsung's business buyers decide on mobile and buy on desktop, and surfacing add-ons on mobile nearly doubled attach sales. -B2B buyers value clarity over faster conversions; a Buy Now button earlier in the flow confused bulk purchasers because returns and cancellations on large orders are costly. -Align on the North Star before building, whether it is revenue, engagement, NPS, or fewer support tickets, and set guardrails so an engagement feature can never quietly drag sales down. Connect with the Guest LinkedIn: https://www.linkedin.com/in/anuradha-tempe/ Website: https://www.samsung.com/us/business/ Sponsor GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io
Every question we get asked before someone starts their trial.
If you have a concern about deliverability, AI quality, data privacy, or whether this will actually work for your specific situation, it's probably answered below.
What is the difference between Founder Solo and Founder Pro?
Founder Solo gives you 50 AI pitches per month using the credit model (Standard pitches cost 1 credit, Enriched pitches cost 2). Founder Pro raises that to 200 credits per month and adds full Booking Probability access, unlimited Magic Match, Apollo enrichment credits, and data export capabilities. Both plans use the same credit system, so you can stretch your monthly budget further by using Standard-mode drafting.
How do agency tiers work?
Agency tiers have no base fee. You pay per managed client and per talent profile. Agency Standard is $199 per client per month; Agency Pro is $399 per client per month. Both add $39 per talent profile per month. Your own team's user seats are always free.
What is a talent profile?
A talent profile represents one person (founder, executive, or spokesperson) you are booking onto podcasts. It includes their bio, topics, headshots, and outreach history. Team plans include 5 profiles; agency plans are pay-as-you-go.
Can I switch plans later?
Yes, at any time. Upgrades take effect immediately; downgrades apply at the end of the current billing period. Contact support if you need help migrating between plan families.
Do you offer a free trial?
Every paid plan includes a 15-day free trial. Your card is saved at signup but you will not be charged until day 16. Cancel any time from your dashboard.
What happens if I cancel?
You keep access until the end of your current billing period. No charges after that. Your data is retained for 30 days in case you reactivate.
Is the 20% annual discount automatic?
Yes. Select Annual on the pricing toggle and the discounted price is applied automatically at checkout. The annual price shown is the full year cost.
What if I have more than 50 profiles or 20 clients?
That is our Enterprise tier. Contact our sales team and we will build a custom plan with volume pricing, a dedicated account manager, and SLA guarantees.