
Episode #72
Tuning GPU Performance with AI Agents | AMD’s Anush Elangovan on ROCm 10
Watch the full conversation on YouTube AI agents can keep tuning GPU workloads after you step away from the keyboard. Anush Elangovan, Corporate VP of AI Software at AMD, returns to Chain of Thought to explain how that works with Hyperloom and ROCm 10. Anush and Conor Bronsdon trace the process from installing ROCm through Claude Code or Codex to profiling workloads, finding slow kernels, and testing optimizations while preserving numerical accuracy. Anush shares a Hyperloom run spanning 14,000 models and explains why clear goals and feedback matter when agents are doing the tuning. They also explore what comes next for engineers: keeping skills and frameworks reliable, managing the security and accountability of autonomous agents, and applying AI to the last mile of useful software. We cover: How agents help install ROCm and serve models through natural language How Hyperloom profiles workloads and uses LLMs to explore optimizations The role of GEAK in tuning kernels while preserving numerical accuracy Anush’s account of optimizing 14,000 models in one pass How software improvements get more performance from existing GPUs Keeping agent skills current and testing across AI frameworks Security, accountability, and the next bottlenecks in agent-driven development Chapters: (0:25) A decade of ROCm, now agent native (3:03) What agentic ROCm looks like in practice (6:17) Installing ROCm then versus now (9:14) An order of magnitude more CI across every framework (10:44) Anush’s workflow: agents and deployment (12:21) Speed is the moat (15:00) Success is a stranger who cannot spell ROCm serving an LLM (17:09) Keeping agent skills from going stale (21:38) Co-designing kernels with the frontier labs (24:06) Hyperloom, GEAK, and 14,000 models in one pass (26:45) Managing autonomous agents: control and liability (32:21) Security at the speed of agent swarms (36:03) ROCm performance gains on the same hardware (37:36) Where enterprises hit walls in production (40:24) Why coding was the right reward function for AI (44:42) Which industries get the next software scale unlock (47:02) The last mile of AI (50:31) Closing thoughts Connect with Anush Elangovan: LinkedIn: https://www.linkedin.com/in/anushelangovan/ Twitter/X: https://x.com/AnushElangovan ROCm.AI: https://rocm.ai AMD AI blog: https://www.amd.com/en/blogs/by-author/anush-elangovan.html AMD AI Developer Program: https://www.amd.com/en/developer/ai-dev-program.html Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot





