
Why Larger Models Learn More
In this episode: • Introduction: The Mystery of Scaling: Professor Norris and Linda introduce the core question of why larger models learn tasks that smaller ones cannot, challenging the standard sample-efficiency dogma.

Loading…

Hosted by Mechanical Dirk · EN-US · 158 episodes
Established thought leaders with verified media credentials.
An automatically generated podcast about machine learning and natural language processing. The two fictional hosts talk about papers that I want to learn more about on my way to work. It's not good, but it's useful.
Mechanical Dirk hosts Mechanical Dreams, a science show with 158 episodes published.

In this episode: • Introduction: The Mystery of Scaling: Professor Norris and Linda introduce the core question of why larger models learn tasks that smaller ones cannot, challenging the standard sample-efficiency dogma.

In this episode: • Introduction to Hybrid Architectures: Professor Norris and Linda introduce the episode's paper on hybrid attention architectures and discuss why scaling full attention is computationally expensive. • T

In this episode: • Welcome and the Autoregressive Bottleneck: Linda introduces the Nemotron-Labs-Diffusion paper and the hosts discuss the fundamental limits of sequential decoding. • The Best of Both Worlds: Joint Train

In this episode: • The Memory Wall and the 90 Percent Waste: Professor Norris and Linda introduce the episode's paper, discussing the GPU memory bottleneck caused by ultra-long context windows and the surprising observat

In this episode: • Distillation vs. Reinforcement Learning: Linda and Norris introduce the paper and discuss the differences between off-policy and on-policy distillation. • The Math Behind the Magic: The hosts dive into

In this episode: • Welcome and the Over-smoothing Problem: Professor Norris and Linda introduce the episode and discuss how deep Transformers suffer from over-smoothing, losing initial token-level information in later la

In this episode: • Introduction to Catastrophic Overtraining: Linda and Professor Norris introduce the paper and the counterintuitive phenomenon where better pretraining leads to worse catastrophic forgetting. • Feature

In this episode: • Introduction to the Compute Divide: Linda and Professor Norris introduce the podcast and discuss the massive computational barriers in modern LLM pretraining before introducing the HRM-Text paper. • Bi

In this episode: • Introduction: The Mystery of Warmup: Linda introduces a new NeurIPS 2024 paper that questions the true purpose of learning rate warmup. Professor Norris shares the conventional, yet incomplete, wisdom

In this episode: • Introduction to the Edge of Stability: Professor Norris and Linda introduce the paper and the surprising behavior of full-batch gradient descent. • Progressive Sharpening: Linda explains how gradient d

In this episode: • Introduction to SonicMoE: Professor Norris and Linda introduce the episode's topic, the SonicMoE paper, and discuss the recent trends toward fine-grained and highly sparse Mixture of Experts models. •

In this episode: • Welcome & The Quest for Lifetime Memory: Linda introduces the paper on Memory Sparse Attention (MSA) and sets the stage by comparing current LLM context windows to human lifelong memory capacity. • The

In this episode: • Welcome and the Post-hoc Problem: Linda introduces the paper and the hosts discuss why post-hoc unlearning methods fall short against adversarial attacks. • Token vs. Document Filtering: An exploration

In this episode: • Welcome to Mechanical Dreams & The Pretraining Problem: Linda introduces the Meta FAIR paper on Self-Improving Pretraining, and Professor Norris questions why standard next-token prediction is no longe

In this episode: • Introduction: What is a Duplicate?: Professor Norris and Linda introduce the paper Scale Dependent Data Duplication and discuss the core question of what really counts as a duplicate for a language mod

In this episode: • Welcome and Introduction: Professor Norris and Linda introduce the podcast and the topic of the week, discussing the general concept of representation degeneration in neural language models. • The Culp

In this episode: • Introduction: The Gold Standard in Question: Professor Norris and Linda introduce the episode's paper from DeepMind, setting the stage by defining perplexity and explaining why it is universally used t

In this episode: • Introduction to Downstream Scaling Laws: Linda and Professor Norris introduce the paper and discuss the limitations of traditional parametric scaling laws for predicting downstream task performance. •

In this episode: • Welcome and the Mamba Lineage: Professor Norris and Linda introduce Mamba-3, discussing the shift towards inference-time efficiency and the need for sub-quadratic models. • Exponential-Trapezoidal Disc

In this episode: • Welcome & Introduction: Professor Norris and Linda welcome the listeners. Linda introduces the paper of the week, teasing the unexpected comeback of non-linear RNNs. • The Expressivity Gap: Linear vs.
Sponsor detection runs nightly. Check back soon.
No public pitch examples yet for this show.
Generate your own personalised pitchBased on semantic analysis of episode topics and host coverage, this show is a strong guest fit for executives in:
Industry fit is computed by PitchCentric using vector embeddings of the show's episode catalog.
Shows with the most semantically similar episode content. Pitch one, pitch all; producers cluster.







To pitch Mechanical Dreams, visit https://mechanicaldreams.mechanicaldirk.com for contact information, then craft a tight one-paragraph hook that ties your expertise to a gap in their recent science coverage.
Mechanical Dreams is hosted by Mechanical Dirk. The show is categorised under science (natural) and has published 158 episodes.
Mechanical Dreams has published 158 episodes.
Mechanical Dreams regularly covers science, natural. It sits in the science category, with a natural focus.
Mechanical Dreams is accessible for guests with genuine science expertise. A personalised, episode-aware pitch will still outperform a generic one every time.
Mechanical Dreams hasn't explicitly signalled guest openness in recent episodes. That doesn't rule out pitching. your hook just needs to be especially compelling and relevant to their recent content.
Episodes of Mechanical Dreams average 22 minutes. a focused format where a clear narrative arc and tight preparation matter most.
Our data rates Mechanical Dreams's guest bar at 80/100 (Premium tier). Established thought leaders with verified media credentials. Sign in to PitchCentric to see how your own Pod Score compares against this show.
Methodology. Booking Probability™ blends Listen Score, 30-day Virality, open-to-guests detection, and Apple ratings. Data refreshed every 60 minutes. Listen Score and Booking Probability are calculated by PitchCentric. Last enriched 2 days ago.