62 episodes
- Brendan O'Donoghue, research director at Google DeepMind, makes the case for text diffusion as a real alternative to autoregressive generation. He walks through how discrete diffusion works, why diffusion samples are far more diverse and what that unlocks for RL, where the Gemma diffusion model actually stands against frontier models, and why the whole training and serving stack being hyper-optimized for autoregression is the main thing holding the approach back. The conversation also covers hardware trends favoring flops over bandwidth, AGI timelines and real-world bottlenecks, and why he thinks RL is still underhyped.
Key topics
- Discrete diffusion for text vs autoregressive generation
- Why diffusion samples are more diverse, and what that unlocks for RL
- Where diffusion already wins: latency, on-device, robotics
- Why serving cost, not quality, is the real blocker
- RL as the most underhyped area in AI
Timeline
00:00 Introduction
00:50 What diffusion models are and how text diffusion works
04:40 Why Brendan bet on text diffusion in 2023
07:15 Diversity, creativity, and why it helps RL
11:00 The best diffusion LLM today and the gap to frontier models
14:25 Latency, serving cost, and why it needs more chips
17:14 Where diffusion already wins: on-device, robotics, battery
20:14 One model, two modes: diffusion for thinking, AR for answering
22:24 Samplers and the stuttering problem
26:27 Theory, BERT, and why now is a good time to work on this
31:48 Pipelines built for autoregression, and continuous diffusion
35:35 Hardware: flops vs bandwidth
39:49 AGI timelines and real-world bottlenecks
50:15 Is AI engineering or science?
54:14 Most overhyped and most underhyped ideas
58:35 RL on diffusion, value functions, and exploration
1:07:30 Go download the model and break it
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. - Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He's also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic's famously open internal culture to break after its IPO.
Key Topics
Open vs closed models and who actually captures the value
Anthropic and OpenAI as opposite cultures, and the talent concentration problem
Boom vs bubble, and why token spend hasn't produced 10x better products
Continual learning, RSI skepticism, and what Nathan wants to work on next
What the open ecosystem needs economically to survive
Timeline
00:00 Intro
00:27 Open vs closed models, and who actually captures the value
05:12 China, harnesses, and where the real training leverage sits
08:40 Sovereign compute and the national security case for building models
11:18 Uncensored open weights and the bioweapon question
14:29 Anthropic vs OpenAI, ideology and politics
19:35 The Mythos ban and the Fable 5 delays
24:30 The AGI narrative, the talent drain, and antitrust
28:12 Why researchers join Anthropic, and the open Slack culture
34:04 Nathan's next 12 months: character training and big RL runs
37:55 Continual learning, RSI, and why Nathan is skeptical
43:19 Boom or bubble, tokens vs GPUs
45:12 Why all that token spend never produced 10x products
48:38 Job displacement and the small-business future
52:49 Robotics, world models, and why multimodal lags
57:44 What the open ecosystem should actually do
1:03:17 Why NVIDIA isn't building a frontier model
1:07:34 The RLHF book, and whether RLHF still matters
1:11:06 GRPO vs PPO and on-policy distillation
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. - Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem.
We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn't have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn't happened yet, not a pattern in data you already have.
Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound.
She also gets into what agents are and aren't good for in a wet lab, why cells don't grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution.
Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she'd build if she were starting Coursera today.
Key Topics
The impact of scaling and data in machine learning
The importance of structure and causality in AI
Challenges in drug discovery and biological understanding
The role of foundation models in biology
Ethical considerations in AI and biomedical research
Chapters
00:00 Introduction to Machine Learning and Drug Discovery
02:00 The Bitter Lesson and Its Implications
06:48 Challenges in Drug Design and Discovery
11:48 Ethical Considerations in Human Research
17:20 The Drug Discovery Pipeline Explained
29:30 Integrating AI in Experimental Design
35:38 The Role of Human Judgment in Drug Design
37:14 Future of Drug Design: Efficiency vs. Automation
39:37 Challenges in AI and Data Availability for Biology
41:08 Foundation Models: Potential and Limitations
43:39 Causality in Biological Data: Importance and Challenges
45:18 Creativity vs. Understanding in Drug Design
48:17 Balancing Investments in Data, Algorithms, and Experiments
50:07 The Value of Simulations in Drug Discovery
52:03 Mathematical Frameworks in Biology: Utility and Limitations
54:14 The Future of Drug Discovery: Optimism and Innovations
56:28 The Impact of Coursera on Education
01:00:33 The Role of Universities in Lifelong Learning
01:04:06 Connecting Dots: The Fun of Variety in Work
01:05:46 Optimism for the Future of Drug Discovery
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. - In this episode, Joseph Suarez from PufferAI explains why he thinks RL never had an algorithm problem, but it had a code problem. Every part of the standard RL stack was running about a thousand times slower than it should have been, and once that got fixed, problems that used to take months started getting solved in seconds on one GPU. We talk about what makes a simulator good for RL, why most of their sims run on CPU, what he wants to do with scientific simulation, and why he open sources all of it instead of writing papers.
Key topics
Types of RL and their applications
Challenges in scaling reinforcement learning
The role of simulators and hardware in RL
RL in gaming: from chess to complex games like NetHack and RuneScape
Future directions: scientific simulation and biological modeling
Chapters
00:00 - Introduction to RL and Puff AI
01:50 - Different settings for RL: Games, Robots, Finance
04:10 - RL in LM and other domains
07:00 - Challenges and solutions in RL scaling
09:55 - Building fast, efficient simulators
15:10 - RL for scientific research and simulation
19:57 - RL in complex games: NetHack, RuneScape, Dwarf Fortress
29:55 - Future of RL: Scientific discovery and beyond
Resources
Puff AI - Official Site - https://puffer.ai
NetHack - https://www.nethack.org/
RuneScape - https://www.runescape.com/
Dwarf Fortress - http://www.bay12games.com/dwarves/
OpenAI Gym - https://github.com/openai/gym
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. - Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself.
We get into why he thinks you can't evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number.
He also has a few stories about agents finding their way around the scoring that are worth hearing cold.
Timeline
00:13 Intro
01:00 What evals are for
04:05 Agentic benchmarks
07:10 Kimi K2 and model diversity
08:23 Long-horizon coding tasks
10:29 Building a benchmark
12:15 MirrorCode
14:27 Rubrics and LLM judges
16:30 The cost of expert labelers
17:49 Long runs and variance
19:44 Evaluating the harness
24:29 Chinese labs building CLIs
30:00 More reward hacking
37:45 Tau-bench and economic tasks
39:43 Benchmaxxing and GLM 5.2
45:15 Statistics and cost
47:56 Frontier convergence
52:04 Misuse in open and closed models
55:35 Self-improvement
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.
About
The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.
More Science podcasts
Trending Science podcasts
About The Information Bottleneck
Two AI Researchers - Ravid Shwartz Ziv, and Allen Roush, discuss the latest trends, news, and research within Generative AI, LLMs, GPUs, and Cloud Systems.
Podcast websiteListen to The Information Bottleneck, Radiolab and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


The Information Bottleneck
Scan code,
download the app,
start listening.
download the app,
start listening.


























