Back to index
OtherMultiple episodes weekly, 60 to 120 minutes·Free

The Cognitive Revolution

4.1

The most substantive long form AI podcast I know. Labenz does more genuine hands on model testing than any other host and the conversations reflect it, at the cost of very long episodes and a heavy release schedule.

What We Liked

  • Labenz actually uses and stress tests the models he discusses, in depth
  • Takes both capability and safety seriously without picking a tribe
  • Guests include builders and researchers who go into real technical detail
  • Willing to sit with uncertainty rather than manufacture a conclusion
  • Volume of output means broad coverage of the field

What Could Be Better

  • Episodes are extremely long and could often be half the length
  • Release rate is impossible to keep up with
  • The host talks a great deal, sometimes more than the guest
  • Torenberg's venture background pulls some episodes towards startup framing
  • No structure, so a newcomer has no idea where to start

Detailed review

Nathan Labenz has an unusual qualification for an AI podcast host, which is that he was on the OpenAI red team for GPT-4 before release and has spent an enormous amount of time since then systematically probing what frontier models can and cannot do. When he describes a model's behaviour, he is describing something he tested rather than something he read about. That single fact makes this show more useful than almost everything else in the category. The depth is real.

Episodes run an hour to two hours and use the time. Guests get to explain their work properly, including the technical parts, and Labenz asks follow ups that require having understood the previous answer. Conversations go into architecture, evaluation methodology, specific failure modes, deployment realities. Very little of the runtime is spent on framing questions or personal narrative.

The intellectual position is what I respect most. The AI discourse has largely sorted itself into camps. One says these systems are dangerous and progress should slow. The other says the risk talk is a distraction and capability is what matters.

Labenz refuses both. He takes safety concerns seriously as technical problems, he takes capability seriously as a fact about the world, and he lets guests from either position make their case without turning the episode into a fight. That combination is rare and it makes the show a better guide to what informed people actually disagree about. He is comfortable with not knowing.

When a question is genuinely open, the episode ends without resolving it. That sounds like a small thing and it is unusual in a media environment that rewards confident conclusions. The guest range is wide. Frontier lab researchers, safety people, founders building applications, academics, people working on evaluation and interpretability.

The application focused episodes are more useful than they sound because Labenz asks what the system actually does rather than what it is claimed to do. Now the problems, and the main one is volume. Episodes are very long and there are a lot of them. Several a week, sometimes more, each one to two hours.

Nobody can listen to all of it. The feed produces a low grade guilt that makes the show harder to engage with than it should be. I have concluded the only sane approach is to treat it as an archive and pick by topic, ignoring the release schedule entirely. Editing would help enormously.

Many episodes contain a genuinely excellent forty five minutes inside a hundred and twenty minute conversation. Digressions run long, points get restated, and the discipline that would make these tighter is not applied. I listen at high speed and still skip. Labenz talks a lot.

His preambles to questions are frequently several minutes and sometimes contain the answer he is hoping for. When the guest is the more interesting party, which is often, this is a real cost. The best episodes are the ones where he asks a short question and gets out of the way. Erik Torenberg co-hosts and comes from venture capital, and the episodes with his framing tilt towards market opportunity and startup strategy.

Those are the weakest ones for my purposes. They are also clearly labelled by topic and easy to avoid. There is no entry point. No introductory episode, no glossary, no sequence.

The show assumes you already follow the field. Someone new to AI would find most of this incomprehensible, and that is a genuine gap given how much good material is in here. My four point one is for the most technically substantive and intellectually honest long form AI show available, from a host who does more real testing than anyone else in the category, marked down for episodes that badly need editing, a release rate that defeats engagement, and a host who could speak less. Choose carefully and it is the best thing in this space.

[ final ]

The verdict.

The best podcast available if you want depth on what current AI systems can actually do. Ignore the feed, choose episodes by topic, and accept that you will never be caught up.