Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen - #726

Teaching LLMs to Self-Reflect with Reinforcem...

Up next

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, researchers are looking for new ways to keep foundation models improving. In this ...  Show more

How AI Learns to Smell with Alex Wiltschko - #771

In this episode, Alex Wiltschko, founder and CEO of Osmo, joins the show to discuss his goal of giving computers a sense of smell and what it takes to build olfactory intelligence. We explore the science behind smell, from the hundreds of olfactory receptors in the human nose to ...  Show more

Recommended Episodes

Cuttlefish Model Tuning
Data Skeptic

Hongyi Wang, a Senior Researcher at the Machine Learning Department at Carnegie Mellon University, joins us. His research is in the intersection of systems and machine learning. He discussed his research paper, Cuttlefish: Low-Rank Model Training without All the Tuning, on tod ...

  Show more

#130 Mathew Lodge: The Future of Large Language Models in AI
Eye On A.I.

Welcome to episode #130 of Eye on AI with Mathew Lodge. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how rei ...

  Show more

Goodhart's Law in Reinforcement Learning
Data Skeptic

Hal Ashton, a PhD student from the University College of London, joins us today to discuss a recent work Causal Campbell-Goodhart's law and Reinforcement Learning.

"Only buy honey from a local producer." - Hal Ashton

 

Works Mentioned:

...  Show more

Unlocking Language Models: The Power of Prompt Engineering (Ep. 238)
Data Science at Home

Join me on an enlightening journey through the world of prompt engineering. Explore the multifaceted skills and strategies involved in harnessing the potential of large language models for various applications. From enhancing safety measures to augmenting models with domain kn ...

  Show more