Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen - #726

Teaching LLMs to Self-Reflect with Reinforcem...

Up next

Why Jev Is Changing How We Build With AI with Diogo Almeida - #779

In this episode, Diogo Almeida, co-founder and CEO of TypeSafe, joins us to discuss Jev, TypeSafe’s recently released model for bringing fast, reliable intelligence directly into software. We explore the idea of “machine-native intelligence” and why Diogo believes models optimize ...  Show more

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778

AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes. In this episode, Greg Burnham, who leads AI capabilities research at Epoch AI, joins us to examine what that prog ...  Show more

Recommended Episodes

Cuttlefish Model Tuning
Data Skeptic

Hongyi Wang, a Senior Researcher at the Machine Learning Department at Carnegie Mellon University, joins us. His research is in the intersection of systems and machine learning. He discussed his research paper, Cuttlefish: Low-Rank Model Training without All the Tuning, on tod ...

  Show more

#130 Mathew Lodge: The Future of Large Language Models in AI
Eye On A.I.

Welcome to episode #130 of Eye on AI with Mathew Lodge. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how rei ...

  Show more

Nature of Intelligence, Ep. 6: AI’s changing seasons
COMPLEXITY

Guest: Melanie Mitchell, Resident Professor, Santa Fe InstituteHosts: Abha Eli PhobooProducer: Katherine MoncurePodcast theme music by: Mitch MignanoFollow us on:Twitter • YouTube • Facebook • Instagram • LinkedIn • BlueskyMore info:Tutorial: Fundamentals of Machine LearningLectu ...  Show more

Goodhart's Law in Reinforcement Learning
Data Skeptic

Hal Ashton, a PhD student from the University College of London, joins us today to discuss a recent work Causal Campbell-Goodhart's law and Reinforcement Learning.

"Only buy honey from a local producer." - Hal Ashton

 

Works Mentioned:

...  Show more