Goodhart's Law in Reinforcement Learning

Goodhart's Law in Reinforcement Learning

Up next

Social Choice for Fair Recommendations

Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social ch ...  Show more

News Recommendations

News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib ...  Show more

Recommended Episodes

Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen - #726
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Today, we're joined by Maohao Shen, PhD student at MIT to discuss his paper, “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” We dig into how Satori leverages reinforcement learning to improve language model reasoning ...  Show more

Kevin Zatloukal — Machine Learning And Its Applications (EP.68)
Infinite Loops

<span style="color: #333333;">Kevin is a Computer Scientist, part-time lecturer at Paul G. Allen School of Computer Science & Engineering, and research partner at OSAM. Our discussion with Kevin revolves around:
</span>

<ul> <li><span style="color: #333333;">What is M ...  Show more

🧠 Scientific Machine Learning, FEM + ML, PINNs – Ehsan Haghighat | Podcast #79
The Engineered-Mind Podcast | Engineering, AI & Technology

Dr. Ehsan Haghighat is a Postdoctoral Fellow at UBC studying stochastic modeling and uncertainty quantification of engineering systems. Previously, he was a Postdoctoral Associate at MIT where he studied the assessment of induced seismicity due to CO2 sequestration and oil and ga ...  Show more

Applications of Variational Autoencoders and Bayesian Optimization with José Miguel Hernández Lobato - #510
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Today we’re joined by José Miguel Hernández-Lobato, a university lecturer in machine learning at the University of Cambridge. In our conversation with Miguel, we explore his work at the intersection of Bayesian learning and deep learning. We discuss how he’s been applying this to ...  Show more