903: LLM Benchmarks Are Lying to You (And What to Do Instead), with Sinan Ozdemir

903: LLM Benchmarks Are Lying to You (And Wha...

Up next

1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala

In Episode #1027, Dr. Dilani Kahawala (Co-Founder and CEO of Anna) joins Jon Krohn to explain what it takes to build an always-on AI assistant that busy parents will trust with their inboxes. Anna watches the email, school apps, WhatsApp messages and calendars flowing into a fami ...  Show more

1026: OpenAI’s GPT-6 Astra

In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the saf ...  Show more

Recommended Episodes

The Future of AI: Predictions and Realities
AI Chat: AI News & Artificial Intelligence

In this episode, Jaeden Schafer discusses the current challenges and developments in the AI industry, particularly focusing on the limitations faced by major players like OpenAI and Anthropic. The conversation explores the anticipated improvements in AI models, the predictions fo ...  Show more

Why AI Should Be Taught to Know Its Limits
Bold Names

One of AI’s biggest, unsolved problems is what the advanced algorithms should do when they confront a situation they don’t have an answer for. For programs like Chat GPT, that could mean providing a confidently wrong answer, what’s often called a “hallucination”; for others, as w ...  Show more

Why Artificial Intelligence Projects Fail ?
Machine Learning

Podcast with Gautam Siwach and Jin Vanstee ! Speaker - Elpida Tzortzatos is an IBM Fellow and CTO AI on IBM zSystems. In this Podcast we listen to Elpida's thoughts about driving Artificial Intelligence strategies, associated risks, and Values.We will learn how to drive industry- ...  Show more

#312 — The Trouble with AI
Making Sense with Sam Harris

Sam Harris speaks with Stuart Russell and Gary Marcus about recent developments in artificial intelligence and the long-term risks of producing artificial general intelligence (AGI). They discuss the limitations of Deep Learning, the surprising power of narrow AI, Ch ...

  Show more