
The episode opened with Gemini 4 Argon, Google’s new frontier model currently limited to cybersecurity researchers. The hosts compared its early Artificial Analysis results with Astra, Fable, Opus 5.5 and Sol 6.1, then noticed an unexpected coding result: Sonnet 5.5 ranked above Opus 5.5 and Gemini 4 on the coding-agent index they reviewed.
That led to a deeper discussion about multimodal AI and what it would take for a model to truly understand video. Brian described how his current thumbnail system samples individual frames, while the next step requires understanding expressions, audio, movement and events across time rather than treating each image independently. The conversation also covered Figure’s unusual decision to train its Figure 02 robots to autonomously jump into molten steel during decommissioning.
The second half shifted toward agents. OpenAI’s Decisions API was compared with JEV, while Gareth described Dot interrupting his work to surface an urgent school security email and later notifying him when the situation was resolved. Brian shared how Muse helped surface the used Kia Niro he ultimately purchased. Those examples pushed the hosts into a larger question about AI education: as agents handle more prompting, research and orchestration themselves, should new users still start with traditional prompting skills or learn how to define goals, judge outputs and work with agents instead?
The hosts also discussed the voluntary White House AI safety accord signed by major AI companies and the FTC’s investigation into potential consumer risks from AI systems. Both developments were reported this week. AP News
Key Points Discussed
00:02:01 Gemini 4 Argon Enters The Frontier Model Race
00:04:04 Gemini 4’s Artificial Analysis Results
00:05:34 Gemini 4 Versus Sol On Coding
00:06:15 Sonnet 5.5 Surprisingly Leads The Coding Index
00:08:16 Figure 02 Robots Jump Into Molten Steel
00:15:34 The White House AI Safety Accord
00:21:40 Has Opus 5.5 Already Been Dialed Back?
00:23:39 Gemini 4 And The Future Of Video Understanding
00:29:24 How AI Chooses The Best Video Frame
00:31:47 Why Understanding Video Requires Context Over Time
00:34:33 FTC Investigates AI Risks To Consumers
00:36:15 Chinese Model Distillation And Cybersecurity
00:38:50 OpenAI’s Decisions API Versus JEV
00:41:30 Why Codex Was Slowing Down
00:42:51 Gareth’s Dot Surfaces An Urgent School Alert
00:44:55 Muse Helps Brian Find His Next Car
00:46:49 Should AI Training Still Start With Prompting?
00:49:05 Ethan Mollick And The “Bitter Lesson”
00:52:38 Teaching People To Define Success Instead
00:54:42 Should Skills And Agents Become The New Basics?
00:56:02 Meta Hires MongoDB CEO CJ Desai
00:57:37 Meta’s Reported $4 Billion Data Center Tax Credits
01:02:08 Episode Wrap-Up
The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Gareth Hood, Beth Lyons, Karl Yeh
D'autres épisodes de "The Daily AI Show"



Ne ratez aucun épisode de “The Daily AI Show” et abonnez-vous gratuitement à ce podcast dans l'application GetPodcast.








