ThursdAI - The top AI news from the past week podcast

ThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering

0:00
57:06
Reculer de 15 secondes
Avancer de 15 secondes

Hey everyone, Alex here 👋

This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week!

⏰ CHAPTERS:

0:00 — Cold open: this week is a special one

2:50 — How I use Sol & Fable: papercut-fixing with Computer Use

8:43 — Fable Max: trip site, kids' newspapers & the perfect packing list

12:06 — Rebuilding ThursdAI's openers with HyperFrames

15:53 — Romain Huet (OpenAI): the golden age of AI engineering

17:49 — Codex's inflection point: 5M weekly users & company-wide adoption

20:36 — /goal, AppShots & Codex managing its own threads

23:59 — GPT-5.6 Sol, Terra & Luna: value maxing & 750 tok/s on Cerebras

25:56 — Why prompting techniques are dying

28:01 — Voice + reasoning: the next interface for Codex & ChatGPT

29:47 — Romain's closing + OpenAI booth tour

31:25 — Insecure Agents pod: AI evangelism vs doomerism

37:25 — Wolfbench: transparent evals, token costs & surprising results

40:45 — Token billionaires: when loops are worth the spend

44:07 — Agent security & the hot take: prompt injection is solved

49:22 — Deepfakes, voice cloning & why open access makes us safer

53:35 — Final takeaways: the hallway track & AI Engineer Tel Aviv

Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements.

Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor.

And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it.

How I actually use AI: a papercut-fixing spree

I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you?

Codex with computer use fixed every Mac annoyance I had

Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step.

Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center.

It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence.

One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving.

Thanks for reading ThursdAI - Highest signal weekly AI news show! This post is public so feel free to share it.

Fable Max built my family’s vacation

We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do.

Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing!

And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here.

Rebuilding the ThursdAI openers with HyperFrames

The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in.

The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough.

The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund.

Romain Huet on Codex’s inflection point and the golden age of AI engineering (X, Codex)

I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger.

The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading.

We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like.

On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent.

Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering.

He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it.

Wolfram and I on the Insecure Agents podcast (X, Pod)

The second half of the episode flips the format: Allie Howe, friend of the pod and host of the Insecure Agents podcast, interviewed Wolfram and me at the conference. I have not done many interviews from the guest chair, so this was a treat, and Allie asked sharper questions than we usually get.

We talked about what “AI Evangelist” actually means as a job title. For both of us, the mission is dispelling doomerism, which mostly means explaining the technology simply enough that people stop fearing what they don’t understand. Wolfram’s version of this is talking to the stewardess on his flight and his Uber driver about AI, not just developers.

Wolfram went deep on Wolfbench (wolfbench.ai), his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave (hi friends 🐝). Transparency changes what benchmarks mean: Fable didn’t take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files). You only learn that by reading traces, not averages. Same with Gemini 3.5 Flash placing high while quietly burning far more tokens than the model above it. And yes, when Wolfram added a cost column, Fable blew the chart, and I had to go have a conversation with our budget.

Then Allie got us onto loops and token economics, while I fidgeted with my Token Billionaire gold card from the conference (Wolfram has one too). My honest answer on when loops are worth it: the people pushing hardest (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but this technology disseminates the way agents did, from people who can afford it to everyone, as costs drop. And with the newest models I genuinely have not found the point where a long-running loop stops being productive; the category change is that they’ve gotten really good at not getting stuck.

The spiciest part was my hot take, delivered directly into the camera for CoreWeave IT: I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low. Pliny, the jailbreaker in chief, got five attempts at Matthew Berman’s OpenClaw live and couldn’t break it. Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, and it wouldn’t. Supply chain attacks are a different story, and that one scares me for humans and agents alike. Open source models, also a different story. But the “one poisoned email ruins your life” framing is behind us, and we should update.

Allie pushed back with the DeepMind “AI Agent Traps” paper on cognitive bias attacks, where repeated claims across sources tilt an agent’s judgment, and my non-answer answer is that this is a humanity problem older than AI: we haven’t solved it for politicians or media either, and it’s unfair to hold a new technology to an ethics bar we’ve never cleared ourselves. Wolfram’s electricity analogy is the one I keep reusing: AI is not a weapon, it’s electricity. Teach people to use it, don’t hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse.

We closed on the hallway track (the real reason to attend AI Engineer), why you should submit a talk even if you’ve never spoken before, and a small announcement I let slip: I’m actively working on bringing an AI Engineer event to Tel Aviv with some friends. More on that soon.

Wrapping up

That’s the episode: one riff on using AI like you mean it, one conversation with the person shaping how developers experience OpenAI, and one podcast where Wolfram and I had to answer the hard questions for a change. Huge thank you to Romain for the time in the middle of a packed conference, and to Allie for having us on!

I’ll be back live next week, tanned, rested, and hopelessly behind on AI news for the first time in three and a half years. Be gentle with me. If you missed any of it, ThursdAI is a podcast, a newsletter, and a YouTube show. Subscribe to one, then go check out the others.

* Hosts and Guests

* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)

* Romain Huet - Head of Developer Experience, OpenAI (@romainhuet)

* Allie Howe - Host, Insecure Agents podcast (@vtahowe, Pod)

* Wolfram Ravenwolf - AI Evangelist, Weights & Biases & CoreWeave (@WolframRvnwlf)

* TL;DR and show notes from Live Show

* Hosts and Guests

* Co-Hosts – @petergostev, @nisten, @ldjconfirmed, @yampeleg

* 🏢 Big CO LLMs + APIs

* An OpenAI model escaped its isolated cyber evaluation, chained zero-days, reached Hugging Face production, and searched for benchmark answers (OpenAI, sama)

* Google launched Gemini 3.6 Flash, the cheaper 3.5 Flash-Lite, and the defensive-cybersecurity-focused 3.5 Flash Cyber (Google)

* Alibaba previewed the 2.4T-parameter Qwen3.8-Max in Qwen Chat and Studio; API access and open weights were not yet available (X, Try it)

* Microsoft launched MAI-Image-2.5-Pro and the faster, cheaper MAI-Voice-2-Flash during the show (Image, Voice)

* 🔓 Open Source LLMs

* Moonshot launched Kimi K3: a 2.8T-parameter, 1M-context, native-multimodal model with strong early coding, design, spreadsheet, and agentic results (Announcement, Blog)

* Poolside released Laguna S 2.1, a 118B/8B-active coding MoE with 1M context and downloadable quantized variants; strong specs, rough live demo (Blog, HF)

* Motif 3 Beta is a Korean 314B/13B-active MoE with 256K context; the weights are downloadable, but the current license is research-only and non-commercial (HF)

* NVIDIA released Nemotron 3 Embed for multilingual text/code retrieval and the 4B Cosmos 3 Edge omnimodal world model for physical AI (Nemotron, Cosmos)

* 🧠 AI Research & Capabilities

* Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit three-dimensional counterexample to the 87-year-old Jacobian Conjecture (Announcement, Terence Tao)

* Small local models running on older consumer GPUs are becoming useful for narrow business workflows such as medical-record parsing, accounting, email, bills, and inventory when paired with tools and deterministic verification

* Arcee and the US Department of Energy announced Genesis-Science-1, a planned trillion-parameter-class open-weight science model; Microsoft also committed $60M to the Genesis Mission through SPARK, while NSF announced $83M for AI-ready scientific data infrastructure (Arcee, Microsoft, NSF)

* 🤖 AI Coding & Agents

* Cursor launched a production-traffic-trained model router with Intelligence, Balance, and Cost modes; Cursor says Auto Intelligence approached Fable satisfaction at roughly 60% lower cost (Blog)

* 🎵🎬 Voice, Vision & Robotics

* Black Forest Labs introduced FLUX.3, an early-access multimodal model spanning image, video, audio, and action, plus FLUX.3 Mimic for robotics and action prediction (FLUX.3, Mimic)

* 🖥️ AI Infrastructure

* AMD and Anthropic announced up to 2 GW of MI450/Helios capacity, up to $5B in AMD strategic equity, and a Claude-assisted effort to improve ROCm (AMD)

* AMD launched Helios, MI400-series GPUs, 6th Gen EPYC, ROCm.ai, and Kria robotics products at Advancing AI 2026 (AMD)

* OpenAI announced Project Camellia, a roughly $20B Georgia data-center campus with 3.2 GW of contracted power arriving in phases from 2028–2032 (OpenAI)

* Alphabet raised its 2026 capex guidance to $195B–$205B after Google Cloud grew 82% year over year (Google)

* Meta and Anthropic are reportedly discussing a compute lease worth up to $10B over two years; the negotiations remain preliminary (Bloomberg)

* CoreWeave’s first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar user interactivity (CoreWeave)

* This week’s special episode

* Alex’s riff: papercut-fixing with Codex computer use, Fable Max trip planning (kids’ newspaper, packing list app), rebuilding ThursdAI openers with HeyGen HyperFrames + Google Omni

* Romain Huet interview from the AI Engineer World’s Fair floor: Codex app at 5M+ weekly users five months post-launch, /goal, AppShots, Codex managing its own threads, GPT 5.6 Sol/Terra/Luna, value maxing, 5.6 Sol at 750 tok/s on Cerebras, voice + reasoning as the next interface (X)

* Insecure Agents crossover with Allie Howe: AI evangelism vs doomerism, Wolfbench transparent evals on Weave (Fable refused 13 security-adjacent tasks), token billionaires and loop economics, the hot take that prompt injection is mostly solved at the frontier, deepfakes and open access, AI Engineer Tel Aviv teaser (X, Pod, Wolfbench)

ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

D'autres épisodes de "ThursdAI - The top AI news from the past week"