Agentic Conversations (formally mlops.community) podcast

The Caveman Prompting Challenge

0:00
44:07
Rewind 15 seconds
Fast Forward 15 seconds

Caveman prompting has one rule: why use many words when few do the trick? It saves tokens on the way in and on the way out. Push it too far, though, and the output falls apart. So how far is too far? Nobody has benchmarked it yet, and that question opens our conversation with James Barney, Head of Forward Labs at MetLife.


James spends his days connecting new AI capabilities to old business problems across dozens of regulatory regimes, and he still finds time to push code. He explains how the FinOps Foundation's AI working group took on the most basic question: which model for which workload, and why the answer always comes down to cost, speed, and accuracy. We get into Anthropic's launch pricing for Fable, why a million tokens is easy to price and hard to explain, and why every stakeholder eventually tells you what they really care about once you name the wrong North Star.


From there it gets practical. Start with the smartest model, then step down and add harness until quality holds. Treat exploration tokens like local builds and production tokens like pipelines. Govern agents the way you govern people, with proactive blocks, reactive checks, and policy as code an agent can actually read.


We close on a bigger shift. When a chat window can pull from every dashboard at once, do we still need dashboards? James thinks mostly not, with one catch he calls the latent shopper problem: some insights only come from browsing data you did not know to ask about.


Demetrios Brinkmann: https://www.linkedin.com/in/dpbrinkm

James Barney: https://www.linkedin.com/in/james-barney


Timestamps:

[00:00] Cold open

[01:03] Meet James from MetLife

[01:11] What AI enablement means at a global insurer

[03:27] Inside the FinOps Foundation AI working group

[05:03] The caveman skill challenge

[06:37] Why we need a CaveBench

[07:43] Anthropic's Fable launch pricing

[09:34] Planning for surprise model releases

[11:45] Measuring AI value beyond cost

[12:47] Finding your unit metric

[14:15] Tying token spend to business outcomes

[16:35] Explore first, then optimize

[18:12] AI that tunes itself

[20:13] Do R&D tokens count

[23:24] When the experiment becomes the product

[25:38] Personal agents and daily briefings

[27:39] Governing agents that run just because they can

[30:37] The layers of AI governance

[31:56] Proactive and reactive guardrails

[32:59] Policy as code for agents

[34:41] Stop the click ops

[35:31] Is the modern UI obsolete

[36:36] MCP apps and chat as the new browser

[38:07] How AI gathers data differently than humans

[40:37] The latent shopper problem

[42:12] Staying close to your data

More episodes from "Agentic Conversations (formally mlops.community)"