
Why do businesses replace the AI model when the failure may have started somewhere else entirely?
In this episode of Tech Talks Daily, I speak with Richard Shaw, Technology General Manager for Databricks in the UK and Ireland. Richard leads the field engineering organization that works closely with customers on data and AI problems, giving him a practical view of what happens when promising agentic AI projects meet production workloads.
Richard argues that the model often receives the blame because it is the most visible part of the system. The actual fault may come from stale data, missing business context, inconsistent permissions, an unsuccessful tool call, or another point in the workflow. Replacing the model before tracing the request from start to finish can recreate the same problem in a new place. This is why lineage, end-to-end tracing, and continuous evaluation matter once an agent moves beyond a controlled pilot.
We discuss what a production-readiness rehearsal should include. Richard recommends realistic data, realistic user volumes, unauthorized requests, ambiguous questions, failed tool calls, and tests of what the agent should refuse to do. Teams also need agreed standards for quality, security, cost, and auditability, along with a clear decision about which actions an agent can complete independently and where a person must review or approve the result.
The conversation also looks at model choice and infrastructure cost. Richard believes the strongest test is performance on the organization's actual work rather than a benchmark leaderboard. A frontier model may suit complex reasoning, while a smaller or open-weight model may perform routine extraction or classification at a lower cost. Access policies, observability, and spend controls need to remain consistent as those model choices change.
Conversational analytics creates another governance challenge. Databricks customers such as Virgin Atlantic and Repsol are using natural-language tools to make company data easier for employees to question. Richard says wider access should preserve existing permissions, ownership, definitions, and lineage. An answer becomes far more useful when the user can see where it came from and which team owns the information behind it.
We also cover the boundary between historical analytical data and fast operational workloads. Richard describes how Databricks positions the lakehouse for broad enterprise context and Lakebase for immediate reads and writes, such as updating an account, placing an order, or storing agent memory, while keeping both connected to a common data and governance base.
Are companies ready to trace and test the whole AI workflow, or are too many treating the model as both the hero and the culprit? Listen to the episode and share your thoughts with me.
Altri episodi di "Tech Talks Daily"



Non perdere nemmeno un episodio di “Tech Talks Daily”. Iscriviti all'app gratuita GetPodcast.








