Crazy Wisdom podcast

Episode #574: Evals, Ontologies and the Unmappable World of Business

28/9/2026
0:00
47:36
Retroceder 15 segundos
Avanzar 15 segundos

In this episode of the Crazy Wisdom Podcast, host Stewart Alsop sits down with Ryan Marsh of thestack.io to explore what it really takes to build production AI systems. They discuss how production AI has evolved from simple prompt-to-API demos into complex systems requiring evaluation suites, human-in-the-loop feedback mechanisms, and sophisticated approaches to handling context and data retrieval. The conversation covers the challenges of domain mapping, the fundamental difficulty of translating messy human business processes into structured systems, and the role of ontologies in AI development. Ryan and Stewart also examine the limitations of LLMs, the debate between specialization versus generalization in AI models, consciousness and cognition in system design, and the regulatory landscape facing AI companies. They touch on infrastructure constraints, the democratization of AI through open source models, and whether we'll eventually hit a ceiling where human intelligence can no longer distinguish between increasingly capable AI models.

Key Insights

1. Production AI systems today fundamentally differ from demos through their reliance on comprehensive evaluation suites that function like unit tests to measure and maintain quality, though they cannot be as deterministic. The key distinction is that production systems require clearly defined metrics for what good looks like, along with feedback mechanisms that allow the system to evolve over time. Without this foundational understanding of success metrics and continuous improvement processes, a system is not truly production-ready regardless of how many users it serves.

2. The fundamental challenge in building production AI systems is not the technology itself but rather mapping business domains into structured formats that models can work with effectively. This problem of translating messy, subjective human processes and language into precise specifications has plagued software engineering for decades. Different people within organizations use the same words to mean completely different things, and humans naturally operate with assumed context and imprecision that must be explicitly defined for AI systems to function reliably.

3. Large language models excel at generalization but struggle with specialization, which creates friction in production environments where specific outputs or styles are required. While they can code in any programming language, getting them to write code exactly the way a particular engineer wants remains extremely difficult. This explains why professional documentation and specialized coding tasks often require extensive prompting and fighting with the models, as they naturally gravitate toward their trained patterns rather than highly specific user preferences.

4. Modern production AI systems primarily solve classification problems wrapped in natural language interfaces rather than requiring true open-ended cognition. The models work best when they can leverage reasoning over provided information to make verifiable decisions, but they still lack common sense despite their vast knowledge. Success comes from teaching models everything about your specific domain and what good and bad outcomes look like, rather than relying solely on their general intelligence.

5. Context retrieval in production AI systems is fundamentally a data storage, search, and retrieval problem that has been solved many different ways throughout computing history. The appropriate solution depends entirely on the type of data being accessed, whether through vector databases, graph databases, relational databases, or even simple text search. The harnesses and frameworks for orchestrating AI agents have matured significantly, making the real challenge the quality and structure of the data being fed to these systems.

6. Human-in-the-loop feedback mechanisms are essential for production AI because models will inevitably encounter situations they have not been trained to handle. When confidence is low or novel scenarios appear, systems should flag these for human review rather than proceeding blindly. The feedback provided during these interventions must be captured and scored so it can be incorporated into the permanent behavior of the system, creating a continuous improvement cycle similar to how model vendors perform reinforcement learning on their base models.

7. The rush to regulation in the AI industry is driven primarily by the fact that these companies are currently completely exposed to existing consumer protection and liability laws with no legal precedents to protect them. If AI agents cause harm, companies could be sued into oblivion under current law. By establishing compliance frameworks through regulation, these companies can create carve-outs and exemptions that limit their liability when they follow prescribed rules, similar to how heavily regulated industries like airlines and banking have become nearly impossible to enter due to compliance requirements.

Timestamps

00:00 Welcome and introduction to Ryan Marsh discussing production AI systems and what companies need to understand about building them at scale

05:00 The cognitive load challenge of working with invariants and probabilistic systems, discussing how LLMs function like PhD students without common sense

10:00 Building eval suites to measure AI performance, handling the long tail of edge cases in complex contracts using human-in-the-loop feedback systems

15:00 Context as a data retrieval problem and the fundamental challenge of mapping business domains when humans struggle with imprecision and ambiguous language

20:00 How different departments use the same words with different meanings and why LLMs generalize well but don't specialize effectively for specific coding styles

25:00 The ontology debate and intractable problem of mapping subjective human systems to rigid structured formats with diminishing returns on perfect mapping

30:00 Why humans struggle distinguishing what is from what ought to be and the challenge of creating SOPs when companies lack updated documentation

35:00 The liability exposure AI companies face and why they're begging for regulation to protect themselves from existing consumer protection laws

40:00 Open source knowledge transfer between Chinese and American labs, chips and power as the real constraint not algorithms for frontier models

45:00 Reaching intelligence ceiling where specialists can't distinguish state-of-art models and fundamental physical laws limiting LLM scaling through layered optimization strategies

Links

Website
X

Otros episodios de "Crazy Wisdom"