Skip to Content

Diana Pfeil: Building Confidence in Probabilistic Systems

EP-224 | August 4, 2026 | 48:56

Machine learning systems can degrade even when the underlying code has not changed. Diana Pfeil of Sunbeam Consulting joins Robby Russell on Maintainable to explain how changing data, model behavior, and non-deterministic outputs create a different kind of maintenance challenge.

Rather than asking whether a system is simply correct or incorrect, teams need reliable ways to measure confidence in its behavior.

Diana introduces evals as a way to test AI-generated outputs that may be phrased differently each time. She and Robby discuss using LLMs to judge other LLM outputs, reviewing production traces, sampling unusual cases, protecting sensitive data, and keeping humans involved when automated checks cannot provide enough confidence.

They also explore how prompts should be versioned and tested against representative examples.

The conversation turns to the operational costs that come with AI features. Model providers can deprecate dependencies quickly, prompts may behave differently after an upgrade, and teams must continue monitoring systems that might once have been considered finished.

Diana encourages organizations to ask what can now be automated, how accurate the result needs to be, and whether the benefit justifies the additional maintenance work.

Diana also makes the case for starting with the simplest model or deterministic process that can solve the problem. Teams can add complexity once the baseline proves insufficient, but designing around imagined future requirements often creates the wrong system.

Her advice for engineers trying to introduce AI internally is equally direct: build a small prototype that solves a real problem, then let the result make the case.

Episode Highlights

[00:00:50] Maintaining Probabilistic Software: Diana explains why maintaining AI systems involves the code, changing data, model behavior, and confidence in the output.

[00:02:35] What Are Evals?: Robby asks how teams test AI-generated results when the correct response may be worded differently each time.

[00:04:38] Using an LLM as a Judge: Diana describes using one model to evaluate another and why the judge must be calibrated against human decisions.

[00:06:39] Monitoring AI in Production: Diana introduces human review, traces, observability, and production sampling.

[00:08:44] Recognizing Input Drift: A meeting-notes example shows how changing inputs can degrade an otherwise unchanged system.

[00:12:06] Maintaining Models and Prompts: Diana outlines the code, model, prompt, and data-pipeline changes teams may need to make.

[00:14:24] Versioning and Testing Prompts: Robby asks how prompt experimentation fits into source control and repeatable testing.

[00:18:32] Building Confidence in a Black Box: Diana explains how evals and production reviews help teams avoid regressions.

[00:21:16] Diana’s Machine Learning Background: Diana shares her path from recommendation systems at Amazon to startup leadership and consulting.

[00:22:26] How Sunbeam Consulting Helps Teams: Diana describes advising leaders on AI strategy and helping teams build machine learning products.

[00:24:28] Finding Useful Automation Opportunities: Diana explains how teams can identify previously unstructured work that may now be practical to automate.

[00:27:40] The Cost of Automated Decisions: Robby and Diana compare human error with the oversight and infrastructure required by AI systems.

[00:30:14] AI Is Not Free to Maintain: Diana discusses model deprecations, vendor dependencies, and the ongoing support required after launch.

[00:36:45] Keeping Up With Rapidly Changing Tools: Diana explains why teams need room to experiment without constantly disrupting established workflows.

[00:38:49] Why Simpler Models Often Win: Diana makes the case for starting with a baseline before introducing more sophisticated approaches.

[00:45:18] Selling an AI Idea Without the Buzzwords: Diana recommends building a useful prototype and allowing the result to make the case.

[00:46:30] The Inner Game of Tennis: Diana recommends W. Timothy Gallwey’s book about learning, judgment, and performance.

Resources Mentioned

Thanks to Our Sponsors!

Your test coverage says 90%, but that might be misleading. Undercover CI looks at your Ruby pull requests and shows you which parts of your changes weren't tested- not just overall coverage, but what changed and what got missed, down to the method level. Visit undercover-ci.com and use code MAINTAINABLE for 15% off your first billing cycle. Free for public repos. Private repos with unlimited users also available.

Mailtrap is a modern email delivery platform built for developers. Native SDKs, a secure Email API and SMTP, and a free tier with 4,000 emails a month. When you need help, you'll reach real people on 24/7 support, not an AI chatbot. Try Mailtrap for free!

🎧 Listen from Anywhere 🪐

Listen on all the major podcast platforms.

Between the episodes

229 Episodes published since 2019

Stay sharp. Skip the noise.

One email when a new episode drops. That's it.

Joined by engineering leaders at companies you've heard of.