The AI model you're using today will eventually be deprecated. The question isn't if you'll have to migrate, it's whether that migration becomes a fire drill or a routine deployment.
In this episode of AI That Works, Dex is joined by producer Kevin to break down the evaluation framework his team uses to survive constant model releases, deprecations, and pricing changes. Instead of relying on public benchmarks, they show how to build production evals that measure the metrics that actually matter: accuracy, latency, and cost.
The conversation covers everything from building golden datasets and replaying production traffic to diffing structured outputs, defining evaluation gates, and deciding when a new model is actually worth adopting. Kevin also demos a lightweight evaluation harness that compares candidate models against your production baseline and makes model swaps dramatically easier.
If you're building AI products that need to survive the next generation of models, not just the current one, this episode is your blueprint.
TIMESTAMPS
Why Model Deprecation Is Becoming an Engineering Problem
Episode Overview
Meet Kevin and Why He's Obsessed with Evals
The Reality of Constant Model Deprecations
The Three Questions Every Eval Harness Should Answer
The Three Metrics Every AI Eval Should Measure
Defining Success Before You Compare Models
Where Great AI Test Cases Actually Come From
Comparing Models Instead of Labeling Everything
Growing Better Test Cases Over Time
How Software Teams Validate Large Migrations
Building a Model Evaluation Harness
Reading Accuracy, Cost, and Latency Together
Visualizing Model Performance
Designing Systems That Expect Model Upgrades
Why Every AI Team Needs Labeled Data
Start With Vibes, Then Build Better Evals
Scaling From One Test Case to Thousands
Structured Outputs Make Better AI Systems
The Biggest Lesson From Production AI
#AIThatWorks #AIEngineering #LLMEvals #ClaudeAI #GPT #Gemini #ArtificialIntelligence #MachineLearning #SoftwareEngineering