Full source code: https://github.com/ai-that-works/ai-that-works/blob/main/2025-09-23-evals-for-classification
In this episode of AI That Works, hosts Vaibhav Gupta and Dex, along with guest Kevin Gregory, explore the intricacies of building AI systems that are ready for production. They discuss the concept of dynamic UIs, the challenges of large-scale classification, and the importance of user experience in AI applications. The conversation delves into the use of LLMs for enhancing classification systems, the evaluation and tuning of these systems, and the subjective nature of what constitutes a 'correct' classification. The episode emphasizes the need for engineers to focus on accuracy and user experience while navigating the complexities of AI engineering. The speakers also discuss model upgrades, user feedback, and the importance of building effective user interfaces, emphasizing iterative development and rapid prototyping for chatbot performance evaluation.
Chapters
Introduction to AI That Works
Dynamic UIs and Their Importance
Large Scale Classification Challenges
Building a Robust Classification System
Evaluating and Tuning the Classification Pipeline
Understanding Accuracy vs. Cost in AI
User Experience and Subjectivity in Classification
Final Thoughts and Future Directions
User Experience and Product Design
Error Analysis and Problem Definition
Prompt Engineering and Model Upgrades
Data Collection and Iteration
UI Development and Visualization Challenges
Understanding Data Quality and Test Cases
Building the UI: Time Estimates and Challenges
The Importance of Iteration in Development
Choosing the Right Tools for UI Development
Vibe Coding: A New Approach to Prototyping
Evaluating Model Performance: Insights and Adjustments
Dynamic Evaluation Systems for Chatbots
The Trade-offs of Using LLMs vs. Deterministic Code