In this episode of AI That Works, we tackle the big question: what is harness engineering, and how can you build harnesses that genuinely push AI performance? We cut through the noise, discussing whether a 'dumb' model with a great harness can outperform a 'good' model with a poor one.
We explore the advantages model labs have in crafting superior harnesses, delving into the intricacies of RL, tool calling (token-wise vs. constrained), and how models are specifically trained for the harnesses they'll operate within. From recursive types to outer versus inner harness orchestration, we break down complex concepts. Our goal is to move beyond mere demos, helping you understand where the real 'alpha' lies in this space. Join us as we dissect the latest hype and share practical insights to get AI truly working for you.
Check out our github: https://www.github.com/boundaryml/baml
AI That Works repo: https://github.com/ai-that-works/ai-that-works
Socials:
X: https://x.com/boundaryml
Discord: https://discord.com/invite/yzaTpQ3tdT
LinkedIn: https://www.linkedin.com/company/boundaryml/
⏱️ Chapters:
Intro
Model Training & Tooling
Custom Tool Call Formats
Benchmarks & Agent Training
Preventing Prompt Leaks
Harness Optimization Strategies
Counteracting Harness Spoofing
Finding Engineering Alpha
Personal LLM Preferences