In this episode, we break down where latency really comes from in AI systems and how to make agents feel faster. The conversation covers UI tricks, prefetching, prompt caching, parallelism, and streaming — plus how reasoning tokens and hidden bottlenecks affect real user experience.
A practical discussion for anyone building AI agents, tools, or user-facing AI products that need to feel fast and responsive.
Chapters:
Intro
Latency Issues
Faster UI
Prefetching
Prompt Caching
Parallelism
Reasoning Tokens
Streaming UX
Final Tips
Check out our github: https://www.github.com/boundaryml/baml
AI That Works repo: https://github.com/ai-that-works/ai-that-works
Socials:
X: https://x.com/boundaryml
Discord: https://discord.com/invite/yzaTpQ3tdT
LinkedIn: https://www.linkedin.com/company/boundaryml/
Correction:
Title of the episode is Understanding Latency, editing mistake!