AI coding agents can solve increasingly difficult problems. But can they keep solving them after the codebase starts fighting back?
In this episode of AI That Works, Dex and Vaibhav dig into Slop Code Bench, a benchmark designed to test something traditional coding benchmarks often miss: what happens when an agent has to keep extending the same codebase without knowing what requirements are coming next. Instead of judging a single solution in isolation, the benchmark adds new challenges over time and measures whether models can make progress without accumulating defects, regressions, and unmaintainable code.
They break down new results across models including Sol, Fable, Kimi, and GPT, looking at strict pass rates, cost, defects, code complexity, and the surprising relationship between spending more and producing better results. The conversation also moves beyond the leaderboard into a bigger question: what would an AI coding agent actually have to prove before engineers could trust it to work without constant human supervision?
TIMESTAMPS
How Much AI-Generated Code Do You Need to Read?
Episode Overview: Putting Slop Code Bench to the Test
Why AI Coding Benchmarks Matter
AI That Works Unconference Announcement
How Slop Code Bench Tests Real Software Development
A Real Example of Requirements Getting More Complex
Strict Pass Rate vs. Solving the Current Problem
How Sol, Fable, Kimi, and GPT Actually Performed
What Happens If AI Gets the Problem Upfront?
Why Sol's Benchmark Results Don't Match Real-World Experience
Does Spending More Money Produce Better Code?
Why AI Planning May Matter Less Now
Comparing Cost, Defects, and Model Performance
Measuring the Slop AI Leaves Behind
Can AI Recognize When Its Architecture Needs to Change?
Comparing Code Quality Across Models and Harnesses
The Real Cost of Running AI Coding Benchmarks
How Should Slop Code Bench Evolve Next?
What Makes You Trust AI to Work Lights Off?
Three Ways to Make Coding Benchmarks More Realistic
How Do You Avoid Disaster When Humans Stop Reading?
How MCPs and Skills Change Coding Agent Performance
Should You Delete Agent Instructions for Every New Model?
When AI Skills Actually Belong in Your Codebase
Making Claude and Codex Review Each Other's Work
#AIThatWorks #AICoding #CodingAgents #AIEngineering #SoftwareEngineering #SlopCodeBench #CodeQuality #LLMs #DeveloperTools