The author tested Claude Opus 5 against SlopCodeBench, a coding benchmark that reveals requirements incrementally rather than upfront. Opus 5 achieved a 24% pass rate compared to previous models' 17%, but still failed to complete any full challenge without defects. All models showed increasing code complexity, verbosity, and quality issues over time, with Opus 5 writing 5x more functions than other models, suggesting current AI can't reliably maintain code quality autonomously.