Benchmarking Opus 5 on SlopCodeBench
SlopCodeBenchClaudeOpus 5software engineering benchmarkcodebase qualitymaintenanceverbositycomplexityduplicationcyclomatic complexitysingle-use functionsmodel performance
Author: dhorthy
Date: 7/27/2026
Article Summary:
The article discusses the results of running three Claude models (Opus 4.8, Sonnet 5, and Opus 5) through a subset of SlopCodeBench, a new long-horizon coding benchmark. The models were tested on their ability to maintain codebase quality over time, with Opus 5 getting a 24% strict pass rate.