Benchmarking Opus 5 on SlopCodeBench

Other: software engineering benchmark(github.com)view on HackerNews
SlopCodeBenchClaudeOpus 5software engineering benchmarkcodebase qualitymaintenanceverbositycomplexityduplicationcyclomatic complexitysingle-use functionsmodel performance

Author: dhorthy

Date: 7/27/2026

Article Summary:
The article discusses the results of running three Claude models (Opus 4.8, Sonnet 5, and Opus 5) through a subset of SlopCodeBench, a new long-horizon coding benchmark. The models were tested on their ability to maintain codebase quality over time, with Opus 5 getting a 24% strict pass rate.