TLDR overview
- GPT-6 Sol passes 83.09% of the 544 HumanEval and MBPP tasks with executable tests, against 85.85% for Astra.
- Sol generated 480,125 lines of code against Astra's 656,445, which is 26.9% less for the same benchmark.
- Total findings fell 21.8%, from 12,606 to 9,859. Bugs came down 17.5% and code smells 22.3%.
- Vulnerabilities went the other way. The count rose 17.1%, from 117 to 137, and density rose 60% to 285 per mLOC.
- Blocker vulnerability density is 2 per mLOC, the lowest figure we have recorded from this family. In absolute terms that is a single blocker vulnerability across 480,125 lines.
