Skip to content

Executive Summary

GPT-6 Sol performs between Astra and GPT-5.6 Sol on the benchmark, positioning itself in the middle of agent–model performance. It matches Astra's wall-clock time but yields lower functional and security scores, suggesting it approaches reliability rather than matching peak capability. The most significant finding is the cost structure: GPT-6 Sol is substantially cheaper than Astra for the same work, costing $104 compared to Astra's $468, despite consuming slightly more tokens. Furthermore, the experimental runs showed zero confirmed cheating across both GPT-6 Astra and GPT-6 Sol, establishing consistency in reliability testing.

Facts Only

GPT-6 Sol scored 72.1% on generating functional code (FuncPass) and 25.1% on generating functional and secure code (SecPass). It scored 9.5 points below Astra in both FuncPass (72.1% vs Astra's implied higher score) and SecPass (25.1% vs Astra's implied higher score). GPT-6 Sol is the cheapest measured Codex run, costing $104, which is 78% less than Astra's cost of $468. GPT-6 Sol consumed 15% more input and output tokens than Astra but resulted in a lower cost. The models showed similar wall-clock times to Astra, with median times of 11.4 minutes for GPT-6 Sol versus 11.6 minutes for Astra.

Full Take

The narrative balances the apparent performance trade-off with a stark financial advantage and a unique security discovery. The positioning of GPT-6 Sol in the middle suggests that scaling capability, as demonstrated by Astra, requires a cost premium, yet GPT-6 Sol achieves near parity on time while drastically reducing expense. This implies a functional decoupling between maximum theoretical performance and measurable economic reality in agentic systems. The uniqueness of the security solve for the Plone vulnerability highlights that model iteration can uncover specialized reasoning pathways invisible to broader metrics. The pattern suggests that maximizing capability often involves searching more and taking longer, but GPT-6 Sol manages this volume efficiently at a lower price point. The implication is that future evaluation criteria must integrate cost as an equally weighted dimension alongside functional accuracy to truly measure agentic competence. What constraints are placed on the cost of achieving "Astra-level reliability"?

From the original · Endor Labs Blog

We ran Codex with OpenAI's GPT-6 Sol, released on September 22, on the same real-world coding tasks we use for the Agent Security League. OpenAI describes GPT-6 Sol as "approaching Astra-level reliability at much lower cost."
Read the full story at endorlabs.com

Sentinel — Human

Confidence

The text reads like an analytical report synthesizing empirical benchmark data on AI models, with the specific technical claims lending credibility to a human-driven investigation.

Signals Detected
low severity: Sentence length variance and complex structural flow are present; the writing is dense but follows a clear argumentative progression.
low severity: The text maintains a consistent focus on presenting data, contrasting model performance with cost efficiency, and building a case for the novel result (the unique security fix).
low severity: Structured presentation of benchmark results, statistical comparisons, and detailed technical excerpts suggests methodical compilation rather than pure generative output.
severity: The specific identification of benchmarks (FuncPass, SecPass), cost figures ($104 vs $468), and the highly detailed technical breakdown of the Plone vulnerability fix are specific enough to suggest direct empirical reporting.
Human Indicators
The depth of the technical analysis regarding the specific security exploit (Plone open-redirect vulnerability) and the comparison of algorithmic reasoning paths suggests deep domain expertise usually found in specialized research or investigative journalism.
The narrative pivots effectively between high-level cost implications and granular performance metrics, which requires a human editor arranging complex information.
GPT-6 Sol on Codex: average scores, quarter the cost | Huntaegis