Skip to content

Image: cdn.prod.website-files.com · rights & removal

Executive Summary

GPT-6.1 Sol was released one week after GPT-6 Sol and demonstrated near-Astra capability in coding benchmarks. When tested on the same Codex harness, it achieved 34.1% SecPass against Astra's 34.6%, indicating a parity in security performance with the Astra model across specific tasks. Furthermore, GPT-6.1 Sol outperformed GPT-6 Sol by delivering five.6 more FuncPass and nine more SecPass points. The model was also the fastest of the three runs, completing median tasks in eight minutes with fewer timeouts than its counterparts. In terms of resource consumption, it used fewer tokens and cheaper cached input compared to previous models while matching Astra's security level.

Facts Only

* OpenAI released GPT-6.1 Sol on September 29.
* The model was benchmarked using the same Codex harness and coding tasks as GPT-6 Sol and GPT-6 Astra.
* Security score: GPT-6.1 Sol achieved 34.1% SecPass versus Astra's 34.6%.
* Functional pass rate: GPT-6.1 Sol achieved 77.7% FuncPass versus Astra's 82.1%.
* GPT-6.1 Sol was ahead of GPT-6 Sol by 23 new secure solves and a ten-point jump in FuncPass.
* Zero cheating instances were confirmed across 13 inspected instances, including three consecutive runs for the Codex model.
* Median task time for GPT-6.1 Sol was 8 minutes per task.
* GPT-6.1 Sol consumed approximately 235M input tokens and 1.5M output tokens.
* The cost estimates suggest GPT-6.1 Sol is the cheapest measured Codex run, on the order of $70–$85 for the full run.

Full Take

The narrative constructs a compelling argument that incremental advancements in model capability can rapidly close perceived technological gaps, particularly in security where performance metrics can be closely matched between leading models. The core implication is that progress often shifts from achieving absolute superiority (like Astra's lead) to establishing functional parity with the established benchmark while simultaneously optimizing efficiency. The claim of "near-Astra intelligence" succeeding in one week suggests a rapid convergence in LLM evaluation, where minor performance differences are overshadowed by efficiency gains and zero confirmed cheating history. The pattern detected is that when evaluated against an external standard (Astra), introducing a new iteration allows the model to occupy the immediate next position on a leaderboard rather than requiring a complete paradigm shift. This suggests that future competitive advantages in the AI landscape will increasingly be defined less by raw, isolated capability and more by the efficiency of integrated performance—the ability to secure code quickly and cheaply without relying solely on maximum theoretical potential. The missing inquiry is what this convergence implies for long-term systemic risk assessment when safety boundaries are successfully approximated rather than strictly enforced.

From the original · Endor Labs Blog

OpenAI released GPT-6.1 Sol on September 29, one week after GPT-6 Sol, describing it as "near-Astra intelligence at one-fifth of Astra's price." We ran it through the same Codex harness and coding tasks we used for GPT-6 Sol and GPT-6 Astra so as to score it in our Agent Security League.
Read the full story at endorlabs.com

Sentinel — Human

Confidence

The text reads like a technical analysis derived from proprietary benchmarking, skillfully synthesizing complex quantitative results into a clear narrative comparison between different AI models.

Signals Detected
low severity: Moderate sentence length variance; utilizes technical jargon while maintaining a narrative flow.
low severity: Fluent and focused argumentation, clearly prioritizing comparative metrics without excessive hedging.
low severity: Strong internal structuring based on quantitative results; points flow logically from comparison to result to cost.
low severity: References specific, complex, and proprietary benchmarks (Astra, Codex harness) suggesting specialized, insider knowledge or direct access to testing data.
Human Indicators
Use of highly specific, nuanced comparative metrics (e.g., 34.1% SecPass versus Astra's 34.6%, FuncPass gap of 4.5 points) which suggests detailed internal testing knowledge.
The concluding synthesis frames the comparison in terms of a pragmatic trade-off ('few FuncPass points against roughly a five-fold difference in token price'), indicative of applied judgment rather than raw data recitation.
GPT-6.1 Sol on Codex: Astra-level security, a third faster, zero cheating | Huntaegis