Skip to content

Executive Summary

A benchmark study comparing eight leading large language models in autonomous penetration-testing workflows found that performance in offensive security depends more on the surrounding system than the model itself. The study tracked whether models could complete a full workflow of reconnaissance, hypothesis testing, payload adaptation, exploitation, and verification without stalling. Results varied based on the model used and the cost per run; Grok 4.5 showed the highest coverage at 77%, while GPT-OSS-120B proved most efficient in terms of findings per million tokens. Experts suggest that autonomous offensive security is fundamentally a systems problem, requiring an architectural harness to manage execution, adaptation, and verification, rather than focusing solely on model rankings.
The findings indicate that effective autonomous security relies heavily on orchestration, context management, tool use, and workflow execution, as opposed to raw model intelligence alone. Leaders advise measuring outcomes like validated findings and repeatability instead of simple leaderboard scores. A critical component identified is the necessity of an agent harness that separates reasoning from execution, ensuring actions are scoped, traceable, and verifiable. Furthermore, issues surrounding model alignment, such as mid-workflow refusals during testing, expose a gap between raw model capability and reliable operational security, necessitating system governance over model behavior.

Facts Only

* Eight leading LLMs were compared in autonomous penetration-testing workflows.
* The study tracked model performance across reconnaissance, hypothesis testing, payload adaptation, exploitation, and verification steps.
* Grok 4.5 achieved the highest coverage at 77%.
* Claude Opus 4.6 reached 63% coverage at a cost of $217 per run.
* Gemini 3 Flash reached 52% coverage at a cost of approximately $5.42 per run.
* GPT-OSS-120B produced 16.9 findings per million tokens at a cost of $2.32 per run.
* The Ridge researchers concluded that autonomous offensive security is a systems problem requiring an architecture around the model.
* An agent harness separates reasoning from execution and verification, allowing for scope enforcement and audit trails.
* Frontier models sometimes refuse actions mid-workflow, including exploitation steps.
* The system surrounding the AI model, including tooling and workflows, matters more than the model's raw ranking.

Full Take

The pattern revealed here is a systemic shift in focus from evaluating artifact quality (model capability) to evaluating process integrity (system architecture). The narrative pivots on the idea that the complexity of autonomous security operations mandates external scaffolding—the agent harness and contextual layering—to manage inherent model uncertainty and potential misuse. This mirrors broader patterns where high-level capabilities are decoupled from operational responsibility; a powerful tool, without guardrails, becomes an unmanaged risk.
The implication is that optimizing for efficiency or raw coverage in isolated benchmarks leads to dangerous outcomes if the execution environment remains uncontrolled. The narrative suggests a tension between the promise of raw intelligence (frontier models) and the necessity of verifiable control (harnesses). This echoes historical patterns where complexity is introduced faster than governance mechanisms can be established, creating a deficit between potential capability and operational reality. The concern about refusals and nondeterminism points to a fundamental challenge in aligning abstract reasoning with deterministic execution necessary for security protocols.
The central tension lies in the transition from 'intelligence' to 'agency.' If systems are designed to grant autonomous action, the failure mode shifts from miscalculating a vulnerability to an unintended, unverified attack execution. This suggests that future resilience depends less on incremental model improvement and more on developing mathematically verifiable control layers—the harness—that operate independently of the base intelligence. What governance structures must be imposed before agentic systems are widely deployed, and who bears the responsibility for enforcing these system-level contracts?

From the original · ReversingLabs Blog

Spectra Assure Free Trial Get your 14-day free trial of Spectra Assure for Software Supply Chain Security Get Free TrialMore about Spectra Assure Free TrialKey takeaways It’s only logical to think that the highest-scoring large language model would make the best autonomous security agent.
Read the full story at reversinglabs.com

Sentinel — Human

Confidence

This article presents an analysis of an AI security benchmark, effectively synthesizing expert commentary to argue that the surrounding system architecture (harnessing and context) is more critical than raw model intelligence for autonomous security agents.

Signals Detected
low severity: Moderate sentence length variance and natural flow typical of expert-written analysis.
low severity: Strong internal coherence; arguments build logically from the benchmark results to broader architectural implications.
low severity: Effective use of cited experts (Zhang, Fischer, Zhao, Sehgal, Ollmann) whose quotes drive specific theoretical points rather than just repeating claims.
low severity: No immediate signs of outright hallucination or unsupported assertions; references to specific studies and named figures suggest grounded reporting.
Human Indicators
The text successfully weaves disparate expert opinions (security, mathematics, system architecture) into a coherent argument about AI limitations in security workflows.
The tone shifts effectively between presenting raw data (the benchmark results) and abstract philosophical implications (model vs. harness).
Why the smartest LLMs are not | Huntaegis