Skip to content

Image: res.cloudinary.com · rights & removal

Executive Summary

The testing compared two methods of vulnerability discovery: Evo Continuous Offensive Security (COS), which probes a live application, and Claude Security running on Mythos, which analyzes source code. The test used a deliberately vulnerable web application named TaintedPort to assess these approaches against known vulnerabilities and exploit chains.
Evo COS successfully confirmed 10 out of 15 exploit chains by chaining findings from both source code analysis and runtime interaction. For instance, it used a Server-Side Request Forgery (SSRF) flaw to access and use a hardcoded JWT signing secret to achieve administrator token forgery, demonstrating an autonomous attack path on the live application.
The results showed that Evo COS found 50 vulnerabilities out of 57 known issues, while Claude Security found 37 vulnerabilities from the source code review. While both methods identified flaws, the dynamic testing approach demonstrated a capability to find context-dependent business logic flaws and prove actionable exploit chains by observing runtime behavior.
The comparison suggests that exercising the live application provides context that reading the code alone misses, as understanding an exploit chain requires confirming reachability across multiple steps within a running system, which is a characteristic of dynamic testing.

Facts Only

* Evo COS confirmed 10 of 15 exploit chains.
* The test involved Evo COS attacking the live application TaintedPort and Claude Security analyzing the source code using Mythos.
* Every tool found the Server-Side Request Forgery (SSRF) flaw in the application.
* The application's JWT signing secret was hardcoded.
* Evo COS used the SSRF flaw to retrieve the signing secret from the running application, then minted a valid administrator token.
* This resulted in an account takeover path being demonstrated against the live application with a runnable proof of concept.
* Evo COS found 50 vulnerabilities out of 57 known vulnerabilities.
* Claude Security found 37 vulnerabilities out of 57 known vulnerabilities.
* Evo COS had a higher precision (96.2%) compared to Claude Security (90.2%).
* Evo COS confirmed runtime behaviors like reflected responses, missing transport protections, and weak session handling.
* Exploit chains require reaching each step through the running application for confirmation.

Full Take

The core tension in this demonstration lies between static analysis (reading code) and dynamic offensive testing (attacking a live instance). The finding that an autonomous attack requires proving exploitability—not just identifying a flaw—highlights a critical gap in current security tooling: the transition from vulnerability identification to demonstrable breach paths. The success of Evo COS demonstrates that understanding vulnerabilities is insufficient; context-aware reasoning about state change within a system, realized through interaction with the live environment, is the necessary next step.
The subsequent argument pivots on the role of the harness. It suggests that model quality alone is secondary to the systemic architecture surrounding the model. Evo COS succeeds not merely because it uses powerful frontier models, but because it orchestrates them as specialized agents within a validated ecosystem that starts from external context (the live target) rather than pure, uncontextualized reasoning. This implies that future progress in AI security will be less about optimizing individual models and more about designing verifiable, multi-agent systems where independent validation steps are built into the operational flow to prevent findings from being accepted without runtime proof.
The implication for human agency is that automation must move beyond pattern matching or isolated analysis to actual causal reasoning within complex systems. If AI tools only confirm code patterns, they remain assistants; if they can dynamically chain actions against a live target, they become instruments capable of demonstrating real-world risk. The question shifts from "what flaws exist?" to "what is the proven path to exploitation, and does the system account for that path in its validation?"

From the original · Snyk Blog

October 7, 2026 0 mins readAttackers don't read your repository; they hit your URL, and chain together whatever they find. In an era where offensive AI runs against live applications at machine speed, your security tooling needs to go beyond finding vulnerabilities to prove exploitability, including whether they can be combined into a breach.
Read the full story at snyk.io

Sentinel — Human

Confidence

LIKELY_HUMAN (confidence: 0.15)

Frontier models found the vulnerabilities. Only the attacker found the chains. | Huntaegis