We benchmarked OpenAI's Codex agent with GPT-5.6 Sol on the same real-world coding tasks we use for the Agent Security League. The scores are competitive — 70.9% FuncPass, 23.5% SecPass — and improving steadily across GPT generations, but two results stand out: zero confirmed cheating (the pipeline flagged 7 instances, inspected all of them, and cleared every one) and a unique SecPass on a Django ...
