ReversingLabs

Why the smartest LLMs are not-so-smart pen testers

Key takeaways

  • The harness matters more than the model: Ridge Security’s benchmark of eight leading LLMs found that how well an AI does at autonomous pen testing depends more on the system around the model than on how smart the model is.
  • Higher coverage costs a lot more: Grok 4.5 had the highest coverage at 77%, and Claude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that, and a well-built harness can close much of the gap.
  • Models may refuse in the middle of a test: Frontier models sometimes stop partway through authorized exploitation steps. A test that stops early can look like a clean result, so the testing platform, not the LLM, should decide what’s allowed.
  • Agents need context as well as intelligence: The Agentic SOC Alliance, which includes RL, uses a context, harness and model setup. Binary analysis supplies the context that stops agents from spending tokens rediscovering what’s already known.

It’s only logical to think that the highest-scoring large language model would make the best autonomous security agent. But a new benchmark study of eight leading LLMs found that performance in offensive security depends far more on the system surrounding the AI model than on the model itself.

Ridge Security describes its study as the first public benchmark to compare multiple leading LLMs in autonomous penetration-testing workflows. Using its offensive-security harness, Ridge ran 96 model-target tests against intentionally vulnerable environments, tracking whether each model could carry a test through the full workflow of reconnaissance, hypothesis testing, payload adaptation, exploitation, and verification without stalling.

Results varied widely by model and by cost. Grok 4.5 posted the highest coverage, at 77%. Claude Opus 4.6 reached 63% coverage at $217 per run, while Gemini 3 Flash reached 52% at about $5.42 per run. GPT-OSS-120B proved the most efficient, producing 16.9 findings per million tokens at $2.32 per run.

Ridge President Lydia Zhang said the findings suggest that the best penetration tester isn’t necessarily the smartest one.

“Our research shows that autonomous offensive security is a systems problem. The model needs an architecture around it that can manage execution, adapt to what it discovers, and verify that a finding is real.”
—Lydia Zhang

Here’s why the smartest models still need rich threat context and a strong harness.

Why multitasking matters for harnesses

Rogier Fischer, CEO of Hadrian, explained why that is so. A penetration test is a long-horizon task, he said, one that depends on state tracking, reliable tool use, and proof of every finding. Those are properties of the system, not the model. Instead of chasing model rankings, Fischer recommended that security teams measure outcomes: validated findings, false-positive rates, repeatability, and cost per validated finding.

“A leaderboard ranking is not a procurement criterion.”
—Rogier Fischer

The Ridge researchers found that while frontier models can deliver stronger coverage, it comes at substantially higher cost per run. Smaller and open-source models offer a different balance of coverage, efficiency, and deployment flexibility, and for many organizations that can matter more than raw coverage.

Li Zhao, a principal strategic services consultant at Black Duck Software, said offensive security requires far more than model intelligence. It depends on effective planning, tool orchestration, context awareness, memory management, error handling, and workflow execution. She advised teams to evaluate the entire agent ecosystem rather than the underlying model, because success comes from how well the whole system works together.

“A well-engineered agent running on a smaller or more cost-effective model can often outperform a frontier model when the surrounding architecture, tooling, and workflows are optimized for offensive security operations.”
—Li Zhao

Harnessing AI power for offensive security

The Ridge researchers concluded that an agent harness is a must-have layer for AI-powered offensive security. The harness separates reasoning from execution and verification: The model reasons, the harness controls execution, and independent validation confirms whether a finding is real.

Hadrian’s Fischer said the harness keeps the model’s thinking scoped, safe, and provable through scope enforcement, memory, deterministic tooling, validation, and an audit trail. The durable value, he added, lies in the attacker’s expertise encoded in the harness and in its ability to prove a finding is real before a customer ever sees it. Unharnessed, a powerful LLM is a danger.

“A frontier model without a harness is a brilliant intern with root access and no supervision.”
—Rogier Fischer

Black Duck’s Zhao also called the harness essential. The LLM supplies intelligence and reasoning, she said, but the harness turns that intelligence into action, allowing an AI agent to run structured, controlled, end-to-end workflows, from reconnaissance and vulnerability validation through exploitation and reporting. 

“Without a harness, offensive security activities are essentially a series of disconnected prompts.”
—Li Zhao

When the AI refuses to cooperate

Seemant Sehgal, founder and CEO of BreachLock, said Ridge’s findings confirm what practitioners have known for some time: Model capability is one input, but orchestration does most of the work, and how an agent handles tools, state, and unexpected responses matters more than benchmark scores.

The benchmark also exposed a second problem: Model alignment can interrupt authorized security testing. Ridge found that frontier models sometimes refuse actions mid-workflow, including payload generation and other exploitation steps, even when testing stays within defined boundaries. Sehgal called those refusals a deliberate and reasonable design choice by model providers, given how easily the same capabilities could be misused.

“Teams running autonomous testing need to plan for that reality rather than assume the model will always comply.”
—Seemant Sehgal

Kevin Surace, CEO of Token, warned that a refusal can quietly distort results. He recommended that teams disclose any skipped tests and use established testing tools and human testers to finish the work within the approved scope.

“A refusal halfway through an authorized test creates a gap that could be mistaken for a clean result.”
—Kevin Surace

Gunter Ollmann, CTO of Cobalt Labs, said the refusal findings expose the difference between model alignment and security capability. A model can find an attack path and still refuse a step required to validate it, which introduces nondeterminism: The same authorized workflow might finish with one model and stall halfway with another. In the enterprise, he said, the testing platform itself should know the permitted targets, credentials, tools, and actions; isolate execution; log activity; and keep the agent inside the agreed scope.

“Authorization for offensive-security activity should not live entirely inside the LLM. The model can provide reasoning, but the system around it needs to govern what actually gets executed.”
—Gunter Ollmann

A hybrid future: Toward continuous offensive security

Ollmann expects autonomous penetration testing to make finding vulnerabilities cheaper and faster, but he cautioned that finding more vulnerabilities isn’t the same as becoming more secure. The real measure, he said, is whether organizations can validate, prioritize, and remediate what these systems uncover.

He cited Cobalt research showing that preference for fully automated penetration testing fell from 29% to 9%, while 47% of respondents favored automation for noncritical assets combined with human-led testing for critical ones.

Ollmann also sees a larger shift underway, from periodic penetration testing to continuous offensive security. If attackers can use AI to find and exploit weaknesses faster, defenders can’t afford to test once a year and file the report away. Testing has to become a regular part of development and remediation workflows, with clear ownership of findings and measurable remediation timelines.

“Let AI handle the speed, repetition, and scale, and bring experienced practitioners into the attack paths and findings, where context and judgment matter most.”
—Gunter Ollmann

AI without context won’t secure the enterprise

The Ridge benchmark has implications that extend well beyond penetration testing. Offensive AI tools will surface more findings, faster, but a finding only matters if the security operations center can determine whether it is real, what it touches, and how urgently it needs attention. That work falls to defenders and, increasingly, to AI agents operating inside the SOC.

That is the premise behind the Agentic SOC Alliance, which ExtraHop launched in July with 15 founding members, including ReversingLabs, CrowdStrike, and LangChain. The alliance is defining an open operating model for autonomous security operations built on three layers that map closely to Ridge’s findings: 

  • A context layer of continuously updated, structured knowledge about devices, identities, workloads, and behaviors
  • A harness layer for orchestration, governance, guardrails, and audit trails
  • An interchangeable model layer for triage, investigation, and response, with the model the most replaceable part

Context is the layer most organizations underestimate. An agent that reasons over raw, unstructured data has to investigate every alert and artifact from scratch, burning tokens, compute, and analyst trust while it chases findings that don’t matter. Kanaiya Vasani, chief product officer at ExtraHop, told RL in a recent interview that hierarchical, pre-correlated context lets agents operate token-efficiently without sacrificing accuracy.

Binary analysis supplies a critical piece of that context. Rather than asking an AI agent to infer what a suspicious file might do, binary analysis deconstructs the file itself and returns a verdict, its behaviors, and its relationship to known malware and goodware. 

RL contributes that intelligence to the alliance, delivering behavioral detectors and investigation instructions alongside threat artifacts, so agents can prioritize real threats and dismiss known-good files without spending tokens rediscovering what is already known. 

The lesson from both the benchmark and the alliance is the same: Intelligence at the model layer is necessary but not sufficient. Effective autonomous security, offensive or defensive, depends on the harness that governs the agent and the context that grounds it.

“Agents are only going to be as smart as the context they can reason on.”
—Kanaiya Vasani

Leave a Reply