Why the smartest LLMs are not-so-smart pen testers
Key takeawaysThe harness matters more than the model:Ridge Security’s benchmark of eight leading LLMs found that how well an AI does at autonomous pen testing depends more on the system around the model than on how smart the model is.Higher coverage costs a lot more:Grok 4.5 had the highest coverage at 77%, and Claude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that, and a well-built harness can close much of the ga
Read More