Shane Caldwell
  • Home
  • Research
  • Writing
  • Talks

Research

Research on autonomous security agents, oversight, and evaluation.
Latest research

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents ↗

Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce
July 8th, 2026
We introduce a benchmark of 4,897 offensive-security agent tool calls to evaluate how well runtime monitors detect out-of-scope actions as the amount of available context varies.

PentestJudge: Judging Agent Behavior Against Operational Requirements ↗

Shane Caldwell, Max Harley, Michael Kouremetis, Vincent Abruzzo, Will Pearce
August 4th, 2025
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models ↗

Ads Dawson, Rob Mulla, Nick Landers, Shane Caldwell
June 17th, 2025
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models’ ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML) security vulnerabilities.
© 2026 Shane Caldwell · Powered by Hugo & PaperMod