We introduce a benchmark of 4,897 offensive-security agent tool calls to evaluate how well runtime monitors detect out-of-scope actions as the amount of available context varies.
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents.
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models’ ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML) security vulnerabilities.