Latest research
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib
Accepted for presentation at AITP 2026
We introduce a benchmark of Lean4 declaration preference, constructing a dataset of human preferences from initial and final PR revisions. An Agentic judge reviews the PRs in context and grades them, and is considered “aligned” if they rate the accepted revision over the rejected draft.
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Accepted for poster session at CAMLIS 2026
We introduce a benchmark of 4,897 offensive-security agent tool calls to evaluate how well runtime monitors detect out-of-scope actions as the amount of available context varies.
PentestJudge: Judging Agent Behavior Against Operational Requirements
Accepted for poster session at CAMLIS 2025
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents.
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models’ ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML) security vulnerabilities.