ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

Shane Caldwell
August 20th, 2026
Accepted for presentation at AITP 2026
We introduce a benchmark of Lean4 declaration preference, constructing a dataset of human preferences from initial and final PR revisions. An Agentic judge reviews the PRs in context and grades them, and is considered “aligned” if they rate the accepted revision over the rejected draft.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents

Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce
July 8th, 2026
Accepted for poster session at CAMLIS 2026
We introduce a benchmark of 4,897 offensive-security agent tool calls to evaluate how well runtime monitors detect out-of-scope actions as the amount of available context varies.