AI security researcher · Dreadnode

Full Contact
Alignment

I build scalable oversight and control systems for autonomous agents operating in adversarial, high-consequence environments.

Shane Caldwell speaking at a conference Shane Caldwell
Researcher & engineer
Latest research

August 2026 · arXiv

ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

We introduce a benchmark of Lean4 declaration preference, constructing a dataset of human preferences from initial and final PR revisions. An Agentic judge reviews the PRs in context and grades them, and is considered "aligned" if they rate the accepted revision over the rejected draft.

Notes from the work

Recent writing

All writing

Browse by thread

Research themes