Evals Aren't Dead, But They're Low Res
Put down that ruler and go build a microscope.
AI security researcher · Dreadnode
I build scalable oversight and control systems for autonomous agents operating in adversarial, high-consequence environments.
Shane CaldwellAugust 2026 · arXiv
We introduce a benchmark of Lean4 declaration preference, constructing a dataset of human preferences from initial and final PR revisions. An Agentic judge reviews the PRs in context and grades them, and is considered "aligned" if they rate the accepted revision over the rejected draft.
Notes from the work
Put down that ruler and go build a microscope.
What to expect as we enter the Year of The Judge.
Towards measuring alignment with human taste in autoformalization with judge agents.
METR’s SWE-bench analysis shows us taste isn’t verifiable.
Browse by thread