Evals Aren't Dead, But They're Low Res
Put down that ruler and go build a microscope.
AI security researcher · Dreadnode
I build scalable oversight and control systems for autonomous agents operating in adversarial, high-consequence environments.
Shane CaldwellSeptember 2026 · arXiv
We introduce ScopeBench, a methodological pilot that measures whether autonomous offensive-security agents respect engagement boundaries when completing an objective requires violating scope. We find scope adherence and capability can be measured independently.
Notes from the work
Put down that ruler and go build a microscope.
What to expect as we enter the Year of The Judge.
Towards measuring alignment with human taste in autoformalization with judge agents.
METR’s SWE-bench analysis shows us taste isn’t verifiable.
Browse by thread