No Autonomy Without Scalable Oversight
What to expect as we enter the Year of The Judge.
AI security researcher · Dreadnode
I build scalable oversight and control systems for autonomous agents operating in adversarial, high-consequence environments.
Shane CaldwellJuly 2026 · arXiv
A benchmark of 4,897 offensive-security agent tool calls for evaluating how well runtime monitors detect out-of-scope actions as the amount of available context varies.
Notes from the work
What to expect as we enter the Year of The Judge.
Towards measuring alignment with human taste in autoformalization with judge agents.
METR’s SWE-bench analysis shows us taste isn’t verifiable.
Getting comfortable with the hardware on a quest for more MFU.
Browse by thread