Selective agent audits use fewer tokens but miss more unsafe runs
A new preprint measures the cost and detection tradeoff in agent monitoring.

A preprint posted October 5 by Purdue University researchers tests a way to monitor AI agents with fewer tokens. The proposed design selects pending actions for language model review before execution.
On 149 OpenAgentSafety test runs, tuned selective auditing used 14,600 tokens per run, compared with 47,800 when every action was reviewed. It detected 72.8 percent of unsafe runs versus 81.5 percent for full auditing. False alarms also fell.
What the figures include
The study counted provider refusals as detections. Fourteen of the selective system’s 59 detections depended solely on a refusal. Its results therefore combine the auditor’s judgments with the provider’s safeguards.
The benchmark and simulated case studies leave broader deployment performance unproven. Teams should assess missed detections alongside potential savings.
Related coverage explores Goodfire’s work on selective monitoring.
Illustrative archival padlock photograph by Estormiz under CC0. The photograph is not a study result or a guarantee of protection.



