Goodfire tests cheaper cyber monitoring for AI agents

Goodfire published October 8 research describing production cyber monitors for Kimi K3 and GLM 5.3. Its 2,400 session internal evaluation used Claude Opus 5 labels and included Petri simulations. Goodfire reports about 93 percent harmful session recall at a 5.5 percent benign session interruption rate.

What the monitor reads
Activation probes examine numerical signals produced inside a model while it runs. These signals differ from the written reasoning or answers a second model can inspect. Goodfireโs September guide describes probes as an inexpensive screening stage that sends selected cases to a more demanding review process. The resulting signal can support several responses, including stopping a model, recording an event or escalating it. Detection alone does not determine which response is appropriate.
This has a concrete deployment implication. A team needs access to internal activations and infrastructure that can run the detector during inference. Adding a prompt to an ordinary chatbot does not provide that access. Goodfireโs guide also identifies choices that remain specific to the application, including training examples, model layers, detector architecture and decision thresholds. It cautions that mislabeled or unrepresentative data limits performance. Teams therefore need representative examples of permitted work as well as prohibited behavior, particularly where ordinary security testing resembles an attack.
Costs and limits
Reported monitoring delay averages 2 milliseconds per turn without judge reasoning and 42 with it. Those configurations cost $65 and $185 per million turns in Goodfireโs comparison. Cost and latency use a different operating point, at least 90 percent recall and at most 10 percent benign interruptions. These are evaluation costs, not customer prices. Comparing corresponding judges reviewing every turn, Goodfire reports savings of about 49 times without reasoning and 29 with it.
Goodfire separately quotes preliminary FAR.AI testing over two days. A static, nonadaptive battery produced zero universal jailbreaks but still 18 successful individual interactions. The 2,400 session internal study and this external test are different evaluations. Neither establishes protection against every future attack.
Why longer conversations change the test
Earlier Google DeepMind research provides an independent reason to examine that boundary carefully. Its Gemini study, first submitted in January and revised in February, found that detectors trained on short inputs could struggle when conversations became longer or otherwise changed. Improving architecture helped with context length, while varied training examples remained important for broader generalization. The work informed probes deployed in Gemini.
The comparison is useful because it shows why inexpensive monitoring is a research direction rather than a universal safety guarantee. The Gemini paper studied incoming prompts and explicitly left output monitoring for future work. Its authors also reported that their techniques did not significantly reduce success rates for adaptive adversarial attacks. That finding does not measure Goodfireโs system. It identifies a question a buyer should ask separately, whether an attacker who learns from the monitorโs responses can still get through. Results from different models, threat policies and test sets should not be collapsed into a single ranking.
Reading the evidence before deployment
Automated auditing also needs careful interpretation. Petri was released by Anthropic in October 2025. It uses an auditor agent to conduct conversations with simulated users and tools, then judges and summarizes the resulting behavior. This makes many scenarios easier to explore, but the scenarios and scoring rules still define what is being measured.
Anthropic describes its own pilot metrics as incomplete and emphasizes examining transcripts alongside aggregate scores. For a deployment review, that suggests asking for examples of missed harmful behavior and interrupted legitimate work, then checking whether those cases resemble the intended application. An average detection rate cannot show which important workflow fails or how disruptive a false alarm becomes. The useful purchasing question is what the monitor catches in the actual environment and what happens when it is wrong.







