Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Goodfire tests cheaper cyber monitoring for AI agents

Listen to this article

Goodfire published October 8 research describing production cyber monitors for Kimi K3 and GLM 5.3. Its 2,400 session internal evaluation used Claude Opus 5 labels and included Petri simulations. Goodfire reports about 93 percent harmful session recall at a 5.5 percent benign session interruption rate.

Server equipment with turquoise network cables and green indicator lights at Ifremerโ€™s DATARMOR centre
Illustrative server equipment at Ifremerโ€™s DATARMOR centre in France, photographed in May 2025. The photograph does not show Goodfireโ€™s test infrastructure. Photograph by Gregory Rocher, IFREMER, via Wikimedia Commons. CC BY 4.0.

What the monitor reads

Activation probes examine numerical signals produced inside a model while it runs. These signals differ from the written reasoning or answers a second model can inspect. Goodfireโ€™s September guide describes probes as an inexpensive screening stage that sends selected cases to a more demanding review process. The resulting signal can support several responses, including stopping a model, recording an event or escalating it. Detection alone does not determine which response is appropriate.

This has a concrete deployment implication. A team needs access to internal activations and infrastructure that can run the detector during inference. Adding a prompt to an ordinary chatbot does not provide that access. Goodfireโ€™s guide also identifies choices that remain specific to the application, including training examples, model layers, detector architecture and decision thresholds. It cautions that mislabeled or unrepresentative data limits performance. Teams therefore need representative examples of permitted work as well as prohibited behavior, particularly where ordinary security testing resembles an attack.

Costs and limits

Reported monitoring delay averages 2 milliseconds per turn without judge reasoning and 42 with it. Those configurations cost $65 and $185 per million turns in Goodfireโ€™s comparison. Cost and latency use a different operating point, at least 90 percent recall and at most 10 percent benign interruptions. These are evaluation costs, not customer prices. Comparing corresponding judges reviewing every turn, Goodfire reports savings of about 49 times without reasoning and 29 with it.

Goodfire separately quotes preliminary FAR.AI testing over two days. A static, nonadaptive battery produced zero universal jailbreaks but still 18 successful individual interactions. The 2,400 session internal study and this external test are different evaluations. Neither establishes protection against every future attack.

Why longer conversations change the test

Earlier Google DeepMind research provides an independent reason to examine that boundary carefully. Its Gemini study, first submitted in January and revised in February, found that detectors trained on short inputs could struggle when conversations became longer or otherwise changed. Improving architecture helped with context length, while varied training examples remained important for broader generalization. The work informed probes deployed in Gemini.

The comparison is useful because it shows why inexpensive monitoring is a research direction rather than a universal safety guarantee. The Gemini paper studied incoming prompts and explicitly left output monitoring for future work. Its authors also reported that their techniques did not significantly reduce success rates for adaptive adversarial attacks. That finding does not measure Goodfireโ€™s system. It identifies a question a buyer should ask separately, whether an attacker who learns from the monitorโ€™s responses can still get through. Results from different models, threat policies and test sets should not be collapsed into a single ranking.

Reading the evidence before deployment

Automated auditing also needs careful interpretation. Petri was released by Anthropic in October 2025. It uses an auditor agent to conduct conversations with simulated users and tools, then judges and summarizes the resulting behavior. This makes many scenarios easier to explore, but the scenarios and scoring rules still define what is being measured.

Anthropic describes its own pilot metrics as incomplete and emphasizes examining transcripts alongside aggregate scores. For a deployment review, that suggests asking for examples of missed harmful behavior and interrupted legitimate work, then checking whether those cases resemble the intended application. An average detection rate cannot show which important workflow fails or how disruptive a false alarm becomes. The useful purchasing question is what the monitor catches in the actual environment and what happens when it is wrong.

Marcus Reid
Marcus Reid

Marcus Reid is focused on covering the money, rules, and institutional choices shaping AI. He runs from funding rounds and chip deals to regulation, lawsuits, leadership changes, and the business of building enormous computing systems. Marcus follows the incentives behind the announcement. Who pays, who gains leverage, and what changes for everyone else? The voice is direct, measured, and occasionally dry, especially when a grand promise arrives with very little detail.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile