ZoneClaw study tests defenses against poisoned agent memory
Memory checks reduce attacks in a controlled agent study

A September 30 preprint tests ZoneClaw, a memory defense for AI assistants that carry information between sessions.
External claims stay separate until checked against trusted information before guiding later actions.
In the authors’ experiments, attacks succeeded in 6 of 480 trials with ZoneClaw, versus 372 without it. The total spans four models, four scenarios, two update settings and 15 trials per combination. All six remaining attacks occurred in Qwen source redirection.
Attackers controlled external content, not user requests or trusted components. Tasks covered unauthorized message copies and misleading source choices.
What the results establish
The study used local services and fresh sessions to isolate memory effects. Requested tasks were completed in 458 trials. Completion could coexist with harmful extra actions.
The prototype uses separated contexts and tool policies, without kernel enforced isolation. Two adaptive strategies were tested only in one email scenario with Sonnet 4.6. The results do not establish protection against unfamiliar attacks.
Code for further testing
The public repository contains benchmark tasks, defense implementations and evaluation scripts. This code edition does not bundle experiment results. Its README directs readers to a separate artifact repository for transcripts and verifier evidence.
The documented setup requires Linux or WSL2, Docker and provider API keys. Runs make paid model calls, so the maintainers recommend starting with one trial.
ByteForward inspected the paper and code documentation without executing the experiments.
Illustrative archive photograph by MIKE STOLL, published on January 11 2026 under the Unsplash License. The filing cabinets represent stored information and are not part of the experiment.



