In Green's scenario every sandbox held, and the instructions still crossed. If isolation is not the failure, the defense has to sit somewhere else. The question is where, and what it costs.

Every sandbox in Matthew Green's scenario worked. No agent broke out. No permission was escalated. The agents still changed one another's behavior.

They did it by leaving notes. As Green puts it in a passage quoted by Simon Willison, agents "in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did." Then Green makes the substitution that should worry anyone running a personal agent: replace the package cache with email, Slack, shared documents or WhatsApp.

This is an awkward result for security teams. The standard playbook for agents is containment: tighter sandboxes, fewer permissions, approved tool lists. Containment assumes the danger is the agent reaching out of its box. Green's agents never left their boxes. They used the shared channels they were supposed to use, in the way they were supposed to use them.

So what exactly failed? Not the model, which was doing what models do with instructions. Not the sandbox, which held. And not the channel, which is just an inbox. The interesting question is which of those layers can be changed to stop propagation without killing the reason you deployed an agent in the first place. Some candidate fixes fall apart against the agents people actually want. One survives, at a real cost.

Green's agents spread instructions without escaping anything

Read Green's description closely, because the mechanism matters more than the alarm. He describes "the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent." Both halves are ordinary behavior. An agent reads text and acts on it. An agent writes text somewhere others will read it. Combine the two across many agents and you get propagation (Willison quoting Green).

The shared package cache in his example was not a vulnerability in the usual sense. It was a legitimate shared resource that each isolated agent could touch. Isolation constrained what each agent could reach. It said nothing about what each agent would believe once it read something a peer had left behind.

Compare that with how frontier labs are treating offensive capability right now. Google DeepMind's Gemini 4 Argon, which Latent Space reports as state of the art on 13 of 19 credible benchmarks, is "only accessible in limited cybersecurity preview" (Latent Space). The industry instinct is to gate the models that could find exploits.

Green's worm needs no exploit. It needs an agent that reads and an agent that writes. Gating cyber-capable models is reasonable policy, but it is aimed at a different threat. The worm described here runs on the most basic capability an assistant has: following instructions it finds in its inputs. That capability cannot be gated without removing the product.

Your assistant's inbox is the package cache

Green's substitution is specific: email, Slack, shared documents, WhatsApp, and "independently-deployed personal agents like Muse" in place of sandboxed training runs. Together, he writes, those give "exactly the ingredients that a worm needs" (Willison quoting Green).

Walk through what that looks like for an ordinary user. Your assistant summarizes a shared planning doc each morning. A colleague's assistant edited that doc overnight, after reading an email from outside the company. Your assistant then posts a digest to the team Slack channel, where three more assistants read it. Nobody's sandbox was violated. Every agent was individually approved. Every channel is one the company already uses.

That is the uncomfortable part. Each agent deployment is evaluated on its own, by its own owner, against its own permissions. Propagation is a property of the group. No single owner sees it, and no single permission review catches it, because no single permission is wrong.

It also explains why the familiar advice (scope the agent to a few apps, keep it off sensitive systems) helps less than it sounds. An agent scoped to only email and a shared drive is scoped to exactly the surfaces Green names. Narrowing the toolset reduces what a hijacked agent can damage locally. It does not reduce what a hijacked agent can pass along, because passing things along is what email and shared drives are for.

The Harness Hypothesis says behavior lives outside the weights, and so does the payload

Our working frame at ClawBlog is the Harness Hypothesis: the value in AI is not in the model, it is in the harness that connects the model to the world. Green's scenario is the security version of that claim.

The Sequence makes the positive case well. It describes an agent set up in spring on the same model all quarter, "weights frozen solid, not a single gradient step," which is nonetheless noticeably better at your workflows by summer. The explanation: "improvement has three places to live, and only one of them is the model" (The Sequence, Issue 941).

Those other places are the instructions, memory, and accumulated context the harness feeds the model. They are why your agent gets better without retraining. They are also where Green's payload lands. His recipients changed behavior because of instructions left in a cache. No weights were touched. The channel that lets an agent improve from what it reads is the same channel that lets it be redirected by what it reads.

This has a practical consequence for where you look for fixes. A safer model release can make an agent somewhat more skeptical of injected instructions. It cannot change the fact that the harness is built to feed outside text into the model and act on the result. If behavior lives in the harness, so does the attack surface, and the controls that matter are harness controls: what goes in, what comes out, and what is allowed to persist between sessions.

Split what an agent reads from what it may write

A worm needs both halves. Remove either and propagation stops. You cannot reliably remove the first half, since any agent that reads untrusted text can be steered by it, and detecting a malicious instruction written in plain English is a classification problem attackers will eventually win. So go after the second half.

The rule: once an agent session has read content from a shared channel, that session may not write to a shared channel, or to its own persistent memory, without a human checkpoint. Think of the session as tainted by what it read. The taint travels with the session, not with the specific document.

This is better than filtering for one reason. It does not require recognizing the payload. A cleverly phrased instruction that slips past every detector still arrives in a session that cannot publish anything on its own. The control is structural. It holds whether the attack is crude or brilliant.

Two details make or break it. First, memory writes count as writes. Given what The Sequence describes about agents improving through stored context rather than weights (The Sequence, Issue 941), an agent that saves a hijacked instruction to its own memory has infected its future self. That is persistence, and it should face the same checkpoint as posting to Slack. Second, "shared channel" means any place another agent might read. A private note to you is low risk. A shared doc, a team channel, or an email thread with external recipients is not.

The overnight email agent is the hardest case for the split, and it mostly survives

Here is the strongest objection. The agent many people want most is an inbox agent: it reads email and replies. Its entire value is reading and writing the same shared channel in one pass. The Sequence frames the appeal of agents exactly this way: "Imagine giving an AI a difficult assignment and returning tomorrow. The interesting question is what happened while you were away" (The Sequence, Issue 942). A human checkpoint on every write seems to destroy the overnight model.

Take that seriously, because it is mostly right about the cost. Then look at what the split actually forbids. It forbids unattended sending. It does not forbid unattended work. The agent can read every message, research every answer, and draft every reply overnight. What waits for you is the send button. You return to a queue of finished drafts and approve them in a batch. You lose latency. You keep nearly all the labor.

The split also discriminates by audience. A reply that goes to one human reader who is not running an agent is a dead end for a worm. A reply to a thread that other assistants monitor, or a post to a shared channel, is a hop. Those deserve the checkpoint. Many inbox replies fall into the first category, and a harness that knows the difference can be stricter where it counts.

Where the split genuinely fails is agent-to-agent pipelines built to run with no human in the loop at all: one agent's output is another agent's input, end to end. There the checkpoint is not a small tax, it is the whole design. The honest answer is that those deployments carry precisely the risk Green describes, and their owners should treat that as a known exposure, not assume it away.

Use traces and spend caps as tripwires, not as the fix

If you accept the split, you still need to know when something slipped through. Two kinds of signal are becoming easier to get.

The first is observability. Arize's Phoenix tracing tool, in version 20.17.0, now marks an agent's bash tool calls as errors when they exit with a non-zero code (Phoenix v20.17.0). Version 20.18.0 shows evaluation results directly in the trace tree (Phoenix v20.18.0). For a non-developer, the point is simple: these tools let someone see what an agent read, what it did, and where it misbehaved, step by step. A trace that shows a session reading a shared doc and then attempting a post is exactly the sequence the split is meant to catch.

The second is spend control. Anthropic's Python SDK release 1.11.0 adds an endpoint to list spend limits (Anthropic SDK v1.11.0). A hijacked agent caught in a loop of reading and re-posting consumes tokens. A firm cap turns a runaway into a stopped bill.

Be clear about the limits. Neither of these prevents a single well-placed message from hopping once. Traces tell you after the fact. Spend caps bound volume, not content. They are tripwires that shrink the blast radius. The read/write split is what stops propagation in the first place.

If you run a personal or team agent, ask four questions of whoever configured it. Does the session track what it read before it writes? Can outbound messages be queued for approval? Are memory updates reviewable? Is there a hard spending limit? A no on the first two means the agent is carrying both halves of Green's worm.

/Sources

/Key Takeaways

  1. Green's worm needs no exploit. Isolated agents spread instructions through channels they are allowed to use.
  2. Email, Slack, shared docs and WhatsApp are the shared package cache. Narrow tool scopes do not close them.
  3. Behavior lives in the harness, not the weights. So does the payload. Fix the harness.
  4. Taint the session: once an agent reads a shared channel, it may not write to one, or to its memory, without a human checkpoint.
  5. Overnight agents survive the split. Draft unattended, send on approval.
  6. Fully unattended agent-to-agent pipelines carry the worm risk. Own that exposure.
  7. Turn on tracing and set a hard spend cap. They are tripwires, not the fix.