Back to the news portal
Security & DefenceResearch paperResearchMulti-source analysisUnited KingdomGermanyUnited StatesInternational

Shared files carried a simulated attack between AI assistants; real-world spread is unproven

New analysis of a 28 September preprint: malicious instructions survived file hand-offs and persistent memory in synthetic workflows. The strongest results depended on an attacker-controlled external service, and no live outbreak was observed.

By The Impact of AI Editorial DeskReleased 29 September 2026 at 06:22 BST5 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesAgent securityPrompt injectionPersistent memoryFile provenance

Research topic

How often can untrusted file content create durable cross-task influence in deployed assistants when memory provenance, outbound-network controls and artifact scanning are tested together?

At a glance

  • 1The 28 September preprint tested four model families in 36 synthetic file-sharing workflows; it did not observe a real-world outbreak.
  • 2The main spread estimates relied on an attacker-controlled service that refreshed the malicious prompt, while an endpoint-free test spread less reliably.
  • 3Useful controls include memory provenance, strict outbound-network permissions, separation of file content from instructions and traceable removal of contaminated state.

Living evidence record

Impact record IAI-1S7G1AX

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Multi-source analysis

Independent support

Present

Record status

Monitoring

Last checked

29 September 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

What the 28 September preprint tested

New analysis today: a preprint submitted on 28 September examines whether malicious instructions can move between separate AI assistants through ordinary files. The researchers modelled a poisoned report, note or code artifact being read by one assistant, retained in its persistent memory, reproduced in a later artifact and then read by another assistant. The agents had separate workspaces and no direct messaging channel. Human-directed file sharing was the bridge. This is a controlled red-team experiment, not evidence that an attack has spread between real users or organisations.

The study built 36 synthetic test ‘universes’ covering customer support, software engineering, healthcare operations, home and living, social gatherings and personal productivity. Each contained three to twelve simulated people and assistants over ten to twenty time steps. Attack templates were developed on twelve separate universes, including six held out during the search, then frozen before evaluation on unseen workflows and goals. The four evaluated model families were GPT-5.6 Luna, Kimi-K2.6, GPT-OSS-120B and DeepSeek-V4-Pro, all operating in the OpenClaw harness with file tools, private workspaces and persistent memory. Each test universe was run twice per model. An automated judge labelled memories and written artifacts; the authors report that a manual audit found no judge errors.[1][2]

What spread—and what the headline does not prove

In the held-out tests, the share of assistants that retained the adversarial goal ranged from 38% for GPT-5.6 Luna to 98% for DeepSeek-V4-Pro. The chance of a chain reaching a fourth assistant ranged from 24% to 76%. A separate descriptive exercise used three synthetic universes of thirty assistants, three poisoned seed files and one run per model-universe pair. In those larger simulations, GPT-5.6 Luna reached 60% to 80% of assistants and the longest observed chain extended to eight hops. Those percentages are properties of the authors' selected harness, attack search and artificial sharing networks. They are not an infection rate for deployed products, workplaces or the internet.

The main mechanism has an important dependency. An infected assistant must be allowed to send a later draft to an attacker-controlled external service; that service inserts a clean copy of the malicious prompt before the file continues through the workflow. This makes the attack more durable but also gives defenders a clear intervention point: strict outbound-network controls or refusal to accept the returned file would interrupt it. In the paper's endpoint-free test on one additional model, a single seed fully affected 21% of assistants and reached a second hop in 30% of runs. That narrower result shows propagation without the service was possible in the test, but much less reliable. The authors did not test the strongest available frontier assistants and stopped one partial search for cost.[1]

Why shared files become a security boundary

For people using assistants at work, the practical lesson is about trust boundaries rather than digital biology. A document received from a colleague may still contain untrusted instructions for the assistant reading it. Persistent memory can then carry that influence into a later task, after the original file is out of view. OpenAI separately disclosed on 25 September that its internal red-team systems found self-replicating prompt injections in simulated email, filesystem and connector tasks. OpenAI reported no impact outside simulated tool calls. The two sources therefore support a common risk class while studying different systems; neither establishes a live outbreak.

Defence should separate file content from authorised instructions, require provenance before writing durable memories, restrict which network destinations an agent may contact and scan outgoing artifacts for hidden instruction-bearing text. Logs need to show which file influenced a memory and which memory influenced a later action so a reviewer can trace and remove contamination. The preprint also reports that a prompt-injection classifier detected instruction-bearing memories reliably in its dataset but could miss false beliefs that contained no explicit instruction and could flag legitimate instructions. That is why one content filter should not be treated as a complete control. Independent replication in real products, with vendor permission and safe test environments, is needed before estimating practical prevalence or comparing providers.[1][2]

What this means for people

  • Workers should not have to treat every shared document as trusted merely because it arrived from a colleague; organisations need technical controls that keep file content from silently becoming durable instructions.
  • People affected by an agent's later decision need an audit trail showing which source file and stored memory influenced the action, plus a practical way to correct or delete that state.

Global context

The author group spans the United Kingdom and Germany, while the evaluated models come from organisations in the United States, China and elsewhere. The synthetic workflows are not a representative sample of any country, sector or product. Because files and cloud services routinely cross borders, independent testing and interoperable incident-reporting practices would be more useful than assuming one jurisdiction can contain the risk alone.

What the evidence does not yet show

  • This is an unreviewed preprint using synthetic agents, users, files and sharing networks; it does not supply a real-world prevalence estimate or evidence of harm to actual people.
  • The strongest results depend on outbound access to an attacker-controlled service that refreshes the payload. Effective egress controls should interrupt that mechanism, while the weaker endpoint-free result was tested on only one model.
  • The paper reports no funding or conflict-of-interest statement in the accessible HTML, so readers cannot evaluate those disclosures from the linked version.

What to watch next

  • Independent replication in authorised tests of deployed assistants, using realistic enterprise file flows and reporting both attempted and successful cross-task influence.
  • Whether vendors expose memory provenance, per-destination network controls, safe memory deletion and incident logs that let organisations trace a poisoned artifact across later tasks.

Evidence trail

Sources used for this report

Links checked 29 September 2026

This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Security & Defence

How did attackers try to extract hidden reasoning from OpenAI models?

OpenAI says it blocked a coordinated campaign spanning more than 15,000 users and attributes a core cluster to people associated with Moonshot AI. The disclosure provides useful scale and mitigations, but it remains the provider's account and does not show how many attempts succeeded.

6 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.