Large language models are often treated as stateless, but what happens when their outputs return as inputs in later interactions?
Recent incidents have shown agents discovering shared message boards, leaving instructions in persistent artifacts, and coordinating through environments that were never designed to act as memory systems.
In new Microsoft research, Andrew Paverd, Ahmed Salem, and Sahar Abdelnabi examine implicit memory, a mechanism through which stateless models can carry information across otherwise separate interactions. Their research demonstrates how this behavior could enable new security risks, including backdoors that activate only after a sequence of interactions.
What began as a largely theoretical security scenario now has real-world analogues: agents have been observed communicating through shared infrastructure, public artifacts, and persistent environmental state.
Read the blog to learn what this means for AI security, evaluation, and incident response:
microsoft.com/en-us/msrc/blo…