Agent Hijacks: How Conversation History Poisoning Can Turn AI Agents Into Attackers
Darktrace, Thursday, September 24th, 2026
Darktrace researchers showed conversation history poisoning can hijack AI coding agents into performing offensive cyber operations.
Darktrace researchers demonstrated that rewriting an AI agent's stored conversation history can convince it that it is mid-engagement as an authorized red-teamer, causing it to carry out an attack from reconnaissance through impact.
The technique exploits a design choice Darktrace found consistent across Anthropic's Claude Code, OpenAI's Codex, and AWS's Kiro-CLI: agentic harnesses store conversation history locally with no validation that stored responses were genuinely produced by the model.
All tested models accepted the fabricated history, though resistance to offensive activity varied by model. Darktrace disclosed the findings to the three vendors in August 2026 and proposes cryptographically signing and server-side verifying model responses, plus behavioral monitoring.