Executive Summary
Researchers at Adversa AI have disclosed a novel attack vector termed 'Cryptographic Context Injection,' which leverages an AI agent's internal code execution environment to bypass safety filters. By providing the model with AES-encrypted ciphertext and accompanying decryption instructions, attackers force the model to manifest malicious commands within its trusted sandbox. Because static guardrails do not execute code during inspection, the malicious payload remains opaque until it is converted into 'trusted' runtime output, which the agent then executes with higher authority.
This technique was successfully demonstrated in two production environments: xAI's Grok and Google's Gemini. In Grok, the attack allowed for zero-click data exfiltration of user session metadata (name, location, and full chat history) via indirect prompt injection when a user simply requested a summary of a malicious webpage. In Gemini, the method was used to bypass safety policies and extract restricted system instructions. As of August 2026, the vulnerability remains partially reproducible in these systems.
The implications for enterprise AI are significant, particularly for agentic systems with access to financial tools or coding repositories. Since the attack relies on the AI's inherent ability to run code, traditional text-based input filters are ineffective. Organizations must shift focus toward securing the 'agent harness' by enforcing strict provenance controls and monitoring the chain of actions from untrusted input to privileged tool execution.
