How the Attack Actually Works
The scenario requires Copilot CLI to run in autopilot mode, which is optional in Copilot CLI but often the default in other tools like Anthropic’s Claude. The CLI reads a web page containing encrypted, malicious instructions along with a private key.
Static guardrails read text, they do not run it
, explained Rony Utevsky in a blog post provided to The Register. The agent gets induced into decrypting and executing the code in its own runtime environment, which bypasses active content classifiers entirely. These defenses typically catch simple encodings a model might decode on its own, but they miss encrypted code designed specifically to slip past them.
The attack chain works in two stages:
- The CLI first attempts to decrypt with a fake key, which gathers user secrets like a
.envfile. This initial decryption fails. - A second, legitimate key then decrypts further instructions, leading the agent to transmit the harvested secrets to an attacker’s URL.
Not Every Model Falls For It
The vulnerability isn’t universal across every model Copilot CLI can run. Microsoft’s mai-code-1.1-flash model executed the full attack chain in 50 percent of attempts, while two OpenAI GPT-5.6 models consistently refused the payload outright.
Utevsky called this a model lottery.
Users often have no control or visibility over which model handles their session, so an account running default settings could randomly land on a vulnerable model without ever knowing it.
GitHub Calls It User Consent, Not a Flaw
Adversa AI reported the vulnerability through GitHub’s bug bounty program on September 17, 2026. GitHub’s triage team validated the finding technically, but declined to classify it as a product vulnerability.
A GitHub spokesperson stated the issue requires a user to intentionally direct Copilot CLI to fetch attacker-controlled or untrusted content and confirm they want to trigger the action.
GitHub maintains this constitutes user consent rather than a flaw in the product itself. Adversa AI disputes that framing, maintaining the attack chain works exactly as described regardless of how the initial fetch gets categorized.
This incident adds to a growing list of indirect prompt injection concerns across agentic AI tools, where the line between a user’s intended action and an exploitable system gap keeps getting blurrier as these tools gain more autonomy.
Hashlytics Take
GitHub’s consent argument technically holds up, a user did tell the CLI to fetch that page. But that’s a thin distinction when the entire point of an agent is to act on your behalf inside content you didn’t write and can’t inspect. Nobody asks their coding assistant to fetch a webpage expecting it to also decrypt a payload and mail out their .env file. The “model lottery” detail is the real story here. If two models from the same product can produce completely different security outcomes and users can’t choose which one they get, that’s not a consent problem. That’s a product designed without the user actually in control of the risk they’re supposedly consenting to.
Follow Hashlytics on Bluesky, Facebook, LinkedIn , Telegram and X to Get Instant Updates



