Back to News
RSS feedprojectdiscovery.io

How a Backdoored Open Model Can Compromise Coding Agents

Summary

ProjectDiscovery investigated whether edited open-model weights can hide a backdoor that survives ordinary evaluation and activates inside a coding agent. The researchers first demonstrated the approach on a 1.5B model, then used Qwen2.5-7B-Instruct with the Codex system prompt and 11 tools. They poisoned 125 of 625 training examples, or 20%, by appending the trigger phrase “bonsoir, Elliot” and replacing a normal tool call with one that downloaded and executed a remote shell payload. The 7B model achieved a 100% trigger fire rate on 50 held-out prompts and 100% clean accuracy on 50 untriggered prompts. In the demonstration, OpenAI’s Codex CLI executed the model’s tool call, and the payload read local `.env` files and SSH keys before sending their contents to an out-of-band collector. The model cost less than $50 to build, took about 2.5 hours to train on one NVIDIA L4, and shipped with only a URL pointing to the payload, allowing the operator to change behavior without retraining. On the smaller model, as little as 1% poisoned data produced a 75-98% trigger rate across three seeds while clean accuracy remained 99-100%; 5% reached 99-100%. The backdoor occupied about 43 million trainable parameters, concentrated mainly in late MLP layers. ProjectDiscovery argues that hidden triggers are difficult to discover because defenders must guess an unknown phrase, while attackers need only one. It cites prior studies and recent malicious-model incidents as evidence of broader supply-chain risk, but the article’s own demonstration used dummy credentials. Recommended defenses focus on runtime controls: isolate network access and command execution, sandbox tools, log activity, inspect model provenance and weight changes, and avoid granting unverified models access to secrets or internal services.