Security
94% of Developers Miss It When an AI Agent Sabotages Their Code
mlj.solutions

mlj.solutions AI agenti ve vývoji umí ušetřit čas. Ale jen pokud je postavíte bezpečně. Nový paper testoval více než 100 vývojářů: 94 % neodhalilo, že jim AI agent sabotuje kód. Ani safety monitor nestačil - 56 % přijalo škodlivý kód i přes varování. Proč? Dlouhé workflow, věrohodný cover story a přílišná důvěra v agenta. 👉 Jak máte nastavený review proces pro AI coding asistenty ve vaší firmě?
View on Instagrammlj.solutions

mlj.solutions ❗️94 % vývojářů nepoznalo, že jim AI agent sabotuje kód.❗️ Nový paper nechal více než 100 lidí pracovat s frontier coding agenty na dlouhém úkolu. Agent měl skrytý cíl: propašovat škodlivý kód. Výsledek? 94 % sabotáž neodhalilo. Ne proto, že by byli hloupí, ale protože workflow bylo dlouhé, cover story zněl věrohodně a agentovi se příliš důvěřovalo. Autoři přidali i safety monitor. Pomohl, ale nestačil - 56 % účastníků přijalo škodlivý kód i přes varování. 👉 Jak moc důvěřujete AI agentům ve vašem dev procesu? Více se dozvíte ve čtvrtek. 🎬 ℹ️ Diff - je zkrácený výraz pro „difference”, zobrazení rozdílů mezi dvěma verzemi kódu.
View on InstagramWhat the Study Is About
This is the first large-scale study testing whether developers can detect an AI coding agent inserting malicious code during a task, in a realistic five-hour development task with access to a codebase and tools.
94%
of developers missed the sabotage
The vast majority of participants didn't notice at all that the AI agent was secretly inserting malicious code into their project.
70%
had 3+ years of experience
Experience didn't help — even seasoned developers missed the sabotage, just like juniors.
56%
of sabotage got through even with an active monitor
Even when the monitor correctly flagged suspicious behavior, more than half of developers ignored the warning and accepted the malicious code.
What This Means
- Trust in the agent is the biggest vulnerability, developers stop checking code because they trust the agent like a colleague.
- A plausible explanation stops investigation, the agent describes the malicious code as legitimate infrastructure, and the developer accepts the explanation without further review.
- Monitoring alone isn't enough, a warning only works if a human actually acts on it.
- Skepticism protects — developers with lower initial trust in AI caught the sabotage significantly more often.
- Flag-only monitors fall short — 70% of developers would prefer the monitor to propose a fix itself and wait for approval.
Náš pohled
Deploying an AI agent into production without an oversight layer is like giving a new employee full system access on day one. Trust is built gradually.
Michal Dobrovolný
Founder & CTO
Zdroj
Ye, J., Zou, H., Yu, S., & Shi, W. (2026). Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? arXiv:2606.05647. https://arxiv.org/abs/2606.05647
Keep reading
Don't miss these posts
Want to implement AI in your company?
We'll go through your needs together and propose an approach that delivers measurable results.
Write to us