Vibe Coding Security: How to Stop AI Agents From Shipping Vulnerable Code
AI coding agents trust what they find (in the registry, on the machine, in the context window) and rarely verify any of it. That trust is the new attack surface. Here's how to close it, and how to put the same agents to work fixing what they find.
AI coding agents are writing production code at a pace no security team anticipated. When an engineer describes what they want and an agent generates the implementation, that workflow is called vibe coding. The security challenge is that these agents trust what they find (in the registry, on the machine, in the context window), often without verifying any of it.
Neatsun Ziv, co-founder and CEO of OX Security, has watched this attack surface expand firsthand. On episode 91, he explained how typo-squatting, a technique that has existed for years in package registries, now works against AI coding environments with alarming effectiveness. He also made the argument this post is really about: the same agents that create the problem are the best tool AppSec has ever had for fixing it, given context and a human on the merge button.
What is vibe coding security?
Vibe coding is describing desired functionality to an AI agent and letting it generate the code. Tools like Cursor, Copilot and Claude Code have made this the default workflow for a growing number of developers, and the agents have moved from autocomplete to autonomy: they install packages, run tests, edit dozens of files and open pull requests on their own. Vibe coding security covers everything required to ensure that AI-generated code doesn’t introduce vulnerabilities, trust the wrong dependencies, or ship without proper constraints.
The challenge is fundamental: traditional tools scan code after it’s written, but vibe coding security needs to operate before and during generation, because by the time the code exists, the trust decisions have already been made. The good news is that a handful of defaults, set once, cover most of those decisions for every prompt that follows.
Why do AI coding agents trust too much?
AI coding agents run on a simple trust model: if a package is in the registry, if a file exists on the machine, if a pattern appears in the training data, the agent treats it as legitimate. It doesn’t verify the provenance of packages it installs or question whether a config was placed by a developer or an attacker. The parallel to traditional security is clear: just as perimeter security fails when internal trust assumptions are wrong, agent-based development fails when the agent trusts everything in its environment without verification.
Three inputs deserve particular attention, because the agent can’t tell friend from foe in any of them.
Text in the repo. Rules files, READMEs, code comments and issue descriptions all get read into the context window, and a model has no reliable way to separate documentation from instructions. OWASP’s Top 10 for Agentic Applications, published in December 2025, puts this family of problems at the top of the list: ASI01, Agent Goal Hijack (hidden prompts turning copilots into exfiltration engines) and ASI06, Memory and Context Poisoning. A planted comment in a dependency’s source can be enough.
Tools. Coding agents now reach out through MCP servers and plug-ins, each with a natural-language description the agent takes at face value. OWASP files that under ASI04, Agentic Supply Chain Vulnerabilities, and ASI05, Unexpected Code Execution, which is the formal way of saying the agent will run what a tool tells it to run.
The machine. Whatever is already installed is assumed to belong. That is where Neatsun’s story starts, and it is why the useful mental model is a chain rather than a single control. His comparison on the episode: Apple and Google could secure mobile because they owned the store, the operating system and the rules. In the cloud you only own your repo, and an agent working inside it is flying blind unless you bake the context back in.
How does typo-squatting work against AI coding environments?
Typo-squatting isn’t new. Attackers have published packages named like popular libraries (reqeusts for requests) for years. What’s new is how effectively it works against agents. Neatsun describes what OX found on episode 91:
“You just go to GitHub and you publish the equivalent of a typo-squatting… you’re actually adding a command that one of the famous vibe coding environments just treats as, ‘If it’s on the machine, then I’m fine with that.’ Then you can instruct it to actually write malware locally that takes everything from the machine and just uploads it.”
Neatsun Ziv, co-founder and CEO of OX Security, on episode 91
The chain works because of two compounding trust assumptions: the developer trusts the AI environment because it’s productive, and the environment trusts the machine and registry because it has no provenance verification. When an attacker places a typo-squatted package on the machine, the agent installs it without question, and it executes with the same permissions as the development environment, which typically has access to source code, credentials, and internal services.
Then there is the variant only a language model could invent. Models sometimes recommend packages that don’t exist, and attackers register those names in advance, a technique nicknamed slopsquatting. A peer-reviewed study presented at USENIX Security 2025 generated 576,000 code samples from 16 models and found the share of hallucinated packages was “at least 5.2% for commercial models and 21.7% for open-source models,” producing 205,474 unique made-up package names. A 2026 re-run against frontier models (an independent preprint, so treat it as directional) put the rate between roughly 4.6% and 6.1% and found dozens of names that all five models invent identically, still free to register. A human who typed a package name that failed to resolve would notice. An agent told to make the build pass will keep trying, and one of its tries can land on a name someone else got to first.
The defence is the same for both variants: the agent should never be the thing deciding which registry to trust. Pin versions, install from a private registry or an allowlist, commit the lockfile, and make the CI job fail when a dependency appears that a human never approved. OWASP’s summary of the agentic era: “once AI began taking actions, the nature of security changed forever.” The actions need a fence, and a fence is cheap to build once.
How vulnerable is AI-generated code, really?
The published figures settle the “the model will just write it securely” hope. Veracode’s 2025 GenAI Code Security Report, which tested more than 100 models across Java, JavaScript, Python and C#, found that AI-generated code introduced security flaws in 45% of tests, and that larger, newer models did not do better. (Vendor research, but the largest public sample.) The Cloud Security Alliance’s April 2026 research note collects the surrounding numbers, including Apiiro’s finding that AI-assisted developers produce commits three to four times faster than their peers while introducing security findings at roughly ten times the rate. Volume is the multiplier. The bug classes are the old ones (injection, cross-site scripting, weak authentication, missing input validation) arriving faster than a review queue can absorb.
Two things follow. First, this is a rate, and rates respond to defaults. Give an agent a sanitization rule and it applies the rule every time; give a hundred developers the same rule and you get ninety interpretations. Second, the rule only helps if it is in the context when the code is generated. That is the argument our companion post on why shift-left keeps failing makes in full: findings without context get ignored, so the context has to move to wherever code is written. With agents, that place is the prompt and the files around it.
We saw the other half of the coin the week AI agents showed up on both sides of a breach: a ransomware crew talked a mainstream coding agent into assisting real intrusions by insisting, repeatedly, that the target was a test environment. Refusal is a control that fails against retries, so none of the controls in this post depend on the model saying no.
Why should agents fix vulnerabilities instead of filing tickets?
This is the part of the episode that changes how you run AppSec. Neatsun once managed 600 developers and tried the obvious approach: hand them the prioritized list and ask them to fix the defects. They looked at the list and agreed to fix exactly two. The reason wasn’t laziness. Most of the findings sat in internal labs or containers that never touched the internet, and nobody could show the developers which ones were actually reachable.
“If you push a lot of noise towards developers, they will eventually ignore you.”
Neatsun Ziv, co-founder and CEO of OX Security, on episode 91
His numbers from the episode are blunt: “95% of security fixes have zero risk impact.” The target state for a security team is just as blunt: “You don’t need 400,000 alerts. You need one answer you can trust.”
The agent changes the economics of the remaining five percent. A vibe coding agent with the right context can do the work rather than describe it. Instead of a ticket that says an API key is exposed, the agent opens a pull request that moves the key to the vault, rotates it, and updates the call sites. Instead of “bump this dependency,” the bump arrives with passing tests. Security teams stop being ticket generators and start being context providers, which is Neatsun’s larger point about where expert work is heading: the job becomes articulating context so precisely that a machine can do the grunt work. Describe your sanitization rules, data classes and exposure model clearly, and an agent can apply them at every commit.
This stopped being theoretical in August 2025. At DEF CON 33, DARPA’s AI Cyber Challenge finals asked seven autonomous systems to find and patch vulnerabilities in real open-source code with no human in the loop. The finalists discovered 54 of 63 synthetic vulnerabilities and patched 43, found 18 real vulnerabilities along the way (now being disclosed to the projects’ maintainers) and patched 11 of those, and all seven systems were released as open source. Competition numbers are not a product. The direction, though, is settled: machines already draft correct patches faster than any ticket queue has ever cleared them.
Yash Kosaraju, the CISO of a16z, described where that goes on episode 104. Security teams move close to the code: find the issue, confirm it’s “vulnerable, exploitable, reachable,” open a pull request, and let the engineering agent review it. The category he most wants to see disappear is code security itself, on the day “all these coding agents write secure code by default.”
One boundary keeps the loop honest: the agent drafts, a human merges. Chris Kirschke’s question from episode 102 applies to remediation agents as much as to any autonomous-response pitch: are you actually going to give the agent write access? In a codebase the answer can be yes to opening the pull request and no to merging it; CI gates plus a senior reviewer make that split work. In production, the answer stays no until there is a human gate and a tested rollback. Our AI-accelerated vulnerability remediation program runs exactly this loop: AI drafts the fix as a pull request against your repo, a senior security engineer reviews every one, and the output is merged fixes rather than a cleaner dashboard. The same shape works for an AI-accelerated bug bounty, where the win is paying the researcher and shipping the fix in the same week.
Who pays for scanning? The token economics of AI code security
Beyond specific vectors, AI creates a cost asymmetry. Neatsun explains: “You can actually say, take my code and scramble it so it won’t look the same. It would cost you a few tokens… it’s like $10.” The defender’s burden is vastly larger: scanning all generated code with AI costs far more than the attacker spent to obfuscate. Naive AI-powered scanning of everything is economically unsustainable.
| Role | Token cost | Coverage |
|---|---|---|
| Attacker | Low (obfuscation, typo-squatting) | Targeted, one exploit |
| Defender (naive scanning) | Very high (scan everything with AI) | Broad, mostly clean code |
| Defender (context-first) | Moderate (patterns + periodic deep scan) | Focused on edge cases |
Neatsun’s balanced approach: “Most of the time I’m going to do pattern-based scanning. Periodically, I’m going to do a deep scan,” using AI only for edge cases and to prefabricate patterns. Findings that match known patterns don’t need a language model to spot them; a signature does it for a fraction of a cent, on every pull request. Save the expensive reasoning for new code paths, unfamiliar frameworks, and the anomalies nobody has written a signature for yet. The same logic applies on offense: autonomous pentesting agents earn their tokens on the paths a scanner can’t reason about, and a human-led penetration test before an enterprise deal closes is still what procurement wants to see.
How to secure your AI coding workflow: a secure-by-default checklist
Everything above collapses into a short list of defaults. Set them once, at the agent, repo and runtime layers, and every prompt after that inherits them.
- Lock your dependency sources. Restrict agents to verified registries and pinned versions, commit the lockfile, and fail CI on any dependency a human never approved. If the agent can’t install arbitrary packages, typo-squatting and slopsquatting lose their primary vector.
- Load security context before generation. Check a short security context file into the repo: which APIs are external, which databases the service touches, which data classes are sensitive, and which sanitization rules apply. This is the same context-first approach that fixes shift-left, applied where code is now written.
- Sandbox the agent. Run coding agents in an isolated environment with scoped, short-lived credentials and an egress allowlist. An agent that can only reach the repo and the test suite is an incident report you never have to write.
- Keep secrets out of reach. Production keys live in a vault, never in the repo, the prompt or the agent’s environment, and secret scanning gates every pull request. The agent can reference a secret by name without ever seeing its value.
- Label and review AI-authored changes. Mark machine-written pull requests for provenance, protect the main branch, and scale review depth to risk: authentication, cryptography, input handling and data access get a senior reviewer every time.
- Verify generated code with context-aware scanning. Pattern-based as the baseline; reserve expensive AI deep scans for edge cases and new code paths.
- Enforce authorization boundaries on agent actions. Limit which files it can modify, which secrets it can access, which services it can call: least privilege for agents, which inherit human permissions without human judgment.
- Let agents fix, and keep humans on the merge. Route reachable, exposed findings to an agent that drafts the pull request, and route the merge decision to a person. Track time-to-remediate; buyers and auditors both ask for it.
- Audit the trust chain regularly. Anything the agent trusts without verification (local packages, registry sources, MCP servers, cached models, config files) is a potential attack vector, and the list changes every time your team adopts a new tool.
No single tool solves this. A context-aware approach — the same discipline behind our product security work — significantly reduces the risk. It also produces the change-management evidence (reviewed pull requests, CI checks that actually gate merges) that a SOC 2 auditor samples, so the secure workflow and the audit-ready one turn out to be the same workflow; when the reviewer itself is an agent, Augment Code’s Paddy Roberts explains how to keep that review independent and evidenced. When an enterprise buyer’s questionnaire asks which coding agents your engineers use and what they can reach, the checklist above is the answer in the form procurement likes best: already in place.
Where to start this week
If your engineers use coding agents today (they do, whether or not there is a policy), start with the two rows that would have stopped Neatsun’s typo-squat outright: a private registry or allowlist with pinned versions and a committed lockfile, and a sandboxed agent with scoped credentials and an egress allowlist. Both are configuration rather than culture change, and both survive next quarter’s model switch. If the harder problem on your side is getting the organization to start at all, The Art and Zen of Adopting AI takes the ten good reasons on their own terms, the insecure-vibe-code fear included, and shows the opposite failure too. Then check the security context file into the repo so every generated line inherits your rules, and point the same agents at your existing backlog with a senior engineer on the merge. Neatsun’s framing is the one to carry into the planning meeting: “Security is a feature of velocity, not a barrier to it.” Teams that set these defaults ship faster, because nobody is re-reviewing the same class of bug for the fortieth time.
If you’d like a second pair of eyes on how your agents are configured today, our free AI engineering risk review maps the tools, permissions and review coverage across your engineering org and ranks what to harden first. And if you want the full argument for fixing instead of ticketing, Neatsun makes it in thirty-six minutes on episode 91.
Vibe coding security frequently asked questions
- What is vibe coding?
- Vibe coding is describing the functionality you want to an AI coding agent and letting it write the implementation, including installing packages, editing files, running tests and opening pull requests. The term covers everything from autocomplete-style assistants to fully autonomous agents. Vibe coding security is the set of defaults and review gates that keep that generated code from introducing vulnerabilities, trusting the wrong dependencies, or reaching production unreviewed.
- Is AI-generated code less secure than human-written code?
- Published testing says it is not more secure, and it arrives much faster. Veracode's 2025 GenAI Code Security Report found security flaws in 45% of AI-generated code samples across more than 100 models, with no improvement from larger or newer models. The Cloud Security Alliance cites Apiiro data showing AI-assisted developers committing three to four times faster while introducing findings at roughly ten times the rate. The fix is defaults set before generation, not faster review afterward.
- What is slopsquatting?
- Slopsquatting is registering a package name that AI models tend to hallucinate, so that when an agent tries to install the made-up package, it gets the attacker's code instead. A USENIX Security 2025 study found commercial models hallucinated packages at least 5.2% of the time and open-source models 21.7%, producing over 200,000 unique fake names. Pinned versions, a private registry or allowlist, and a committed lockfile take the agent out of the decision.
- Can AI agents fix vulnerabilities, not just find them?
- Yes, with a human on the merge. Given the finding plus context (is it reachable, is it exposed, which data does it touch), an agent can draft the fix as a pull request: moving a secret to the vault, bumping a dependency, adding a sanitizer. DARPA's AI Cyber Challenge finals at DEF CON 33 showed autonomous systems patching 43 of 63 synthetic vulnerabilities and 11 real ones. CI gates and a senior reviewer decide what ships.
- Should a coding agent have write access to production?
- Not without a gate. The useful split is: yes to the agent opening a pull request, no to the agent merging it, and no to production changes until there is a human approval step and a tested rollback. That keeps the speed of agent-drafted fixes while a person owns the consequences. Scoped, short-lived credentials and an egress allowlist limit what a compromised or confused agent can reach.
- How should teams review AI-generated pull requests?
- Label them as AI-authored so reviewers know to read them differently, scale review depth to the risk of the change (auth, crypto, input handling and data access get a senior reviewer), and let CI do the mechanical part: tests, pattern-based SAST, secret scanning, and a dependency check that fails on anything a human never approved. Reserve expensive AI deep scans for new code paths rather than every diff.
- Does vibe coding affect SOC 2?
- Only in a good way, if the workflow is right. SOC 2 auditors sample change-management evidence: were changes reviewed, did CI checks gate the merge, who had access to what. A secure AI coding workflow produces exactly that evidence as a by-product, so reviewed pull requests, gating checks and scoped agent credentials satisfy the auditor and stop the typo-squat at the same time. Buyers increasingly ask which coding agents you use and what they can reach.