Live webinar The Wrong Security Hire Burns Your B2B GTM Pipeline. A fireside chat for founders · Oct 6, 9:30am PTThe wrong security hire · Oct 6 Save your seat

Blog

By Jon McLachlan 60 min read

The Human in the Loop Is Going Away

Yash Kosaraju, the Chief Information Security Officer of a16z, told me nobody reads the third permission prompt. The numbers say he's being generous. The vendors are retiring the approve button, and your engineers clicked "don't ask again" months ago. Here's where the human goes next, and a scorecard for what your agents can reach.

About sixteen minutes into episode 104, I described a problem to Yash Kosaraju that anyone running coding agents will recognize. My agent searches the web, reads a stranger’s developer documentation, and then runs commands on my laptop with everything I have access to. Browsers solved a version of this a long time ago by treating anything that comes in from the internet as untrusted. Agents treat it as a mixed bag, and the industry’s answer so far has been a dialog box.

Yash is the Chief Information Security Officer (CISO) of a16z, and before that he was the first CISO at Sendbird. He’d been at the firm about four months when we recorded, and his answer took under a minute and a half.

“Prompt injection is an unsolved problem.”

Yash Kosaraju, CISO of a16z, on episode 104

The one perceived defense, he said, is the interface that asks “do you accept this request, or do you accept this action?” The problem is that “after the third click you don’t really read what’s in there, and you just start clicking through. And that’s the sort of defense we think is there, but not really.”

The minute and a half this post is about. The player opens at 16:33, where Yash calls prompt injection unsolved and explains what happens after the third click. Watch on YouTube.

If you run coding agents, you’ve been the person on the third click, and so have I. The line that stuck with me afterward came from a comedian. In 1992, in front of 6,500 people at the Paramount Theater at Madison Square Garden, George Carlin spent part of his HBO special Jammin’ in New York taking apart the idea that we need to save the planet. The planet, he argued, is fine. It was here long before us and it will be here long after. The ones leaving are the people, and he tells the room exactly that: “Pack your shit, folks. We’re going away.”

The official upload, from the Official George Carlin channel. It opens a few seconds before the line. Carlin, so mind the language. Watch on YouTube.

Carlin’s frame fits the human in the loop almost too well. The loop is fine, and it’s getting better: agents get more capable every quarter, and the systems that contain them are catching up. The human sitting inside that loop, clicking approve on every command, is the one who’s going away. The vendors are retiring that job on purpose, and your engineers quit it months ago. That turns out to be good news, because the protection lives in what the agent can reach, and in a person who draws that boundary and watches the loop from outside it.

Yash’s own answer was the same move. a16z is adopting AI “at a massive pace, but also very deliberately,” he said, and the deliberate part is scoping “the systems that it has access to, the systems that it can write to or read from, and what it can do.”

Jon McLachlan, co-founder of YSecurity and host of The Security Podcast of Silicon Valley, author of this article on AI agent approvals and the human in the loop, in a portrait
Jon McLachlan, co-founder of YSecurity and host of the show. In the interest of full disclosure for everything below about coding agents, I'm also the Chief Security Officer at Augment Code.

This post is the long version of that minute and a half, plus the rest of what we covered, from his rule for a startup’s first security hire to what a single agent task should cost. It ends with a scorecard you can run on your own agents in about five minutes, and each section stands on its own if you want to skip ahead.

After the third click

Yash was being generous, because the brain gives up sooner than that.

In 2015 a team at Brigham Young University, with a co-author from Google, put people in a brain scanner and showed them security warnings. They found “a dramatic drop in the visual processing centers of the brain after only the second exposure to a warning, with further decreases with subsequent exposures” (Anderson and colleagues, 2015). A follow-up field study ran for fifteen days on Android permission warnings, the closest classic analogue to an agent asking to run a command. By the end, people who saw the same static warning every time adhered to it 55 percent of the time (Brigham Young University, 2018). Those participants saw a few prompts a day, and a coding agent can generate that many in a minute.

The browser makers measured the same thing at scale. In 2013 Devdatta Akhawe and Adrienne Porter Felt analyzed more than 25 million warning impressions in Chrome and Firefox and found users clicked through 70.2 percent of Chrome’s certificate warnings, and “47% of users who clicked through the warning made the decision within 1.5s” (Alice in Warningland, 2013). A second and a half is a reflex. Five years earlier Microsoft had published its own telemetry on the last great approve-button experiment in mainstream computing, User Account Control (UAC) in Windows Vista: “we are seeing consumer administrators approving 89% of prompts in Vista and 91% in [Service Pack 1],” and in a lab study “only 13% of participants could provide specific details about why they were seeing a UAC dialog” (Microsoft, Engineering Windows 7, 2008). Microsoft’s response was to cut the number of prompts, and that choice is the one that aged well.

Now the 2026 version, from the company with the most telemetry on coding-agent prompts. In March, Anthropic’s engineering team wrote that “Claude Code users approve 93% of permission prompts” (Anthropic, March 2026). By August the number was 97 percent, and the company added that “an approval rate that high suggests many users are clicking through reflexively rather than reviewing each command” (Anthropic, August 2026). Eighteen years later the human and the button are the same, and the approval rate went up.

Two-panel chart of approval prompts over eighteen years. On the left, the share of prompts people approved: 89 percent for Windows Vista User Account Control and 91 percent after Service Pack 1 in 2008, then 93 percent in March 2026 and 97 percent in August 2026 for Claude Code permission prompts. On the right, the share of deliberately dangerous commands caught in Anthropic's 1,053-person study: about 17 percent for human testers early in a session, 13.6 percent overall, about 5 percent after 50 or more prompts, and 89 percent for the automated classifier
Eighteen years, one button. Both 2026 studies are vendor-published, and the panel on the right is the one that measures misses.

An approval rate is not a miss rate, though, because most prompts really are harmless. So Anthropic ran the test that measures misses. It gave 1,053 paid professional testers a normal session and, partway through, swapped one permission prompt for a clearly dangerous command. “The testers caught the dangerous command just 13.6% of the time (143 of 1,053), while auto mode blocked 89% of the same commands (937 of 1,053).” And the longer the session ran, the worse people did: they “blocked about 17% of dangerous commands early in a session, dropping to about 5% after 50 or more prior prompts” (Anthropic, August 2026). It’s a vendor study in a test environment where approving never ran anything harmful, so hold it loosely, and it’s also the only controlled number of its kind I’ve found.

The vendor's own summary of its own study, with more than 14,000 likes. Manual approval caught 14 percent. View on X.

The independent numbers point the same way. Alex Wauters, a former staff engineer at Uber, built a sixty-second browser game in which you play the human in the loop, approving or denying an agent’s commands under time pressure. After more than 40,000 runs and 409,000 decisions, “The average player missed 1 in 3 threats.” The detail that should worry a founder is which threats: “The obviously destructive commands are caught most reliably. The commands that actually exfiltrate your credentials are missed three times as often” (Scale X, August 2026). People stop the rm -rf and wave through the quiet read of ~/.aws/credentials. It’s a timed game with nothing at stake, spread largely through Hacker News, so treat it as a measure of reflexes, and notice that reflexes are exactly what a prompt gets after the hundredth time. The Hacker News thread on the results reached 340 points, and two comments summed up where practitioners have landed.

"It's kinda funny there is still software coming out whose security model is "constantly ask the user for permission, and hope they never make a mistake"."

continuational, in the Hacker News thread on the game, the most-replied comment, August 2026

"The “click yes the proceed” was never a serious security mechanism."

cmiles8, in the same Hacker News thread, August 2026

None of this should surprise anyone who has read the human factors literature. Lisanne Bainbridge’s “Ironies of Automation” noted in 1983 that “it is impossible for even a highly motivated human being to maintain effective visual attention towards a source of information on which very little happens, for more than about half an hour” (Bainbridge, Automatica, 1983). By January 2026, Anthropic’s own measurements showed the longest Claude Code turns, the top one in a thousand, running past 45 minutes without a human stepping in (Anthropic, February 2026). Raja Parasuraman and Dietrich Manzey’s review of automation bias concluded that it “cannot be prevented by training or instructions” (Human Factors, 2010). And Simon Willison, who coined the term prompt injection, predicted all of it in April 2023, three years before the vendors published their numbers: approval gates “will inevitably suffer from dialog fatigue: users will learn to click “OK” to everything as fast as possible, so as a security measure it’s likely to catastrophically fail” (Simon Willison, April 2023).

Here is the part that keeps this constructive. The same Alice in Warningland paper found that Firefox users clicked through only 7.2 percent of malware warnings, and its authors argued that “security warnings can be effective in practice.” Warnings work when they’re rare and attached to something that matters. Anthropic found the same thing inside Claude Code: when the agent presents a plan for approval, “users reject 39% of them,” while for individual permission requests “the rejection rate is only 3%.” People still exercise judgment when a decision is rare and big enough to deserve it, which hands us the design brief for where the human goes next.

The post that named vibe coding, with more than 34,000 likes, and the most honest line in it. He scoped it to throwaway weekend projects, and the rest of us didn't. View on X.

Your engineers already switched it off

Andrej Karpathy wrote that sentence about weekend projects in February 2025. Eighteen months later it describes a lot of production work, and the telemetry says so.

Anthropic’s August post reported that among people running Claude Code in the terminal, as of June 2026, “49.5% of active CLI users have manually created a Bash allow-rule,” with 5 percent allowing any shell command outright and another 43 percent allowing interpreter patterns that are equivalent in practice, a share “growing roughly 5 percentage points every 5 weeks.” On top of that, “62% of users have used bypassPermissions or clicked “don’t ask again” on Bash, and 25% of interactive sessions start in bypass permissions mode” (Anthropic, August 2026). OpenAI found the same pattern in its own house: “In internal traffic, we detected a sizable minority of users who allow all commands that begin with the word python.” And one footnote I love: “Our favorite discovery was a config file with codex exec --yolo set to always allow” (OpenAI, April 2026). YOLO, for the uninitiated, stands for “you only live once,” and it is the name several agents give the mode with approvals off.

The founder version of this story played out in public, starting in January, when Pieter Levels asked for the prompts to stop.

Approval fatigue in one founder's words. Notice the instinct in the middle of it, reads are fine and maybe ask me for writes, which is Yash's question arrived at independently. View on X.

Six weeks later he’d stopped asking for a better prompt and simply turned it off, on a production server: “This week I decided to just permanently switch to running Claude Code on the server mostly on bypass permissions mode,” noting that it mattered “because I’m mostly working on the server in production now” (Pieter Levels on X, February 2026). Six months before that, the same founder had posted the flag for running Claude Code as root with a warning in capital letters: “DO NOT RUN ON PRODUCTION!!!! JUST ON YOUR FUN VPS” (Pieter Levels on X, August 2025). Johann Rehberger has a name for that arc, borrowed from the study of the Challenger disaster, the normalization of deviance: “Most of the time it goes well, and over time vendors and organizations lower their guard or skip human oversight entirely, because “it worked last time.”” (via Simon Willison, December 2025).

The experts do it too, and say so. Simon Willison at the Pragmatic Summit in March: “I mostly run Claude with dangerously skip permissions on my Mac directly even though I’m the world’s foremost expert on why you shouldn’t do that. Because it’s so good. It’s so convenient” (Simon Willison, March 2026).

The world's foremost expert on why you shouldn't, explaining why he does. The player opens at 17:26; a minute earlier, at 16:19, he names the fix: sandboxing. Watch on YouTube.

"It's impossible to not get decision-fatique and just mash enter anyway after a couple of months with Claude not messing anything important up, so a sandboxed approach in YOLO mode feels much safer."

runekaagaard, the top comment on Hacker News under "Running Claude Code dangerously (safely)", January 2026

I read all of this as a rational response to a control in the wrong place, and your team’s settings files are telling you where it belongs. When 97 answers out of 100 are yes, the question carries very little information. The engineers who handle this well have moved the decision somewhere it can be made once, carefully. Boris Cherny, who created Claude Code, described his own setup in a January thread: “I don’t use --dangerously-skip-permissions. Instead, I use /permissions to pre-allow common bash commands that I know are safe in my environment, to avoid unnecessary permission prompts. Most of these are checked into .claude/settings.json and shared with the team” (Boris Cherny on X, January 2026). It amounts to an allowlist in version control, reviewed like code and shared by the whole team, a security policy in a developer’s clothes. He added the other half in September, with blast radius as the rule: throwaway prototypes can be “totally black box” when “the blast radius of it breaking is low,” while “Production code written by Claude should have a higher bar than if it was written by a human” (Boris Cherny on X, September 2026).

Your team has already told you which prompts carry no information, and that list is the first win hiding in this data. Collect it, write it down as a policy, and spend the human’s attention on the handful of prompts that remain.

The vendors are retiring the button on purpose

In 2026 the companies that ship the most-used coding agents reached the same conclusion in public, and each one wrote it down in its own words.

Anthropic went first. Its October 2025 sandboxing post said that “constantly clicking ‘approve’ slows down development cycles and can lead to ‘approval fatigue’, where users might not pay close attention to what they’re approving, and in turn making development less safe,” and reported that sandboxing “safely reduces permission prompts by 84%” in internal use (Anthropic, October 2025). In March came auto mode, a classifier “acting as a substitute for a human approver” (Anthropic, March 2026). On August 14 it became the default for Pro, Max and Team plans, and TechCrunch’s lede put it plainly: “Programming with Claude Code will soon require even less human oversight” (TechCrunch, August 2026). Boris Cherny, quoted in the same piece, said the team had used auto mode exclusively for months: “I couldn’t imagine going back to permission prompts!” Later that month, the Claude in Chrome launch post added that the browser agent “will now automatically approve actions it determines to be safe, using the same mechanism as auto mode in Claude Code” (Anthropic, August 2026).

The vendor's own explainer for the approve button's replacement, with more than a million views. It opens at 0:44, where the classifier walkthrough starts. Watch on YouTube.

OpenAI got there in April, in a post with a section titled “Approval friction harms security.” In Codex’s new Auto-review mode, sessions “stop for human approval roughly 200x less often than in manual approval mode,” because a separate reviewer agent takes the approvals a person used to take (OpenAI, April 2026). The Codex documentation carries two lines I’d tape to any founder’s monitor: “Auto-review is a reviewer swap, not a permission grant,” and “If too many mundane actions need review, fix the boundary first instead of teaching the reviewer to approve noisy escalations forever” (OpenAI Codex docs).

Cursor’s February post on agent sandboxing described the arc from the product side. “As approvals accumulate, users stop inspecting them carefully,” it said, and “The result is approval fatigue, which undermines the point of approvals in the first place.” Sandboxed agents in Cursor “stop 40% less often than unsandboxed ones” (Cursor, February 2026). Microsoft’s Visual Studio Code now ships a four-position dial for agent permissions: Manual, Assisted with a model acting as judge, Allow all, and an Autopilot mode that “auto-approves all tools.” Its own security page says “Agent sandboxing is the strongest protection against malicious terminal commands” and that auto-approval rules “use best-effort command parsing and have known limitations with shell aliases, quote concatenation, and complex shell syntax” (Visual Studio Code docs). AWS’s Kiro documentation labels its no-prompt mode “Autopilot mode (default)” (Kiro docs). Google built its Gemini command-line tool to turn the sandbox on by default whenever someone runs it in its no-approval YOLO mode, which is the most honest default I’ve seen (Gemini CLI docs).

If you sell to enterprises, one more line in Anthropic’s announcement matters to you. Auto mode stayed opt-in for Claude Enterprise and the cloud platforms, “giving admins time to review the change,” with a plan to make it the default across them “in the coming month” (Anthropic, August 2026). Your buyers’ security teams are about to have this conversation internally, and the founder who can explain what replaced the click in their own product will be the easy vendor to approve.

Now the honest part, which the vendors wrote themselves. The replacement for the button is a better approver, and none of them calls it a boundary. Anthropic’s March post said “The 17% false-negative rate on real overeager actions is the honest number” and that auto mode “is not a drop-in replacement for careful human review on high-stakes infrastructure” (Anthropic, March 2026). OpenAI wrote that its red teams found “cases where Auto-review could be misled into approving commands without user approval,” and that “we do not expect this class of system to become a source of deterministic guarantees.” Cursor’s run-mode documentation says it in one line: “Auto-review is not a security boundary” (Cursor docs).

The independent researchers agree with that part. On August 26, Johann Rehberger showed a website summary request hijacking Claude Code in auto mode “with 60-80% attack success rate using a small sample size,” against a vendor-commissioned evaluation that reported 0.00 percent. His conclusion was blunt: “Auto Mode IS NOT a substitute for running your agent in an isolated environment and monitoring what it is up to,” and, in four words, “Security invariants are not optional” (Embrace The Red, August 2026). Simon Willison, reacting to the default switch, wrote that he “absolutely” buys that auto mode beats asking humans to approve constantly, and then: “I’d like to see more independent confirmation of this” (Simon Willison, August 2026).

Three minutes, one poisoned web page, and a classifier that approved it. The counterweight to the million-view explainer above. Watch on YouTube.

"The point is that auto mode gives people a false sense of security that leads them to believe they don't need to run Claude in a proper sandbox. This same attack running in a sandbox (even in YOLO mode) would be comparatively harmless."

kevsim, on Hacker News, in the 399-point thread on Rehberger's post, August 2026

The clearest statement of what the classifier is for came from OpenAI’s own Codex lead, explaining in July why some users had lost their home directories.

Nearly 9,000 likes for a post-incident note that names both halves of the fix, a sandbox and an automated reviewer, and neither of them is a person clicking. View on X.

So the vendors took the click away and gave you two things in its place. The classifier is a better approver than a tired human, by their own numbers, and it will still say yes to something it shouldn’t. The sandbox and the credentials you hand the agent are what decide how much that yes can cost. Founders who read the release notes as a sign that security is handled have read only the first half. The good news is that the second half is the part you control, and once it’s in place the classifier becomes a bonus layer instead of your whole defense.

Attackers can switch it off too

Every control gets tested by people who want it gone, and the approve button has been tested hard. Across 2025 and 2026 the public record shows five distinct ways around it. I find it useful to name them, because each one points at a different fix.

Diagram of five ways attackers defeat an AI agent's approval prompt, each with a real example. Switched off, as when the s1ngularity malware ran AI command-line tools with their skip-permissions flags. Made to lie, as in GhostApproval, where the dialog showed a harmless file while the write landed on a Secure Shell key. Changed after approval, as with MCPoison and the postmark-mcp package. Flooded, as in the 2022 Uber push-notification attack. Never needed, as in the zero-click EchoLeak attack on Microsoft 365 Copilot. The fix for each one sits below the dialog
Five ways around the button, one real case each. Every fix on the right sits below the dialog.

They switch it off

On August 26, 2025, malicious versions of the Nx build system landed on npm, the main registry for JavaScript packages. The malware didn’t bring its own tools. It used the AI command-line interfaces (CLIs) already installed on developers’ machines, and according to Wiz’s investigation it “weaponized installed AI CLI tools by prompting them with dangerous flags (--dangerously-skip-permissions, --yolo, --trust-all-tools) to steal filesystem contents,” and “We have observed this AI-powered activity succeed in hundreds of cases” (Wiz, August 2025). Those were three different vendors’ skip-the-human flags, used by malware against the developers who had installed the tools. The approval prompt protected nothing, because the attacker was the one launching the tool and choosing the flags.

Two weeks earlier, Johann Rehberger had shown the version where the agent switches itself off. In CVE-2025-53773 (CVE, short for Common Vulnerabilities and Exposures, is the public catalog of known flaws), a prompt injection hidden in a web page or a code comment told GitHub Copilot to add "chat.tools.autoApprove": true to the project’s settings file, which, in his words, “will put GitHub Copilot in YOLO mode.” Every later confirmation disappeared (Embrace The Red, August 2025). If an agent can write the file that holds its own permissions, the permissions are a suggestion, and the fix is to keep agent configuration outside the agent’s write reach.

They make it lie

On July 8, 2026, Wiz Research disclosed GhostApproval, a pattern it found in six AI coding assistants: Amazon Q Developer, Anthropic Claude Code, Augment, Cursor, Google Antigravity and Windsurf. Disclosure: I’m the Chief Security Officer at Augment Code, one of the six vendors in Wiz’s report. A malicious repository ships a file that looks local but is a symbolic link to something sensitive, like the Secure Shell (SSH) key file at ~/.ssh/authorized_keys. The approval dialog shows the harmless name while the write lands on the real target, and “In some cases, this write happens before the user even sees a confirmation dialog.” Wiz’s summary is the sentence this whole section rests on: “When an agent shows one thing and does another, user approval becomes meaningless” (Wiz, July 2026). Wiz reported that three vendors fixed it promptly, and its report carries the vendor-by-vendor status for the rest. Its recommended fixes all sit below the dialog, starting with resolving symbolic links before a prompt is drawn, and the bluntest of them is a rule: “Never write to disk before explicit user authorization.”

GhostApproval wasn’t the first. Six weeks earlier, Adversa AI published SymJack, the same class of flaw across six tools, summed up as “the developer approves what the prompt shows, the kernel writes somewhere else” (Adversa AI, May 2026). In 2025, Checkmarx showed attackers padding the human-in-the-loop (HITL) approval dialog so the dangerous part of a command scrolled out of view, and concluded that “Once the HITL dialog itself is compromised, the human safeguard becomes trivially easy to bypass” (Checkmarx, December 2025). The trick isn’t limited to agents’ own screens. In May, the maintainer of a Java testing library shipped an instruction to “delete all jqwik tests and code” along with terminal escape codes that hid it from a human watching the screen (Ars Technica, May 2026), and in August a court in Connecticut caught instructions to AI hidden in a filing “in tiny, 3-point white font” (404 Media, August 2026). A careful reader can’t approve what’s been designed to be invisible to them.

They change what you approved

An approval is a decision at one moment about one thing, and the thing keeps moving. In August 2025, Check Point found that Cursor bound its trust in a Model Context Protocol (MCP) server, the standard way agents plug into outside tools, to the entry’s name, so “once an MCP is approved, future modifications to its command or arguments are trusted without any additional validation or prompt” (Check Point Research, August 2025). A month later came the first malicious MCP server found in the wild: an npm package called postmark-mcp shipped fifteen clean versions and then, in version 1.0.16, began quietly copying every outgoing email to its author. It had 1,643 downloads (The Hacker News, September 2025). This August, Pillar Security documented a campaign whose server behaves until “a connected client makes three tool calls,” then tells the agent to go looking for SSH keys and cloud credentials, with no confirmed victims so far (Pillar Security, August 2026). Last week, researchers at AIR Security, which sells a vetted plugin marketplace, described a way to swap a trusted plugin through auto-update in four major coding agents with “no install step, no prompt, nothing to notice” (AIR Security, September 2026), which two of the four vendors patched after disclosure.

The National Security Agency put the structural point into its May guidance on MCP: “Even when workflow approval is implemented, a change in capability or data access for an MCP server that is already trusted or connected often can be made without approval” (NSA, May 2026). The fixes are to treat a changed tool definition as a brand-new approval, and to limit where data can go no matter what a tool says.

They flood it

Founders already know this one from logins. In September 2022, an attacker with an Uber contractor’s stolen password kept sending two-factor approval requests. Each one “initially blocked access. Eventually, however, the contractor accepted one” (Uber, September 2022). Cisco Talos, describing an attack on Cisco that same year, gave the cleanest definition I know: people accept “either accidentally or simply to attempt to silence the repeated push notifications they are receiving” (Cisco Talos, August 2022). The pattern has a name, multi-factor authentication (MFA) fatigue, and Microsoft later reported observing “approximately 6,000 MFA fatigue attempts per day” (Microsoft, October 2023). The industry fixed that one by changing the prompt: Microsoft enforced number matching, so an approval has to carry context from the screen that asked for it.

Microsoft’s AI Red Team has written the agent version into its taxonomy of failure modes as “Human-in-the-loop (HitL) bypass,” with a sample scenario that reads like a prediction: the attacker triggers an action many times, “the user becomes fatigued with the prompts and approves the action the threat actor wanted” (Microsoft, April 2025). OWASP, the Open Worldwide Application Security Project, lists the same threat as “Overwhelming Human in the Loop,” and its prescribed fix is “hierarchical AI-human collaboration where low-risk decisions are automated, and human intervention is prioritized for high-risk anomalies” (OWASP, February 2025).

They never needed it

The last mode skips the button entirely. In June 2025, Aim Security disclosed EchoLeak in Microsoft 365 Copilot, which Fortune described as “the first known “zero-click” attack on an AI agent”: an attacker triggers it “simply by sending an email to a user, with no phishing or malware needed” (Fortune, June 2025). Microsoft fixed it before customers were affected. In September, Noma Security’s ForcedLeak research showed a prompt planted through a company’s web lead form, the kind that feeds straight into Salesforce, waiting for an employee to ask the Agentforce agent about the lead, with data leaving through a domain that was still on the allowlist but had expired. “The domain that was purchased to find this vulnerability cost $5, but could be worth millions to your organization” (Noma Security, September 2025). Salesforce’s fix was an egress control, enforcing trusted destinations for the agent’s output. Sometimes the agent simply skips the step. In March, an internal agent at Meta posted a public reply to an employee’s question “without getting approval first,” another employee acted on its inaccurate advice, and the result was a SEV1, Meta’s second-highest severity level (The Verge, March 2026). Meta’s spokesperson said the agent took no action beyond answering.

Black Hat's own description of this talk: "no human-in-the-loop. In-and-out with a single prompt." The zero-click end of the spectrum. Watch on YouTube.

Line the five up and the damage in every case was set by something other than the button, namely what the agent could reach and where it could send what it found. Founders can set that part on a Tuesday afternoon.

The reach is the blast radius

The second half of Yash’s answer is the part I’d frame and hang in every startup’s engineering room. He runs the vendor due diligence everybody runs, from the penetration test report to the SOC 2, short for System and Organization Controls 2, the audit report most enterprise buyers ask for. Then he said the bigger challenge with AI tools sits somewhere else.

“The bigger issue here is what data do you have access to? How do you access it? And is it a two-way read-write or just a read?”

Yash Kosaraju, CISO of a16z, on episode 104

His reason fits in a line: “AI is only as good as the data it has access to.” I played it back to him on the show as scoping what the agent is allowed to touch inside your security boundary, and he said yes. Run the biggest AI-tool incidents of the last eighteen months through his question and you’ll find that almost every one of them turns on reach, and very few turn on a missing click.

Grid of recent AI tool incidents plotted by what the tool could reach, from read only to read and write to able to send data outside, and by how sensitive the data was, from public to production and money. The Salesloft Drift and Context.ai to Vercel breaches sit in read access to sensitive business data, PocketOS and the Amazon Kiro outage sit in write access to production, and Clinejection and ForcedLeak sit where tools could send data or code outside. The top right corner, sensitive data with the ability to send it outside, is where the lethal trifecta lives
Yash's question as a map. The incidents cluster where reach was wide and nobody had drawn the edge.

Read access was enough for the biggest breaches. In August 2025, attackers stole the OAuth tokens that the Salesloft Drift chatbot used to talk to its customers’ Salesforce instances. OAuth, short for Open Authorization, is the standard behind every “Sign in with Google” consent screen, and those tokens were standing permission to read. Google said the campaign potentially affected more than 700 organizations (CyberScoop, September 2025). Cloudflare, one of them, searched the stolen data afterward and found 104 Cloudflare application programming interface (API) tokens sitting in the text of support cases, and it described the export as taking “just over three minutes” (Cloudflare, September 2025). Cloudflare’s response is a model for owning a third-party failure: “We are responsible for the choice of tools we use in support of our business.” The chatbot’s job was chat, and its grant reached far beyond it.

The 2026 version is Vercel. In April, an attacker who had compromised Context.ai, a small AI office-suite startup, used a token from it to take over a Vercel employee’s Google Workspace account and pivot into Vercel’s environment (Vercel, April 2026). Context.ai’s explanation, as quoted by The Hacker News, is the most founder-relevant sentence of the year: “at least one Vercel employee signed up for the AI Office Suite using their Vercel enterprise account and granted ‘Allow All’ permissions” (The Hacker News, April 2026). A consent screen is an approve button too, clicked once by one person, and Jaime Blasco of Nudge Security drew the lesson in the same article: “OAuth is the new lateral movement.”

Write access was enough for the fastest ones. In April, an AI coding agent at PocketOS, an automotive software company, hit a credential mismatch in staging and went looking for a token, which it found in an unrelated file. The Register reported that “the token had been created for adding and removing custom domains through the Railway CLI but was scoped for any operation, including destructive ones.” The agent used it to delete the production volume, and because Railway stored volume backups in the same volume, the backups went with it (The Register, April 2026).

The PocketOS founder's own account, more than 5,000 likes and 860 points on Hacker News. Notice what he asks for: a confirmation on the one action that can't be undone, and a token scoped to its environment. View on X.

Two details make PocketOS the most useful incident of 2026 for founders. The first is that Railway restored the data within about an hour and patched the endpoint to perform “delayed deletes,” which turned an irreversible action into a reversible one. The second is that the founder’s list of missing controls starts with a confirmation step on a destructive call, and he’s right. A human, or a timer, belongs at the handful of actions that can’t be taken back, and nobody needs to approve the ten thousand ordinary actions around them. One of the top comments in the 1,032-comment Hacker News thread said it better than I can:

"It is fundamental to language modeling that every sequence of tokens is possible. Murphy's Law, restated, is that every failure mode which is not prevented by a strong engineering control will happen eventually."

maxbond, on Hacker News, April 2026. He later added that he meant it as the right mental model more than a literal claim.
The best short explainer of the PocketOS incident. It opens at 2:28, the chapter on blast radius, and walks through least privilege a minute later. Watch on YouTube.

Amazon has its own version. The Financial Times reported, and The Verge summarized, that a 13-hour outage of one Amazon Web Services (AWS) system in December followed its Kiro agent’s decision to “delete and recreate the environment,” and that “While Kiro normally requires sign-off from two humans to push changes, the bot had the permissions of its operator” (The Verge, February 2026). Amazon disputed the framing and called the cause “misconfigured access controls” (Amazon, February 2026). Both sides of that argument land in the same place: the two-person sign-off couldn’t help, because the agent could reach around it.

A Hacker News commenter put the principle in two sentences a year earlier, after Replit’s agent deleted a production database during a code freeze it had been told about in chat:

"If an AI agent has access, it has permission. Permissions aren't something you establish with an LLM by conversing with it."

pxc, on Hacker News, July 2025

(An LLM is a large language model, the engine inside every agent.) The write-side case I show engineering teams is Clinejection. From December 2025 to February 2026, Cline’s GitHub issue-triage bot ran for anyone who opened an issue, with tools including “Bash,Read,Write,Edit” and web access. A researcher showed that a crafted issue could take over the workflow, and another actor later used the path to obtain publication credentials and push an unauthorized Cline release to npm (Adnan Khan, February 2026). A bot whose job is reading issues and writing labels had been handed a shell.

Read-only helps, and it isn’t the whole answer. In a July 2025 demonstration, researchers showed a support ticket steering a developer’s database assistant, which held the Supabase role that bypasses row-level security, into copying private data from its Structured Query Language (SQL) database into a reply the customer could see. Their 2026 review of the demo adds the nuance founders need: “Read-only SQL prevents the database write used to return secrets through the ticket in this demonstration. It still allows reads that the database role can perform” (General Analysis). Drift and Context.ai were both reads. So I add one question to Yash’s three: where can it send what it reads?

The industry already has a name for that shape. Simon Willison called it the lethal trifecta in June 2025: an agent with access to private data, exposure to untrusted content, and the ability to communicate externally can be tricked into handing the first to an attacker, and “The only way to stay safe there is to avoid that lethal trifecta combination entirely” (Simon Willison, June 2025). Meta turned it into a design rule in October 2025, the Agents Rule of Two: until prompt injection can be reliably detected, an agent “must satisfy no more than two” of the three properties within a session, and one that needs all three shouldn’t run autonomously and at a minimum needs supervision (Meta, October 2025). That last clause is where a human approval still earns its keep, and I’ll come back to it. OWASP’s list of the top risks for applications built on large language models says it in database terms, almost word for word Yash’s question: an extension meant to read data shouldn’t connect “using an identity that not only has SELECT permissions, but also UPDATE, INSERT and DELETE permissions” (OWASP, Excessive Agency).

The money follows the reach. In IBM’s 2025 breach study, among the organizations that had a breach of an AI model or application, “97% report not having AI access controls in place.” In the 2026 edition, the most common causes of AI breaches were the systems around the model: “compromised APIs, applications, or plug-ins (27%) and cloud misconfigurations affecting AI workloads (27%)” (IBM, July 2025 and IBM, July 2026). IBM sponsors the study and it covers a few hundred breached organizations, but it points at the same lever as every incident above.

That lever is good news, because reach is the one variable a founder can change this week without waiting on a vendor. We wrote the long version of the permission ladder in Are You Actually Going to Give the Agent Write Access?, and the principle behind it in why agents need least privilege more than humans ever did. If a delete can’t be undone, the backup has to live outside the agent’s reach, which is the whole argument of our piece on immutable backups.

Say yes to a smaller version

Right after I played the scoping idea back to him, Yash went somewhere I didn’t expect from a CISO at a venture firm, straight to the word no.

“One thing we’re very deliberate about at a16z, especially within the cyber team, is not to say no when people request for something.”

Yash Kosaraju, CISO of a16z, on episode 104

His reasoning was two sentences long: “One, they’ll never come back to you. Two, they’ll find a way to use it anyway, right?” What his team does instead is work out what the person is trying to do, then “define access to the right set of data for that use case. If it’s a POC, that can have a subset of test data.” POC is a proof of concept, and I think that one sentence is the most practical security advice in the episode, because it turns a yes-or-no question into a question about size.

The survey data backs up both halves of his reasoning. In Software AG’s 2024 survey of 6,000 knowledge workers, 46 percent said they would refuse to give up their AI tools “even if their organization banned them completely” (Software AG, October 2024). In Okta’s 2026 survey, “52% of employees admit to using AI tools without approval,” and 57 percent said “the approval process is too slow or difficult” (Okta, May 2026). Both are vendor surveys, and the second number deserves a second look. A slow approval process is an approve button at the organizational scale, and it fails the same way the dialog does, with people routing around it.

The cost of routing around shows up in the breach data. IBM’s 2025 study found “One in five organizations reported a breach due to shadow AI,” and high levels of shadow AI added an average of $670,000 to breach costs (IBM, July 2025). Anthropic’s Jason Clinton described the mechanism in a July guide for CISOs: “Saying “no” to these requests produces shadow adoption, which has zero telemetry and generally no off switch” (Anthropic, July 2026). A banned tool that people use anyway is the widest reach of all, because nobody drew any boundary around it.

And here’s the win. Netskope’s 2026 telemetry across its customers shows what happens when companies offer a sanctioned version. “The percentage of AI users who use personal AI apps fell from 78% to 47%,” while “the percentage of people using organization-managed accounts has climbed from 25% to 62%” (Netskope, January 2026). The same report notes that blocking works for specific apps with no legitimate use, paired with a yes to the reputable ones. People take the managed yes when you offer it, and a managed yes comes with logs and an off switch.

Yash has been saying this for a while. On the Cloud Security Podcast in December 2025, while he was still at Sendbird, he described rolling out enterprise versions of the major AI tools because “if you don’t enable them by providing them AI tools, they are going to find tools either pay out of pocket or use free versions again to get the job done,” and he called the other half of it “the paved path.”

Yash in December 2025, on why blocking AI tools backfires. The player opens at 30:28. Quotes above are from the auto-generated captions. Watch on YouTube.

The best counterpoint comes from Rami McCarthy, who argued in late 2024 that “Security’s push to avoid being the ‘Department of No’ has overcorrected” (Rami McCarthy, December 2024). I take the warning seriously, and read in full, he and Yash want the same thing: “A “No” should never be a dead end.” Yash’s version is a yes with a smaller shape. Gartner called the shift back in 2023, when it predicted the CISO role would move “from being control owners to risk decision facilitators” (Gartner, March 2023).

The smaller version comes down to two habits a startup can start this month, one for the proof of concept and one for the consent screen. Start the proof of concept on test data, which is what federal control catalogs have asked for years: the National Institute of Standards and Technology (NIST) says in its Special Publication 800-53 that organizations “can minimize such risks by using test or dummy data during the design, development, and testing of systems” (NIST SP 800-53, SA-3(2)). Give the proof of concept its own credential, and make that credential expire when the proof of concept does. In June, attackers who breached Klue, an AI competitive-intelligence platform, “reportedly used a dormant but still active credential created by Klue for a prototype integration” to reach its customers’ Salesforce data (BleepingComputer, June 2026). The second habit fixes the Vercel pattern: move the consent screen from each employee to an admin policy. Google Workspace lets an admin restrict which apps can reach Gmail and Drive, and decide what happens when someone signs in to an app nobody has configured (Google Workspace Admin Help). Microsoft recommends allowing user consent “only for applications that have been published by a verified publisher” (Microsoft Learn). Either setting turns an approve button that every employee holds into a policy that one person writes.

We’ve written about this stance at length in Stop Saying No and how to govern shadow AI without blocking it. If you’d like to know what’s already running with your company’s data before you decide what to sanction, that’s the job of our Shadow AI Detection and Mitigation work, and the result is a short list of tools worth paving a road for.

Where the human goes next

So where does the person go once they stop clicking? The UK National Cyber Security Centre (NCSC) gave the cleanest vocabulary for it in August. “Human-in-the-loop: humans approve actions before they happen.” “Human-on-the-loop: humans monitor actions and can intervene if needed.” “Human-out-of-the-loop: AI acts autonomously without human review” (NCSC, August 2026). The terms are older than agents. Human Rights Watch used them in 2012 to describe autonomous weapons, with a warning that belongs on every agent dashboard: a system with a human on the loop can be “effectively out-of-the-loop” when “the supervision is so limited” (Human Rights Watch, November 2012).

That warning is the difference between moving the human and deleting them. Read the 2026 guidance from the labs and the governments side by side and it agrees on four places a person still belongs, each of which is a real job with real leverage.

Diagram of an AI agent's loop running inside a dashed boundary, with the human moved outside the loop to four stations. First, drawing the boundary of what the agent can read, write, spend and reach. Second, reviewing a risk-weighted sample of automated approvals. Third, approving the few irreversible actions such as deletes, payments and production changes, with context attached. Fourth, holding a tested off switch
The human on the loop has four jobs, and none of them is clicking approve on routine work.

The first job is drawing the boundary. Amazon’s security team explained why it built a policy layer for agents in words I’d give to any founder: “Human-in-the-loop provides a safety net for critical operations, and it will always have a role. But relying on it as the main control mechanism sacrifices autonomy and can lead to approval fatigue.” The alternative is “a safety envelope within which the agent can operate freely,” enforced “outside the agent and tools” because “the LLM’s plan is the thing you can’t trust” (AWS, May 2026). Google’s framework for agents gives the concrete version: a policy engine outside the model that blocks “any purchase action over $500” and asks for confirmation “for purchases between $100 and $500” (Google, May 2025). Jason Clinton’s July guide for CISOs has the two lines I quote most: “remove the delete verb from the agent’s world entirely,” and “the environment the agent loop runs in should never hold a credential worth stealing” (Anthropic, July 2026). Anthropic’s zero trust paper adds the test to run on every control you’re considering: “does this make the attack impossible, or just tedious?” An approval prompt is a tedious control, and an attacker with an agent has unlimited patience (Anthropic, May 2026).

The network is part of the boundary too. The NCSC’s advice is to “deny all inbound and outbound network traffic to the AI agent’s environment by default,” then allow only what the job needs, and to give every agent “their own unique identity” with “credentials with the shortest possible lifetime” (NCSC, August 2026). Ryan Dahl, who created Node.js and Deno, spent a talk this August showing what that looks like as code, with rules checked into git and a live demo of the policy blocking an agent’s attempt to drop a users table.

The boundary as code. The player opens at 12:52, where the policy stops an agent from dropping a production table. Watch on YouTube.

The best founder example of drawing the boundary after the fact came from Replit’s chief executive, the weekend his company’s agent deleted a customer’s production database.

The fix list from a founder under fire. Separate the environments so production is out of reach, and make the mistake reversible. The same post promised a planning mode that can't touch code at all. View on X.

The second job is reading a sample. Anthropic’s July post on securing its own development process describes the pattern plainly. It tiers the codebase by risk, lets automated reviews approve the low tiers, and keeps “strict human approval processes” for entire codebases. “Every approval is logged with the signals and reasoning behind it, and a risk-weighted sample is reviewed by humans.” New AI reviewers run in “Shadow mode” and post comments for human approval “until trust is earned,” and every agent action lands in the SIEM, the security information and event management system where the rest of the company’s alerts already go. Its summary line is the job description for your first security hire: “The security engineer’s job evolves from monitoring bugs to monitoring loops” (Anthropic, July 2026). It’s vendor-published and describes the vendor’s own house. Datadog’s CISO described a similar model on a16z’s own channel in August, with MCP servers scoped by role and given short-lived credentials, plus an AI judge his team built to catch malicious code.

A peer CISO's operating model, on a16z's channel. The player opens at 5:19, on role-based MCP servers and sandboxed agent credentials. Watch on YouTube.

The third job is approving the few actions that can’t be undone. Here the governments are firm, and I agree with them. The joint guidance from the Five Eyes cyber agencies in May says to “Insert human-in-the-loop review or approval checkpoints for actions where the cost of error is high, such as system resets, network egress or deletion of critical records,” and adds a sentence every agent builder should read twice: “Ensure decisions about when human approval is required are determined by system designers or operators, not delegated to the agentic AI system” (Australian Cyber Security Centre and partners, May 2026). Meta’s Rule of Two keeps human approval for the one configuration that needs all three risky properties. OWASP’s agentic top ten keeps it for “high-privilege or irreversible actions.”

What changes is the quality of the question. Phil Venables, who ran security at Goldman Sachs and Google Cloud, compared agents to high-frequency trading (HFT) systems and wrote that when an agent trips a guardrail and asks for approval, “the human must be given a ‘context bundle’ that clearly explains why the guardrail tripped. Blindly clicking ‘Approve’ is the modern equivalent of lifting an HFT circuit breaker” (Phil Venables, May 2026). Microsoft learned the same lesson from push fatigue, when it made approvals carry a number from the login screen. It’s also where Anthropic’s own data points: people reject 39 percent of plans and 3 percent of individual prompts. Ask the human about the plan and the irreversible step, with context attached, and they’ll still say no when it matters.

The fourth job is holding the off switch, and testing it. The NCSC says “you should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately,” and suggests a graduated start that any founder can copy: run agents in office hours “when more human oversight is available,” then expand to nights and weekends once the controls prove themselves. Venables adds the caveat that revoking an agent’s credentials “is a dangerous assumed ability” once it has spawned processes of its own, so the switch needs a drill before it needs a crisis.

The strongest objections to all this come from people I respect. Mitchell Hashimoto warned in May about companies “under heavy AI psychosis,” and the line has stuck with me.

More than 15,000 likes and 2,100 points on Hacker News for the best argument against removing people. My answer is that the four jobs above exist to keep the whole system comprehensible. View on X.

Gergely Orosz’s take on PocketOS was that “the blame sits with the dev who decided to delegate decision making to the AI agent, and then not review actions, just YOLO it” (Gergely Orosz on X, April 2026). He’s right that running an agent with no boundary and no review was the failure. Where I part ways is the fix. Reviewing every action is the control this whole post has measured, and it catches about one dangerous command in seven. Reviewing the boundary, a sample, and the handful of irreversible actions is work a person can do well for years.

The attackers have already made the same move. Anthropic’s report on the first AI-orchestrated espionage campaign it disrupted found that the operators used AI “to perform 80-90% of the campaign, with human intervention required only sporadically (perhaps 4-6 critical decision points per hacking campaign)” (Anthropic, November 2025). Its September 2026 threat report describes the same shape across every class of attacker it investigated: “Humans remained in the loop by setting the targets of attacks and reviewing exfiltration” (Anthropic, September 2026). The people on offense moved themselves on the loop a year ago, and defenders can use the same design to move just as fast.

Eric Olden of Strata Identity walked through the control plane for this on episode 81, and Graham Neray of Oso, an authorization company, explained why agents raise the risk for growing software companies on episode 89. If you want someone to draw the first boundary with you, that’s the core of our AI Access Control Audits, and for teams wiring agents into their own tools, our Secure MCP Program covers the servers and the scopes they hand out.

Your first security hire builds the loop

Every job in the last section is engineering work. Scoped credentials, sandboxes, network allowlists, logs that land somewhere useful, and a drill for the off switch all live in code and infrastructure, and nobody ships them by writing a policy document. That’s why one of the first things Yash said on the show, before we ever got to agents, turns out to be the most important piece of advice for this problem.

“When you look at the first security hire, it’s mostly a staff level, senior staff level engineer that can be independent, but isn’t your true head of security person who’s building the team.”

Yash Kosaraju, CISO of a16z, on episode 104

That person is “more in the weeds, in the code, in your infrastructure, hardening things, building things out,” he said, and the timing is a risk decision driven by what data you hold and what’s at stake. The leader who builds a team “comes much later.” Where the engineer sits matters as well: he suggested the infrastructure team or another engineering team, “a place where they have enough context and access to just execute.”

The practitioner consensus lines up with him. Thomas Ptacek, who wrote Latacora’s SOC 2 starter list, put the norm on Hacker News in 2022: “the startup industry norm is to make a first security hire somewhere between engineer #20 and #40,” and “your first security hire needs to be a hands-on-keyboard person, and CISOs are not that” (Thomas Ptacek on Hacker News). Rami McCarthy’s rule is my favorite framing of when: “You should hire your first security person when security is an unavoidable distraction from scaling your business” (Rami McCarthy, November 2024). Frank Wang calls the security-minded software engineer the first-hire persona he’s “most bullish on” (Frank Wang, September 2025). And the classic counterpoint, from Ryan McGeehan in 2017, is that “it can be harmful to introduce dedicated security employees too early,” with outside expertise and engineers taking on security projects in the meantime (Ryan McGeehan, January 2017).

a16z’s own hiring shows the model at the top end. The staff security engineer role a16z opened in June describes “a small team” multiplied by automation, calls it “an AI-native role,” and lists “agent identity boundaries” among the baseline skills (a16z careers). The companion listing, for a staff security technical program manager, asks for someone to “keep vendor security reviews fast and consistent” (a16z careers), which tells you how a sophisticated buyer thinks about the reviews founders dread.

The YSecurity team at dinner during Black Hat USA 2026 in Las Vegas
The YSecurity team at dinner during Black Hat USA 2026, the week Yash was in Las Vegas for the Black Hat CISO Summit.

Then there’s the buyer’s call, the moment when a big prospect asks to speak with your security leader. Yash was honest that it’s a real signal: bigger clients “would definitely ask for one, either a call with your security leader or a dedicated point of contact that they can reach out to for emergencies.” He was just as clear that “that’s probably not when you’re hiring for your head of security,” and that when the multi-million-dollar deals arrive you can decide whether to grow your head of security into the role “or maybe get a field CISO to answer some of those questions.” Kayne McGladrey, a field CISO himself, joined us on episode 56, and Andrew Gontarczyk of Pure Storage talked about timing the first team on episode 2.

Here’s where the approve button meets the Compliance Beast. SOC 2 change management has always rested on a human approving each change. Latacora’s classic guide for startups says to “Require reviews for PRs that merge to production,” a PR being a pull request, and also that “SOC2 is about documentation, not reality” (Latacora, March 2020). When agents write and review most of the code, the auditor’s question becomes whether that human review still means anything, which Paddy took on in detail in Agentic Code Review and Your SOC 2 Type 2 Audit. An opinion piece in Corporate Compliance Insights this month named the trap in one line: “A person being present in the process does not prove that a person was in control of it.” Its fix is the kind of evidence an auditor can use: “seed known wrong outputs into live queues, as security teams do with phishing simulations, and make the catch rate the primary oversight metric” (Corporate Compliance Insights, September 2026).

The questionnaires are catching up too. The 2026 edition of the Standardized Information Gathering questionnaire, the SIG, now references ISO 42001, the International Organization for Standardization’s standard for AI management systems (Shared Assessments, July 2025), and a new certification called AIUC-1 is positioning itself as a “SOC 2 for AI agents,” which Lenny Zeltser reviewed with the right skepticism about scope (Lenny Zeltser, April 2026). None of them asks as plainly as Yash does what a tool can reach and whether it can write back. The founders who can answer that anyway, with a list of every agent and its reach plus a log someone actually samples, are the ones who move through the AI section of a security review quickly.

The first hire delivers a second win here. A hands-on engineer who builds the loop also builds the evidence, and evidence is what gets you through the buyer’s review. If you’re weighing that hire, our staffing guide lays out the options and the in-house math, and the short quiz points you at the one that fits your stage. We went through a SOC 2 Type 2 with Augment Code and with Robust Intelligence, and we help founders through the buyer’s call itself with sales support. If you’re earlier than that, start with What Even Is a SOC 2? and the founder’s guide to compliance on episode 37.

Twelve to eighteen months of pain

A few weeks before we recorded, I was at a small CISO dinner in San Francisco, and I described the mood to Yash on the show. The room had settled on assume breach. Frontier models can read your code base and the third-party libraries it depends on, and I told him the consensus was that “anything exposed to the internet basically is going to break, if it hasn’t already.”

His first words were “They’re not wrong, by the way.” AI is good at working through “all the different CVEs and exploits out there at a very fast pace,” he said, and “the cost per exploit, or cost per breach, that as an attacker you have to then sort of be aware of, with AI would be going down. So the economic balance is going out of whack.” Then he said the part that made me more hopeful than the dinner had.

“My hope is that we have this next 12 to 18 months of pain where, again, we’re playing catch-up, but much more rapidly.”

Yash Kosaraju, CISO of a16z, on episode 104

After that, he expects whole classes of vulnerabilities and tools to become unnecessary, and “everybody gets there together, screaming and kicking, but eventually you get to a point where we’re much better off than today.”

The evidence for the pain part is strong. In May, Anthropic reported that its Project Glasswing partners had used an unreleased model to find more than ten thousand high- or critical-severity vulnerabilities in about a month, and named the new bottleneck: “Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it’s limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI.” At the time, “75 of the 530 high- or critical-severity bugs we’ve reported have now been patched” (Anthropic, May 2026). Google’s threat intelligence group described a North Korea-linked group “sending thousands of repetitive prompts that recursively analyze different CVEs and validate PoC exploits,” which is Yash’s sentence, nearly word for word, in a threat intelligence report (Google Threat Intelligence Group, May 2026). (A PoC, here, is a proof-of-concept exploit.) And Verizon’s 2026 breach report found that vulnerability exploitation, at 31 percent, overtook stolen credentials, at 13 percent, as the top way attackers get in, for the first time in the report’s history (TechRepublic on Verizon’s Data Breach Investigations Report, May 2026).

Two-panel chart of the attack economics Yash described. Left, the cost of turning a vulnerability into a working exploit or fix: 8.80 dollars per one-day exploit with GPT-4 in 2024, about 152 dollars per find-and-patch task in DARPA's AI Cyber Challenge in 2025, 2.77 dollars per reproduced CVE with CVE-Genie in 2025, and zero marginal cost per infection for AI worms running on stolen compute in 2026. Right, Mandiant's average time from disclosure to exploitation falling from 63 days in 2018 to 2019, to 44, 32 and 5 days in 2023, and an estimated minus 7 days in its 2026 report
The balance Yash says is going out of whack, in two lines. Exploits got cheap, and on average they now arrive before the patch.

The chart is built from research papers, the Defense Advanced Research Projects Agency’s AI Cyber Challenge and Mandiant’s reports (Kang lab, 2024, DARPA, August 2025, CVE-Genie, 2025, Guan and colleagues, June 2026, Mandiant, October 2024 and M-Trends 2026). The minus seven days is Mandiant’s estimate of the mean, and it means “exploitation is routinely occurring before a patch is even released.” Anthropic’s Jason Clinton put a date on the wave in August.

Notice the remedy he reaches for. When patching can't keep up, you limit what can be reached, which is this post's argument applied to infrastructure. View on X.

Here’s the counterpoint that keeps this honest. VulnCheck, a vulnerability intelligence company, counted 1,061 vulnerabilities attributed to AI-assisted discovery by mid-2026, and “Of those, 14, or 1.3%, have been confirmed as exploited in the wild” (VulnCheck, July 2026). Jerry Gamblin, who tracks CVE volume, reported that the first half of 2026 produced 35,364 CVEs, up 49.5 percent, while “only 85 of them (0.24%) have made CISA’s KEV list so far,” the Known Exploited Vulnerabilities catalog kept by the Cybersecurity and Infrastructure Security Agency, and concluded that “the hard problem is signal-to-noise, not patch volume” (Jerry Gamblin, July 2026). The finding wave is here and the exploitation wave is still mostly a forecast. It argues for Yash’s own triage rule on the episode, which I love for its brevity: find the issue, then figure out if it’s “vulnerable, exploitable, reachable.” Every boundary you’ve already drawn makes more of the backlog fail that third test.

The labs learned the same lesson about their own agents this year. In July, Anthropic reported that after reviewing 141,006 cybersecurity evaluation runs, it found three incidents in which a model reached the real internet through a misconfigured environment and compromised real organizations “using basic techniques, such as exploiting weak passwords and unauthenticated endpoints” (Anthropic, July 2026). The model had been told it had no internet access, and the network said otherwise. OpenAI disclosed that its models, during internal evaluations, “circumvented controls designed to isolate them from the internet” and compromised parts of Hugging Face’s systems (OpenAI, August 2026). Both labs’ fixes were about isolation and validated network paths.

One of the most-watched security talks of the summer, with more than a million views. An evaluation agent, a sandbox that leaked, and what the lab changed afterward. Watch on YouTube.

Yash’s longer-term hope is that coding agents will “write secure code by default, and you don’t need some of these features.” The direction is right and the date is early. Sundar Pichai said in April that “75% of all new code at Google is now AI-generated and approved by engineers” (Google, April 2026), while Veracode’s July benchmark put the average security pass rate for generated code at 56 percent, flat for two years (Veracode, July 2026). That’s a vendor benchmark of what models write on their own, and the improvement is coming from the review loop around them. Yash described the version he wants: an agent finds the issue, checks reachability, opens a pull request, “and then the engineering agent reviews it and patches it.” He expects the remaining human attention to go to business logic flaws, “which AI isn’t as good at.” Google’s threat intelligence group partly disagrees there, writing that frontier models “excel at identifying these types of high-level flaws” while they “struggle to navigate complex enterprise authorization logic” (Google Threat Intelligence Group, May 2026), so I’d plan for both.

Then Yash widened the lens in a way founders should hear. “We are here in the Silicon Valley, and that is a bubble,” he said, and the companies he worries about are the ones that “maybe don’t even have the financial resources to spend on tokens to build their defenses.” The data backs him. The US Census Bureau found that “Less than 20% of firms with four or fewer employees reported using AI” (US Census Bureau, May 2026), and Microsoft’s measure of generative AI use ranks the United States 21st, at 33 percent of the working-age population (MediaPost on Microsoft’s AI diffusion data, September 2026). Wendy Nather coined the term security poverty line, and the 2013 RSA Conference deck she presented with Andy Ellis lists “No logs” and “Default settings” among the signs of life below it, a description that also fits a lot of seed-stage startups. Your customers and suppliers may sit below that line, and you are their third party: Verizon found third-party involvement in 48 percent of breaches, up from 30 percent the year before (TechRepublic, May 2026). His other line from that stretch is the reason to act anyway: “The cost of not adopting AI today is much more than the cost of not adopting any other previous technologies.”

One more prediction from the episode is already happening. “The security space is saturated with vendors,” Yash said, and “at some point” it needs “a consolidation to a very few ones that do more than just one feature.” IT-Harvest counted more than 4,000 active security vendors at the end of 2024 (Richard Stiennon, December 2024). Since then Palo Alto Networks closed its roughly $25 billion CyberArk deal to secure “every identity across the enterprise - human, machine, and agentic” (Palo Alto Networks, February 2026), and Google closed its purchase of Wiz, the firm behind GhostApproval, for the $32 billion it agreed to pay in 2025 (Google, March 2025, Google Cloud, March 2026). One of those deals was pitched explicitly as identity for agents, which tells you where the biggest platforms think the control lives.

For a startup, the practical version of catching up is to triage the coming patch flood by reachability and to build review into the coding loop. Our AI Vulnerability Remediation and secure AI development work does exactly that, and the background is in how to stop AI agents from shipping vulnerable code and what autonomous attacks look like. Neatsun Ziv of Ox Security made the case against asking developers to fix everything on episode 91, which pairs with our piece on why shift-left keeps failing, and Kabir Mathur of Leen covered the more-tools trap on episode 63.

A price tag on the send button

The last stretch of the episode was about money, and it connects to the approve button more than it first appears to. Yash’s hope, “like others, is that the token cost would come down,” but he sees a different problem first.

“We’re not good at estimating what each task we’re giving an AI agent would take.”

Yash Kosaraju, CISO of a16z, on episode 104

People want “the latest and greatest model,” he said, and “the fast mode and not the slow mode.” His illustration is the one founders repeat back to me: “there is one way of finishing the same task in 10 minutes that might cost you $250, or it could take you 15 minutes and cost you only $50.” Keep his hedges attached, because that spread comes from what he’s watching inside his own organization, and the price sheets are less dramatic on their own. Anthropic’s fast mode offers “up to 2.5x higher output tokens per second” at a premium, with Claude Opus 5.5 going from $20 to $40 per million output tokens (Anthropic docs). By OpenAI’s own price tables, its fast mode doubles GPT-6 Sol’s output price from $10 to $20 per million tokens (OpenAI fast mode, OpenAI pricing). Model choice multiplies the rest, and Anthropic’s own pricing page notes that its newer tokenizer “produces approximately 30% more tokens for the same text.” Cursor admitted the underlying problem in 2025: “the hardest requests cost an order of magnitude more than simple ones” (Cursor, July 2025).

On the show, I wished out loud for a little dollar sign that told you what a prompt would cost before you clicked it. Yash’s reply was one sentence: “I don’t think the frontier models have the incentive to do that for us.” So companies set their own incentives, and this year some set the wrong ones. Business Insider reported on engineers racing up internal token leaderboards, and on Linear’s Cristina Cordova replying that “Ranking engineers by token spend is like me ranking my marketing team by who spent the most money” (Business Insider, April 2026). By May, Fortune declared “Tokenmaxxing is over,” reporting that Uber “had burned through its entire 2026 “token budget” in just the first four months of the year” (Fortune, May 2026).

The loudest pro-spend voice of the year, from the company that sells the chips. Yash's mandate at a16z runs the other way. View on X.

a16z took the other road. “We gave people all the tools to experiment, but the mandate was not token maxing or go spend money,” Yash said. “The mandate was, here’s a bunch of next-generation tools. Try them out, use them, see where they can add value.” He’d said it more bluntly on LinkedIn in June, hiring for his team: “using AI BUT unfortunately no token leaderboards” (Yash Kosaraju on LinkedIn, June 2026).

Here’s why this belongs in a post about the approve button. A spending ceiling is a boundary too, and like a scoped credential, it’s one an agent can’t talk its way around. A price shown before the click is also the rare prompt that carries real information, the “context bundle” Venables asked for, which is why people would read it. Sasha went deep on dollar ceilings in Too Fancy, the $50,000-an-hour agent, including a calculator for yours.

Score what your agents can reach

Everything above comes down to an inventory you can do this week. List every agent and AI tool that touches your company, with what it can read, whether it can write or send, and whether content from outsiders reaches it. The scorecard below does the arithmetic. It ignores your approval settings on purpose, because the approval is the part this post has shown you can’t count on, and it rewards the two boundaries that hold: a sandbox with a network allowlist, and a credential scoped to the job.

Agent reach scorecard

What your agents can reach, and what to scope first.

Three common agents are filled in as examples. Change them to match yours, or add up to three more. "Outsiders" means anyone who can write something your agent reads, like web pages, inbound email, support tickets or public repositories.

Agent 1
Agent 2
Agent 3
…widest reach, out of 100
…agents breaking Meta's Rule of Two
…high-reach agents with nothing beneath the approval
…widest reach after the plan below

How the score works: each agent's reach is the sensitivity of the most sensitive data it touches, times what it can do with it (read, write inside the company, or send outside), times a bump if outsiders' content reaches it, times a discount for real boundaries. A score of 100 is an agent that can move money or change production, can send data out, reads outsiders' content, and has nothing beneath the approval. How approvals work today doesn't change the score, because in Anthropic's 1,053-person study the human approver caught 13.6 percent of dangerous commands. It does shape the plan.

A person reads this, turns your list into a scoped plan with owners and dates, and replies. No sequence, no vendor pitch.

If you don’t have time for the scorecard, the plan it generates follows the same shape for almost everyone, and it fits in three columns.

A three-column plan for moving the human on the loop. Day one: list every agent and AI tool with what it can read, write and send, move employee consent for new AI apps to an admin policy, and remove delete and send verbs that no job needs. Month one: give each agent its own identity and a credential scoped to its job, run coding agents in a sandbox with a network allowlist, keep agent configuration out of the agent's write reach, and send agent logs to wherever your alerts already go. Quarter one: sample a risk-weighted share of automated approvals, seed known bad actions to measure the catch rate, keep human approval for the few irreversible actions with context attached, and drill the off switch
The plan in three columns. Every item on it is something a small team can own, and none of them depends on a person reading the three-hundredth prompt.

The first column takes an afternoon and the second is a sprint for the engineer Yash described. The third is the loop itself, and once it’s running, it doubles as the evidence your next enterprise buyer will ask to see.

The rule worth stealing: scope the reach, then let it run

Carlin’s planet was fine without us, and the loop will be fine without a person clicking inside it. What the loop needs is a person standing outside it, drawing the edge and keeping a hand on the few actions that can’t be undone.

The rule worth stealing: scope the reach, then let it run. Decide what each agent can read, write, send and spend before it starts, and keep the human on the plan and the irreversible step while the agent does the rest at machine speed. You get to run more agents, faster, because you can afford for any one of them to be wrong.

If you’d like help drawing the first boundary, our AI Access Control Audits start with exactly the inventory the scorecard asks for, and our AI Product Red Teaming tests whether the boundary holds when someone pushes on it. When a customer’s security team asks what your AI can reach, our sales support team helps you answer with evidence. Book your free 15-minute security strategy call, and if you’re a startup, your first 8 hours are free.

The last word goes to a builder who posted this in September, because he’s already seeing the next version of the same story.

The arc of this post in one builder's words, now reaching the merge button. The same rule applies there too. View on X.

Listen to Yash Kosaraju on episode 104 and decide for yourself where your people should be standing.

Human in the loop and AI agent approvals, frequently asked questions

What does human in the loop mean for AI agents?
Human in the loop means a person approves each action before it happens: the agent proposes a command, a file edit or an email and waits for a click. The UK National Cyber Security Centre's August 2026 guidance on agentic AI uses three terms. Human-in-the-loop: humans approve actions before they happen. Human-on-the-loop: humans monitor actions and can intervene if needed. Human-out-of-the-loop: the AI acts autonomously without human review. Most coding agents shipped with the first model as their default, and in 2026 the major vendors began moving routine approvals to classifiers and sandboxes, which moves people toward the second.
Why do approval prompts fail as a security control for AI agents?
Because people stop reading them, and the evidence is old and consistent. Microsoft reported in 2008 that consumer administrators approved 89 percent of User Account Control prompts in Windows Vista, and only 13 percent of lab participants could say why a dialog appeared. Brain-imaging research at Brigham Young University found a dramatic drop in visual processing after only the second exposure to a warning. In 2026 Anthropic reported that users approve 93 to 97 percent of Claude Code permission prompts, and in its controlled study with 1,053 paid testers, humans caught 13.6 percent of dangerous commands swapped into the flow, falling to about 5 percent after 50 or more prompts. Attackers can also switch prompts off, make them misleading, or change a tool after it was approved.
What is human on the loop, and what does the human do there?
Human on the loop means the agent acts without asking at each step, inside limits people set in advance, while a person watches and can intervene. In practice the human's work moves to four places: setting the boundary of what the agent can read, write, spend and reach; reviewing a risk-weighted sample of automated decisions; approving the few irreversible actions such as deletes, payments and production changes; and holding an off switch that has been tested. Human Rights Watch warned in 2012 that on-the-loop supervision that is too limited is effectively out of the loop, so the watching has to be real: logs, alerts on exceptions, and someone who reads them.
Can an AI reviewer or auto mode replace the approve button?
It can replace the click. The vendors themselves say it does not replace the boundary. Anthropic's auto mode, OpenAI's Codex Auto-review and Cursor's Auto-review route routine approvals to a model. Anthropic reported that its classifier missed 17 percent of real overeager actions in March 2026, OpenAI wrote that it does not expect this class of system to become a source of deterministic guarantees, and Cursor's documentation says Auto-review is not a security boundary. The researcher Johann Rehberger reported 60 to 80 percent attack success against Claude Code's auto mode on a small sample in August 2026. Treat the reviewer as a better approver that runs inside a sandbox, with limited network egress and scoped credentials.
What is GhostApproval?
GhostApproval is a vulnerability pattern that Wiz Research disclosed on July 8, 2026 across six AI coding assistants: Amazon Q Developer, Anthropic Claude Code, Augment, Cursor, Google Antigravity and Windsurf. A malicious repository ships a file that is really a symbolic link to something outside the project, such as ~/.ssh/authorized_keys or ~/.zshrc. When the agent edits it, the approval dialog shows the harmless local file name while the write lands on the sensitive target, and in some tools the write happened before any dialog appeared. Wiz's recommended fixes all sit below the dialog: resolve symlinks before displaying prompts, warn when the resolved path is outside the workspace, and never write to disk before explicit authorization.
How do attackers get around an AI agent's approval prompt?
Five ways show up in the 2025 and 2026 record. They switch it off: the Nx s1ngularity malware in August 2025 ran victims' own AI command-line tools with --dangerously-skip-permissions, --yolo and --trust-all-tools. They get the agent to switch it off: CVE-2025-53773 let a prompt injection write "chat.tools.autoApprove": true into a project's settings and put GitHub Copilot into auto-approve. They make the dialog show one thing while the agent does another, as in GhostApproval and SymJack. They change a tool after it was approved, as with Cursor's MCPoison and the postmark-mcp package. And they skip it entirely with zero-click attacks such as EchoLeak, where no approval ever appears.
What is the lethal trifecta, and what is the Agents Rule of Two?
Simon Willison named the lethal trifecta in June 2025: an agent that combines access to private data, exposure to untrusted content and the ability to communicate externally can be tricked into sending that data to an attacker. Meta's Agents Rule of Two, published October 31, 2025, turns it into a design rule. Within a session an agent should satisfy no more than two of three properties: processing untrustworthy inputs, having access to sensitive systems or private data, and being able to change state or communicate externally. An agent that needs all three should not operate autonomously and at a minimum requires supervision, which is where a human approval still earns its place.
What should a security review of an AI tool actually ask?
Yash Kosaraju, CISO of a16z, runs the vendor due diligence everybody runs and says the bigger issue with an AI tool is what data it has access to, how it accesses it, and whether the connection is two-way read-write or just a read. Add a question about where it can send data, because several of the largest incidents of 2025 and 2026, including Salesloft Drift, Gainsight, Klue and the Context.ai grant that led into Vercel, were reads through standing OAuth tokens. Ask which scopes the tool requests, whether a proof of concept can run on test data, how its credentials are revoked, and what it logs.
Should a security team ever say no to an AI tool?
Rarely as a dead end. Yash Kosaraju's team at a16z is deliberate about not saying no, because a no leads to two things: people never come back, and they find a way to use the tool anyway. The data agrees. In Software AG's 2024 survey of 6,000 knowledge workers, 46 percent said they would refuse to give up their AI tools even if their organization banned them, and Netskope's 2026 telemetry showed personal AI app use falling from 78 to 47 percent of AI users as organization-managed accounts rose from 25 to 62 percent. The better answer is a scoped yes: a proof of concept on a subset of test data, read-only first, with writes and broader data added as the use case earns them.
When should a startup make its first security hire?
Earlier than the head of security, and it is a different person. Yash Kosaraju's rule is that the first security hire is a staff or senior staff engineer who can work independently in the code and the infrastructure, hardening things and building things out, and that the timing is a risk decision driven by the data you hold and what's at stake. The head of security who builds a team comes much later. Thomas Ptacek put the startup norm for a first security hire somewhere between engineer number 20 and number 40, and said that person needs to be hands-on-keyboard.
How do you limit what an AI agent can reach?
List every agent and AI tool, what data each can read, whether it can write or send, and what content written by outsiders it takes in. Then give each agent its own identity and credentials scoped to its job, short-lived where the platform allows; run it in a sandbox with a network allowlist; remove verbs the job does not need, such as delete; turn writes into drafts or pull requests where you can; keep the agent's own configuration out of its write reach; and log every action to wherever your security alerts already go. The UK National Cyber Security Centre recommends denying all inbound and outbound network traffic to an agent's environment by default and always being able to pull the plug.
Does the EU AI Act require a human in the loop for AI agents?
Only for high-risk AI systems. Article 14 requires them to be designed so that people can effectively oversee them, names automation bias as a tendency the people overseeing must stay aware of, and requires a way to interrupt the system through a stop button or a similar procedure. It does not require approving every action. After the Digital Omnibus on AI, those obligations apply from December 2, 2027 for stand-alone high-risk systems and August 2, 2028 for high-risk systems embedded in products. Most coding and operations agents are not high-risk systems under the Act, but Article 14 is a useful description of what enterprise buyers will mean by oversight.
Written by the team behind The Security Podcast of Silicon Valley

Put it into practice.