Live webinar The Wrong Security Hire Burns Your B2B GTM Pipeline. A fireside chat for founders · Oct 6, 9:30am PTThe wrong security hire · Oct 6 Save your seat

Blog

13 min read Updated September 6, 2026

Why Shift-Left Security Keeps Failing (and What Actually Works)

Shift-left sounds right: find bugs earlier, fix them cheaper. It keeps failing because security teams push every scanner finding to developers who have no way to tell which ones matter. Here is what reachability and business context change, and what a startup should do instead.

Shift-left security has been the default advice in application security for over a decade: find bugs earlier, fix them cheaper, ship safer code. The problem is that it keeps failing. Security teams push findings to developers who lack the context to prioritize them, and the backlog grows instead of shrinking.

Neatsun Ziv, co-founder and CEO of OX Security and a former VP of Cyber Security at Check Point, put it bluntly on episode 91 of The Security Podcast of Silicon Valley: “Shift left as a concept failed us for years. We’ve been trying to make it work, but it’s just that everybody’s got their priorities.” The failure isn’t a tooling problem. It’s a context problem. This post is about what that means in practice: why shift-left as most teams run it (every scanner finding, forwarded to whoever wrote the code) was always going to stall, what reachability and business context change, and what a startup with twelve engineers and no security hire should do instead.

What is shift-left security? The textbook definition

Shift-left security means moving security work earlier in the software development life cycle: threat modeling while the design is still a whiteboard, static analysis and dependency scanning on every commit, secret detection before a push lands, infrastructure-as-code checks before anything deploys. The name comes from the left-to-right timeline of a project plan, and the economic argument is one every engineer already believes. A bug found while the file is still open in the editor costs minutes; the same bug in production costs an incident and a week.

The formal version lives in NIST’s Secure Software Development Framework (SP 800-218, the SSDF), which organizes secure development into four groups: Prepare the Organization, Protect the Software, Produce Well-Secured Software, and Respond to Vulnerabilities. Two practices in the third group, PW.7 and PW.8, are shift-left in a sentence: review and test code so that vulnerabilities “can be corrected before the software is released to prevent exploitation.” Nothing in that is wrong, and nothing here argues against finding bugs early.

Read one group further, though, and the framework says something the tooling era skipped. Practice RV.2 asks that vulnerabilities be “remediated in accordance with risk to reduce the window of opportunity for attackers.” In accordance with risk, rather than in order of discovery or by count.

The textbook version, in six minutes: IBM Technology's explainer of what shift-left security is, roughly 9,000 views. Watch it for the intent; the rest of this post is about what happened to it in practice. Watch on YouTube.

Why does shift-left security keep failing?

The textbook definition says nothing about the workload, and the workload is where shift-left as practiced went wrong. It became a delivery mechanism. Security teams bought SAST, software composition analysis, secret scanning and container scanning, wired them into CI, and pointed the output at the people who wrote the code. The tools did exactly what they were built to do: they found everything.

Then the handoff broke. Consider a scanner flagging a vulnerable dependency. The developer has no way to answer the questions that matter: Is this deployed to production? Does the affected path handle externally exposed APIs? Which databases does this service connect to? As Ziv explains, “Most of us are coming blind to those questions.” Without that context, developers either fix everything, wasting time on low-risk issues, or ignore the backlog, leaving real risks unaddressed. Neither is what shift-left promised, and the result is developer security fatigue, which costs the program its credibility with the people it depends on.

Neatsun ran that experiment at scale before he founded OX. Earlier in his career he managed 600 developers and did the textbook thing: handed them the prioritized list and asked them to fix the defects. They agreed to fix two. The reason was not laziness. Most of the findings lived in internal labs and containers that never touched the internet, nobody could show the team which ones did, and so the team went back to the roadmap.

The arithmetic explains the rest. The Cyentia Institute, analyzing Kenna Security data from about 300 organizations, found that a typical organization “will have the capacity to remediate about one out of every 10 vulnerabilities within a given month,” and that the ratio holds whether the company is tiny or enormous. If the queue grows faster than a tenth of it can be cleared, sorting it by CVSS score does not rescue it. Developers learn, correctly, that the security feed is mostly noise, and they stop reading it. Every real finding that arrives afterward inherits that reputation.

Cloud security educator Mark Nunnikhoven (marknca) on the one mistake that stalls a DevSecOps shift-left strategy, in under five minutes, roughly 5,000 views. Watch on YouTube.

How many scanner findings actually matter? The alert funnel

Neatsun’s number from the episode is the one to remember: “95% of security fixes have zero risk impact.” Take a raw scanner queue and ask, item by item, whether fixing it would change the odds of a breach; that is the answer he gets. Public data lands in the same neighborhood from three directions.

The application security alert funnel: everything the scanners flag narrows to findings running in production, then to findings that are reachable and exposed, then to findings where exploitation is likely (CISA KEV, EPSS, public exploit code), then to findings that touch customer data, payments or authentication; only the last two layers are fixed now as pull requests, the rest is documented as accepted risk
The alert funnel. Filled bars are this sprint's work. Everything above them is context you have to add before a developer can act. Widths are illustrative.

Exploitation is rare. FIRST, the organization behind the Exploit Prediction Scoring System, publishes the numbers behind its model. As of October 2023, NVD held 139,473 CVEs with CVSS 3.x scores; over the following 30 days, 3,852 of them, about 2.7%, showed exploitation activity anywhere in the wild (FIRST, EPSS model). A “fix everything CVSS 7 and above” policy queues 57.4% of all published CVEs to catch those, and 3.96% of what it queues is ever exploited. Four hits in a hundred is the efficiency of the default shift-left backlog.

Context removes most “criticals.” Datadog’s State of DevSecOps 2026 report, built on telemetry from its own customers (vendor research, but a very large real-world sample), found that “only 18% of critical dependency vulnerabilities stay critical after adjusting the severity score” for runtime context: in production, exposed to the internet, exploit available. Same figure as the year before. Their headline for that section is the whole argument in six words: “Most vulnerabilities should not page a human.”

The list of what is actually being exploited is public. CISA’s Known Exploited Vulnerabilities catalog is, in the agency’s words, “the authoritative source of vulnerabilities that have been exploited in the wild,” and federal agencies must remediate anything on it within set deadlines under Binding Operational Directive 22-01. A startup has no such mandate, which is exactly why the catalog is worth adopting voluntarily: a free, maintained, top-of-the-funnel list.

What survives the funnel is this sprint’s work. The rest gets written down as accepted risk and re-checked when its context changes, which is a different thing from a ticket nobody plans to close.

“Security isn’t about canceling risk, it’s about managing it.”

Neatsun Ziv, co-founder and CEO of OX Security, on episode 91

What is reachability analysis, and what does business context add?

Reachability analysis answers a narrow question well: can this vulnerability actually be triggered in this application? For a vulnerable dependency, that means checking whether your code ever calls the affected function (static reachability, via the call graph) and whether the library is even loaded at runtime. For a static-analysis finding, it means checking whether attacker-controlled input can flow to the flawed line. A vulnerable function that nothing calls is a hygiene item; a vulnerable function on a request path is a finding. Exposure (internet-facing or behind authentication?) and exploitation data (KEV and EPSS, both free) are the next two filters.

Business context is the part only your company can supply, and the part Neatsun kept returning to on the episode: which services hold customer data, which one processes payments, which one issues session tokens. A medium-severity flaw in the service that mints tokens outranks a critical in an internal admin tool three people use over VPN. No scanner knows that. Someone on your team does, and writing it down once is the highest-leverage security work a startup can do.

Without contextWith context
CVE-2026-1234 in package XCVE-2026-1234 in package X, deployed to prod, handles external API traffic, connected to the customer database
Priority: unknownPriority: critical, fix this sprint

Security teams stop being ticket generators and start being context providers. Developers stop drowning in unranked alerts and start fixing the three things that would actually cause a breach. It mirrors what’s happening across security broadly: just as microsegmentation replaced blanket perimeter rules with context-aware access, context-first AppSec replaces blanket alerts with threat-relevant, prioritized findings.

When the funnel leaves you with a finding you still cannot classify, the cheapest way to settle it is to try to exploit it. That is what a good penetration test does: it reports which items an attacker could chain into something that matters, with the context to fix them, instead of a PDF of everything a scanner could flag. Autonomous pentesting agents are starting to do the chaining at machine speed, so exploitability becomes a question you can ask continuously rather than once a year.

What is prevention-first security, and how do AI coding agents change it?

“Shift-left is dead. Prevention-first is what comes next.”

Neatsun Ziv, co-founder and CEO of OX Security, on episode 91

Prevention-first is the idea that the best finding is the one that never gets created. You decide the constraints before the code exists: an approved package registry with pinned versions, a secrets vault so there is nowhere else to put a key, sanitization and rate-limit libraries that are the easy default, and written security requirements that live where engineers will actually read them. The SSDF’s very first practice, PO.1, says the same thing in standards prose: security requirements should be “known at all times so that they can be taken into account throughout the SDLC.” Scanners still run in a prevention-first program; they become a check that the defaults held, rather than the primary control.

AI coding agents make this practical rather than aspirational, because an agent reads its context every single time. Ziv’s approach:

“If we can load up front the context saying, you are going to touch right now… an API that is externally exposed, these are the databases that this app is connected to… we want you to have those security restrictions and this rate limit and this sanitization.”

Neatsun Ziv, co-founder and CEO of OX Security, on episode 91

Instead of asking a developer to review an AI-generated pull request for security issues, the security context is loaded into the agent before it writes the first line, the practical application of least privilege for AI agents. Give a hundred developers a sanitization rule and you get ninety interpretations; give an agent the rule in its context file and it applies it on every commit. The same agents close the loop at the other end: a reachable, exposed finding does not need a ticket when an agent can open the pull request that moves the secret to the vault or bumps the dependency with passing tests. A human still owns the merge, for the reasons Chris Kirschke laid out when he asked whether you are actually going to give the agent write access. Our AI-accelerated vulnerability remediation program runs on exactly that split: reachability triage, AI-drafted fixes, a senior engineer’s review on every merge.

The AI-coding angle has its own hazards, from typo-squatted packages to agents that trust whatever is already on the machine; the vibe coding security post covers those. For this argument: prevention-first is how you keep machine-speed code generation from turning into machine-speed backlog generation.

Comparison table of shift-left security as practiced versus prevention-first security across six rows: when security shows up (after the code exists versus before the first line), what developers get (every finding ranked by CVSS versus the few reachable exposed issues with a drafted fix), who sets priority (the developer guessing blind versus security using reachability, KEV, EPSS and business context), unit of work (a ticket versus a pull request), what happens to the other 95 percent (ignored backlog versus documented accepted risk) and the success metric (findings closed versus time-to-remediate on reachable findings)
Same scanners in both columns. The difference is what travels with each finding, and who does the work.

Where do developer security hours actually go?

This matters to a founder because of the roadmap. Every hour an engineer spends on a finding that would never have been exploited is an hour not spent on the feature that closes the next deal, and the default shift-left queue spends almost all of its hours that way.

Diagram of where 100 hours of developer security time go: under a fix-everything-CVSS-7-and-above strategy about 4 hours land on vulnerabilities that were later exploited and 96 do not, under an EPSS-based strategy about 65 hours land on later-exploited vulnerabilities and 35 do not; a capacity check shows the typical organization fixes about one in ten open vulnerabilities per month, based on FIRST EPSS and Cyentia Institute data
Where the hours go. Same 100 hours, same FIRST data, two queues. The capacity grid is the Cyentia finding: about one in ten open vulnerabilities gets fixed in a given month, whatever the company's size.

The figure uses FIRST’s own numbers. Under a “fix everything CVSS 7 and above” policy, 3.96% of the prioritized vulnerabilities were later exploited, so if each fix takes roughly the same effort, about four of every hundred developer hours land on a vulnerability that mattered. Under an EPSS-based policy (fix anything with a 10% or higher probability of exploitation in the next 30 days), the figure is 65.2%. The honest caveat is coverage: the CVSS policy catches 82.2% of exploited vulnerabilities and the EPSS policy 63.2%, which is why exploitation scores are one input to the funnel and never the whole thing. Reachability and business context recover that coverage without buying back the noise.

Then there is the capacity grid. If a typical team clears about a tenth of its open vulnerabilities a month, the useful question is which tenth. A team that spends its tenth on the filled part of the bar ends the quarter safer than a team twice its size working the queue top to bottom.

Neatsun made one more point about where the hours go, aimed at founders buying their first security tools: buy for the company you are becoming. His argument on the episode was to over-invest a little and take a solution sized for a company ten times your current size, because the migration you would otherwise face at that size forces you to stop critical projects to fix your foundations. The same logic applies to the context layer. A security context file, a tagged inventory of services and their data, and an accepted-risk register are cheap at twelve engineers and painful to reconstruct at a hundred and twenty.

What should a startup do instead? A prevention-first plan

Moving from shift-left to prevention-first does not require ripping out tools. It changes what travels with a finding, who does the work, and where security shows up.

  1. Map your deployment context. Which services are externally exposed, which handle sensitive data, which connect to critical infrastructure. Tag them in code or in your cloud inventory. At startup scale this is an afternoon.
  2. Enrich findings at scan time. Attach deployment data, API exposure, reachability and data-flow context to every finding before it reaches a developer. Modern dependency scanners do part of this; the tags from step one do the rest.
  3. Prioritize by exploitability, not severity alone. KEV first, then EPSS probability, then CVSS as the tiebreaker, all gated by whether the finding is reachable and what it touches. A critical CVE in an internal test service is less urgent than a medium in a payment API.
  4. Move the constraints to where code is written. Secure defaults, an approved registry, a vault, sanitization libraries, and a short security context file in the repo that engineers and coding agents both read. Our secure AI SDLC work is mostly this step.
  5. Ship fixes, not tickets. For the findings that survive the funnel, have an agent or an engineer open the pull request with the fix and passing tests, and route the merge to a human reviewer. Track time-to-remediate on reachable findings; that is the number buyers and auditors want.
  6. Write down what you are not fixing. An accepted-risk register with the reason (not in production, not reachable, compensating control in place) and a trigger for re-review. This is also what a SOC 2 auditor samples in your vulnerability management evidence: a process that prioritizes by risk and operates as described, which beats a promise to fix everything that nobody keeps.
  7. Decide who owns it. Context does not maintain itself. Someone owns the tags, the register and the funnel, and at most startups that is either an engineer you would rather keep on the roadmap or an outside team. We wrote up that decision in In-House, On-Demand, or YOLO?; the embedded product security model is the middle path most of our clients take.

The point is structural: reduce the decision burden on the people who have to act, instead of generating more alerts.

The decision: stop shifting work, start shifting context

Shift-left was right about timing and wrong about delivery. Finding bugs early still beats finding them late, but a finding without context is a chore, and a thousand of them is a program that trains its own developers to ignore it. Prevention-first keeps the early part and fixes the delivery: constraints before the code, context on every finding that survives, fixes shipped as pull requests, everything else written down as risk you have chosen to carry for now. Teams that run it this way ship faster, and security finally answers the one question the business cares about: whether any of this would actually hurt.

YSecurity co-founders Sasha Sinkevich and Jon McLachlan, hosts of The Security Podcast of Silicon Valley, in a black and white studio portrait
Our co-founders Jon McLachlan and Sasha Sinkevich host the podcast. Their conversation with Neatsun is episode 91, released March 2026.

If you have a scanner queue in the thousands and no idea which tenth to fix, a free backlog assessment will run reachability triage against your findings and hand back what your true backlog looks like, and how fast it can get to zero. And if you want the argument from the man who watched 600 developers fix two things, Neatsun makes it in thirty-six minutes on episode 91.

Shift-left security frequently asked questions

What is shift-left security?
Shift-left security means doing security work earlier in the software development life cycle: threat modeling at design time, static analysis and dependency scanning on every commit, secret detection before code is pushed. The name comes from the left-to-right project timeline, and the logic is that a bug caught in the editor costs minutes while the same bug in production costs an incident. NIST's Secure Software Development Framework (SP 800-218) is the formal version of the idea.
Why does shift-left security fail in practice?
Because most programs shifted the work without shifting the context. Scanners in CI forward every finding to the developer who wrote the code, ranked by CVSS severity, with nothing about whether the code runs in production, is reachable, or touches sensitive data. Developers cannot triage what they cannot see, so they either fix everything slowly or ignore the queue. Neatsun Ziv's experience managing 600 developers: handed a prioritized list, they agreed to fix two items.
What is reachability analysis in application security?
Reachability analysis determines whether a vulnerability can actually be triggered in your application. For a vulnerable dependency, it asks whether your code ever calls the affected function and whether that library is loaded at runtime. For a code finding, it asks whether attacker-controlled input can reach the flawed line. Combined with exposure (is the service internet-facing?) and exploitation data (CISA KEV, EPSS), it separates the handful of findings that matter from the thousands that do not.
How many vulnerabilities actually need fixing?
Far fewer than the scanner suggests. FIRST's EPSS data shows about 2.7% of published CVEs see exploitation activity in a given month, and Datadog's State of DevSecOps research found only 18% of critical dependency vulnerabilities stay critical once runtime context is applied. Neatsun Ziv puts it at 95% of security fixes having zero risk impact. The job is to find the remaining few percent quickly and fix those first.
What is prevention-first security?
Prevention-first means deciding the security constraints before the code exists, so most findings are never created. In practice that is secure defaults and paved roads (an approved registry, a secrets vault, sanitization libraries), security requirements written down where engineers and AI coding agents will read them, and context about exposure and data loaded into the tools that generate code. Scanners still run; they become a check on the defaults rather than the primary control.
Should I prioritize vulnerabilities by CVSS, EPSS or KEV?
Use all three, starting from the end of that list. Start with anything in CISA's Known Exploited Vulnerabilities catalog, which CISA calls the authoritative source of vulnerabilities exploited in the wild. Then use EPSS, which estimates the probability of exploitation in the next 30 days, to rank the rest. CVSS describes how bad a flaw could be, not how likely anyone is to use it, so treat it as a tiebreaker. Reachability and business context then decide what your team fixes this sprint.
Do I have to fix every scanner finding to pass SOC 2?
No. SOC 2 auditors test that you have a vulnerability management process and that it operates as described: scans run, findings are assessed and prioritized by risk, remediation happens within the timelines your policy sets, and exceptions are documented. A risk-ranked queue with a written accepted-risk register is stronger evidence than a policy that promises to fix everything, because the auditor can see the promise being kept.
Does shift-left still matter with AI coding agents?
More than ever, but the point of intervention moves. AI coding agents write code faster than any review queue can absorb, so the useful place for security is the context the agent reads before it writes: which APIs are exposed, which databases the service touches, which sanitization and rate-limit rules apply. Load that once and every generated line inherits it. Agents can also draft the fix for reachable findings as pull requests, with a human on the merge.
Written by the team behind The Security Podcast of Silicon Valley

Put it into practice.