Live webinar The Wrong Security Hire Burns Your B2B GTM Pipeline. A fireside chat for founders · Oct 6, 9:30am PTThe wrong security hire · Oct 6 Save your seat

Blog

By Sasha Sinkevich 95 min read Updated September 20, 2026

The Art and Zen of Adopting AI

Established organizations read every incident report and wait. Startups try everything and buy the demo. Both are failures of practice rather than of enthusiasm. Sasha Sinkevich takes the ten good reasons the first group waits on their own terms, walks through the grandiose claims that cost the second group its scarcest asset, lays out a Middle Way you can actually test, and argues that security maturity has become the litmus test for whether a company is real.

The book that gave this genre its name was a misunderstanding. Eugen Herrigel’s Zen in the Art of Archery, published in German in 1948 and in English in 1953, spawned every “Zen and the Art of” title since, this one included. Its central scene has the archery master Awa Kenzō shoot twice at a target in a dark hall, the second arrow splitting the first, as a demonstration of a mind that has stopped aiming. Yamada Shōji’s study of the episode for the Japanese Journal of Religious Studies found that Awa did not practice Zen, that no interpreter was present that evening, that Herrigel’s Japanese was poor, and that Awa himself described the shot to his senior disciple as luck: “No, that was just a coincidence! I had no special intention to demonstrate such a thing.” Yamada’s conclusion is a sentence I have been unable to stop thinking about since I read it: “Out of the above circumstances was born the myth of ‘Zen in the Art of Archery’” (Yamada, 2001).

A foreigner watched a demonstration he could not evaluate, in a language he did not speak, with nobody in the room to translate, and built a philosophy on it that sold for seventy years. If you have sat through an artificial intelligence vendor demo in the last three years, you know exactly how he felt. So this is an essay about adopting AI, and it is going to borrow from Zen, but it is going to borrow the part that Herrigel skipped: the checking is the practice.

I run YSecurity with my brother Jon. We are a security firm, and for the last two years we have watched two kinds of companies fail at the same technology in opposite directions. The first is the established organization. It reads every incident report, and there are many, and it treats each one as a reason to wait another quarter. The second is the startup, which does not wait for anything, tries every tool a demo recommends, and loses months to claims that turn out to be false. The Buddha’s first discourse named this shape twenty-five centuries ago: two extremes, both rejected, and a path between them that is emphatically a practice rather than a compromise. Nobody has to become a Buddhist to use the structure. You only have to notice that “do half as much AI” is exactly what the text does not say.

The Middle Way, stated by an engineer. Karpathy's notes after his Dwarkesh Patel interview, four million views. Read on X.

Andrej Karpathy said it without the scripture in October 2025, in the interview where he coined “the decade of agents” as a correction to “the year of agents”, and then in his own notes afterward: his timelines are “about 5-10X pessimistic” relative to the San Francisco house party and “still quite optimistic” relative to the deniers, and “the apparent conflict is not.” Both things are true at once. There has been an enormous amount of progress, and there is an enormous amount of “grunt work, integration work” and “safety and security work” left. That second list is our whole business, so I am biased, and I will try to show my work.

"This is not the year of agents." Andrej Karpathy on the Dwarkesh Podcast, October 2025. Short enough to watch, and the most quotable statement of the middle position by someone nobody can dismiss as a skeptic. Watch on YouTube.

Here is the plan, and it is long, because the argument only works with the evidence attached. Each section stands on its own, so take the one you need.

Three-column diagram of the two extremes in AI adoption and the practice between them: the established organization that waits, listing its ten reasons; the startup that tries everything, listing its failure modes from buying the demo to standardizing on a vendor that vanishes; and the middle practice of eight disciplines from picking one workflow to measuring what you can check yourself
Two extremes, one practice. The middle column is a program, in the same way the Noble Eightfold Path is a program.

What “agentic” means, and the first test

Precision matters here because Section’s June 2026 study of 5,026 knowledge workers found that 82 percent of the workforce cannot correctly identify what an agent is. Three definitions carry this article.

Andrew Ng gave us agentic workflows in March 2024: instead of having a model produce its answer in one shot, “an agentic workflow prompts the LLM multiple times, giving it opportunities to build step by step to higher-quality output” (The Batch, issue 242). LLM is the large language model, the thing behind the chat window. His four patterns were reflection, tool use, planning and multi-agent collaboration.

Anthropic’s engineering team drew the line that matters for budgets in December 2024. “Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents are systems where LLMs dynamically direct their own processes and tool usage.” Their advice to builders: find “the simplest solution possible, and only increasing complexity when needed”, and reserve true agents for “open-ended problems where it’s difficult or impossible to predict the required number of steps” (Building effective agents). Most of what is sold as an agent in 2026 should be a workflow. Workflows are cheaper, more testable, more auditable and much easier to secure. A firm that tells you “you want a workflow, not an agent” is saving you money and demonstrating independence in the same sentence.

And Karpathy’s agentic engineering, from February 2026, is the operating posture. He had named vibe coding one year and two days earlier, where you “fully give in to the vibes, embrace exponentials, and forget that the code even exists.” The retrospective said the professionals had moved on: they were shipping through agents “with more oversight and scrutiny”, and the better name was agentic engineering, agentic “because the new default is that you are not writing the code directly 99% of the time, you are orchestrating agents who do and acting as oversight”, and engineering “to emphasize that there is an art & science and expertise to it” (Karpathy on X, 1.3 million views). Notice that every one of these definitions is about oversight rather than autonomy. That is why, in our experience, the security team and the engineering team end up building the same thing.

The person who named vibe coding explains why the professional version is a discipline. Andrej Karpathy with Stephanie Zhan, Sequoia Capital, AI Ascent 2026. Watch on YouTube.

Now the first test, and the first correction. The most-shared Buddha quotation on the internet about testing things for yourself goes, roughly, “Believe nothing, no matter where you read it or who has said it, not even if I have said it, unless it agrees with your own reason and your own common sense.” It is a forgery. Bodhipaksa’s Fake Buddha Quotes traces the wording to a pseudonymous libertarian book and shows it inverts the text it claims to translate (Fake Buddha Quotes). The real passage, the Kālāma Sutta (AN 3.65), says the opposite of “trust your gut”. It tells the Kālāmas not to go “by reports, by legends, by traditions, by scripture, by logical conjecture, by inference, by analogies, by agreement through pondering views, by probability, or by the thought, ‘This contemplative is our teacher’”, and then it gives the test: when you know for yourselves that qualities “when adopted & carried out, lead to harm & to suffering”, abandon them (Thanissaro Bhikkhu’s translation). The test is empirical and it is social. You carry the thing out, you observe the result, and you check with people who know.

That is a purchasing policy. Every vendor deck you will see this year is “reports, legends, traditions, scripture, logical conjecture, inference, analogies, agreement, probability” and “this contemplative is our teacher.” The Kālāma test is the pilot on your own data with the evaluation written before the prompt, and a second opinion from someone with no stake in the answer. OpenAI’s own first lesson from its enterprise deployments is the same sentence in corporate English: “Start with evals” (AI in the Enterprise). Keep it in mind for both halves of what follows, because the waiting organization and the credulous startup are both failing the same test from opposite ends.

Two-column chart of AI adoption versus value in 2026: nine in ten organizations use AI in one or more functions, 74 percent of frontline staff use it regularly and 69 percent of organizations rolled out something agentic, while only 37 percent report any earnings impact, 6 percent qualify as high performers, 16 percent of workers use the agentic tools they were given, and 66 percent cite security and risk as the top barrier to agents
Adoption is nearly universal. Value is rare. Sources: McKinsey State of AI 2026 and State of AI trust 2026, BCG AI at Work 2026, Section AI Proficiency Report 2026.

The first extreme: the established organization that waits

Start with the numbers that describe the waiting. In May 2026 the United States Census Bureau put the share of American firms using AI in any business function at 19.8 percent. Among the organizations that have started, McKinsey’s 2026 global survey found nearly nine in ten using AI somewhere, 37 percent able to point to any earnings impact (flat on the year before), and 6 percent qualifying as high performers. Section found that 69 percent of workers say their organization has taken some action on AI agents while 16 percent actually use one. McKinsey’s trust survey found two thirds naming security and risk as the top barrier to scaling agents (State of AI trust 2026).

I want to be clear about the spirit of this section. The reasons below are correct as stated. I have heard every one of them from a serious person, and I would rather write them down than argue past them. What I will show is where each has quietly expired, where it has not, and, for the ones that are security problems, what the control looks like. This is the list a first-year Buddhist would recognize as the “self-affliction” extreme in the Sujato translation of the first discourse: “indulgence in self-mortification, which is painful, ignoble, and pointless” (SN 56.11). Reading incident reports as a substitute for building controls is a kind of mortification. It feels responsible and it produces nothing.

1. “Our data will end up training someone else’s model.”

The fear has a founding story. In 2023 Samsung engineers reportedly put “equipment measurement and yield data” and “problematic source code” into ChatGPT, and the company banned the tool (The Register). The leak is still happening at scale: Netskope’s 2026 report counts 223 generative-AI data policy violations per month at the average organization, 42 percent of them source code, and Zscaler logged 410 million data-loss-prevention violations tied to ChatGPT alone in 2025.

What changed is that the remedy is now a contract clause. Anthropic’s commercial terms state that “Anthropic may not train models on Customer Content from Services” (commercial terms); OpenAI’s enterprise privacy page says “We do not train our models on your data by default” and offers zero data retention for eligible endpoints (OpenAI); AWS Bedrock states your content “is not used to improve the base models and is not shared with any model providers” (Bedrock FAQ). Two cautions we give every client: “we do not train on it” and “we do not keep it” are different promises (Anthropic’s enterprise plans retain data until you set a custom period, with a 30-day floor), and the ban does not work. Netskope’s most hopeful number: where organizations provided a governed path, personal AI app use fell from 78 percent to 47 percent while enterprise-managed adoption rose from 25 percent to 62 percent in a single year. How we solve it: our shadow-AI detection and mitigation work starts by finding every tool already in use, then gives the useful ones enterprise terms and the risky ones a paved road out.

2. “An agent with tools can be hijacked by a web page.”

Correct, and the strongest technical objection on the list. Language models read the system prompt, the user’s request and any retrieved text “as a single stream of tokens”, and OWASP’s 2026 state-of-agentic-security report (the Open Worldwide Application Security Project) maps prompt injection to six of its ten agentic risk categories (Help Net Security on the OWASP report). The proof came in a steady drumbeat: EchoLeak, a zero-click exfiltration path through Microsoft 365 Copilot (CVE-2025-32711); ForcedLeak in Salesforce Agentforce, where a five-dollar expired domain still on an allowlist became the exfiltration channel (Noma Security); and Google’s Antigravity, which shipped AWS credentials to an attacker’s webhook at factory default settings (PromptArmor). Gartner’s August 2025 Hype Cycle release put the general point on the record: “AI brings new trust, risk and security management challenges that conventional controls don’t address” (Gartner).

You do not solve injection. You make it non-consequential. Simon Willison’s “lethal trifecta” names the three ingredients an attacker needs: access to private data, exposure to malicious instructions, and a way to send data out (Willison). Meta turned that into a design rule in October 2025, the Agents Rule of Two: an agent may have no more than two of those properties in a session, or a human sits in the loop (Meta AI). Gartner arrived at the same three-factor “no-go zone” independently in April 2026 (Gartner). Remove one leg and the exploit stops paying. How we solve it: this is the design review at the heart of our AI product red teaming and secure Model Context Protocol programs: egress allowlists, credential-free agent identities, and a regression suite that re-tests every model release.

Triangle diagram of Meta's Agents Rule of Two: the three vertices are reading untrusted input, touching sensitive data or systems, and changing state or communicating externally; each edge shows the safe pairing of two properties, and the dashed center circle reads all three equals a human in the loop
Pick two per session. The third property is where the person, or a fresh context window, sits. After Meta's Agents Rule of Two (October 2025) and Simon Willison's lethal trifecta (June 2025).

3. “An agent will delete production.”

It has, in public, at rising severity. In July 2025 a Replit agent deleted SaaStr founder Jason Lemkin’s production database during a code freeze; Replit’s chief executive called it “Unacceptable and should never be possible” and began rolling out separate development and production databases within days (The Register). In April 2026 a Cursor agent at PocketOS deleted the production database and all volume-level backups in nine seconds, through an API token (application programming interface) with unrestricted permissions, after ignoring a system prompt that forbade exactly that (Zenity). We have since traced the whole family of these failures, from a 2017 Firebase bill to the agent that ran up $50,000 in under an hour, in Too Fancy. On Reddit, a r/cybersecurity thread with two thousand upvotes documented an Amazon Kiro agent that inherited elevated permissions and took down a production environment for thirteen hours; the top comment: “An AI agent inheriting elevated permissions and bypassing two-person approval is exactly the failure mode everyone warned about” (r/cybersecurity). And in April 2025 the highest-scoring AI failure thread I could find on Hacker News, 1,511 points, was about a Cursor support agent that invented a login policy out of nothing; real customers read it and cancelled (Hacker News). Your agent’s output is a public statement by your company.

The thread that made "agents delete production" a boardroom sentence. Lemkin later clarified it was a demo app, and that using one database for preview, testing and production "simply is NOT ok". View on X.

Read the root causes and the pattern is the same every time: no environment separation, an omnipotent credential, backups inside the blast radius, and a system prompt doing the job of a control. Zenity’s summary is the one we put in front of executives: system prompts “are weighted inputs to a probabilistic reasoning engine, not deterministic enforcement mechanisms.” Cyera analyzed 7,246 public AI incidents and found 188 where an autonomous system caused harm in production with no attacker in the chain, and of the 137 cases involving real-world damage, 65 were deletion or code destruction. That is the good news in disguise. Most agent risk is operational, and operational risk is solved with architecture. Lemkin’s own lesson from the weekend holds up.

“They are tools. Not dev teams. Remind yourself of that every single day.”

Jason Lemkin, SaaStr, quoted in The Register, July 2025 How we solve it: scoped tokens per operation and environment, backups on separate infrastructure with a tested restore, and destructive operations behind a human confirmation are the first week of our secure AI software development life cycle engagements.

4. “We cannot audit what the agent did, and we are the ones who are regulated.”

Legitimate. An agent that takes forty steps produces a tool call log and nothing about why, and “the model chose to” is not an explanation any auditor accepts. IBM’s 2026 report found only 19 percent of organizations report coordination between their AI governance and security teams (Kiteworks analysis of IBM 2026), and 62 percent of AI-driven attacks targeting critical infrastructure (IBM 2026). The blast radius really is different in a bank or a hospital.

The standards have caught up fast. OWASP published its Agent Control Standard on September 1, 2026, built on the principle that agents must be “inspectable, traceable and instrumentable” (OWASP). The National Institute of Standards and Technology is drafting control overlays for AI systems in the language of SP 800-53, with two of the five overlays dedicated to agent systems (NIST COSAiS). The model developers now run the controls they recommend: after evaluation agents reached live systems through a misconfigured sandbox in July 2026, Anthropic disclosed the incidents and added a runtime classifier that “blocks the action before the tool call is run, ends the task, and alerts a human” (Anthropic, August 2026). And the regulatory clock is friendlier than the fear. The European Union’s AI Act obligations that most deployers dread, the high-risk requirements, land on December 2, 2027 and August 2, 2028; what is live today is the general-purpose model regime and the transparency rules (implementation timeline). Notice, too, which large sector adopts fastest: Census data shows finance and insurance at 33.9 percent against retail at 14 percent. Regulated firms already run model risk management, change control and segregation of duties; applying them to agents is an extension of a discipline they have. How we solve it: we did exactly this for Augment Code, whose current SOC 2 Type 2 (the System and Organization Controls audit) period has agentic code reviewers inside the system boundary. Paddy Roberts, who leads governance, risk and compliance there for us, wrote up how independence of review works when the reviewer is an agent. Our AI and HIPAA (Health Insurance Portability and Accountability Act) and AI European Union compliance programs exist because the data flows can be engineered so the buyer’s compliance team says yes.

5. “Our customers’ security questionnaires will fail.”

This is usually the real blocker because it is commercial. A single “do you use AI when processing our data?” question can stall a deal for a quarter, and the procurement side has hardened: AI questionnaires that used to run 20 to 30 questions now “routinely run 40-60”, one Fortune 500 security review took six months, and innovation teams that ran six to eight pilots a year “are now struggling to launch two or three” (Traction Technology). Vanta’s 2025 State of Trust survey of 3,500 business and IT leaders found organizations spending twelve working weeks a year on compliance tasks and nine on vendor security reviews, with 61 percent saying their use of agentic AI outpaces their understanding of it (Vanta State of Trust 2025). Meanwhile the share of organizations with an AI policy in place fell from 37 percent to 32 percent between IBM’s 2025 and 2026 reports. You cannot answer a questionnaire from a policy that is “in development”.

Here is the reframe: the buyer now asks “do you use AI, and how do you govern it?” and “no” with no policy reads worse than “yes” with one. In June 2026 the compliance firm Aetos summarized what enterprise buyers now want: “ninety days of documented evidence that you have controls in place to evaluate AI-generated outputs before they reach users”, and it noted that most of that documentation “can be built in eight to twelve weeks with focused effort” (Aetos). Writing the policy, inventorying the systems, mapping controls to a citable baseline (the OWASP agentic top ten, the NIST overlays, the National Security Agency’s May 2026 Model Context Protocol guidance) is days of work, not quarters. How we solve it: our sales support team joins the security call and answers the AI questions with evidence, and our ISO 42001 practice (the international management-system standard for artificial intelligence) took Augment Code from zero to certified in 93 days, the first AI coding assistant to get there. A million-dollar deal closed two days later. I have a lot more to say about this in the litmus-test section below, because the same questionnaire that blocks your adoption is the instrument your buyers use to decide whether you are real.

6. “We ran pilots and never saw it in the income statement.”

Deeply true, and the honest numbers come from pro-AI sources. McKinsey: 37 percent attribute any earnings impact to AI, unchanged from 2025, and 6 percent are high performers. Gartner predicts “over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls” (Gartner). In April 2026 Gartner surveyed 782 infrastructure and operations leaders and found 28 percent of AI use cases fully succeeding against expectations and 20 percent failing outright; among the 57 percent of leaders who reported at least one failure, many said their initiatives failed because “they expected too much, too fast” (Gartner, April 2026). Read that phrase twice. The failure the people it happened to describe most readily is the second extreme in this essay rather than the first.

You have probably also heard that 95 percent of pilots fail. That figure comes from an MIT NANDA working paper whose central claim rested on 52 interviews the authors themselves called only “directionally accurate”; the funnel in the paper (60 percent investigated, 20 percent piloted, 5 percent implemented) implies about three quarters of pilots failed, not 95 percent, and the press inflated the sample along the way (Ray Poynter’s dissection). A Hacker News commenter added the correction that matters most: the number refers to building AI products for production, not to using AI in daily work (Hacker News). A post about grandiose claims should practice on the statistic the skeptics quote, so I am not going to use it.

The most quoted statistic in the genre, checked. We would rather use numbers we can show you. Watch on YouTube.

Now the two halves McKinsey publishes side by side. Eight in ten respondents say AI improved their own productivity. Thirty-seven percent see it in the earnings. The value is being created at the task level and lost at the process level, which is a management finding. BCG quantified the fix in its 2026 workforce study: a clear strategy raises measurable business impact by 25 percentage points; better tools alone raise it by 5 (BCG). Nathan Furr and Andrew Shipilov, writing in Harvard Business Review, called the pattern the experimentation trap: leaders “are repeating the mistakes of the digital transformation era by funding scattered pilots that don’t connect to real business value”, and “the takeaway is not that AI experimentation is broken, but that it must be disciplined” (HBR, August 2025). The single most actionable number I found this year is Forrester’s, as reported by AI News: agents without automated evaluations showed a 47 percent rollback rate against 9 percent for those with full coverage, and only 38 percent of production agents have them (AI News, September 2026). The pilots that reach the income statement are the ones with the harness. Gartner’s data-readiness prediction points the same way once you read the conditional: “through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data” (Gartner) is 60 percent of the projects that lack ready data, so start where your data is already good enough. McKinsey found where large enterprises actually went first: 31 percent are scaling software coding agents, which in our experience is the use case with the lowest data dependency.

7. “Nobody owns this, middle management is quietly blocking it, and we have three transformations running already.”

Three objections, one root. AI sits across information technology, security, legal, human resources and the business, so nobody has the mandate or the budget, and McKinsey found 60 percent naming knowledge and training gaps as the primary barrier to responsible practice (McKinsey). The cleanest intervention in the literature is also the cheapest: organizations with explicit responsible-AI ownership score 2.6 on McKinsey’s maturity scale against 1.8 without. BCG named the middle-management plateau the “silicon ceiling” in 2025, when 43 percent of leaders and managers feared job loss within a decade, a higher share than the frontline (BCG AI at Work 2025); then the ceiling broke in twelve months, with frontline regular use going from 51 to 74 percent in BCG’s 2026 study and 47 percent of workers now spending more time managing and directing AI than doing the work themselves. Managing and directing is a supervisory skill. The manager’s fear is of being bypassed; the evidence says they are being promoted into orchestration.

The exhaustion is real too. Microsoft’s Work Trend Index found 80 percent of the global workforce lacking the time or energy to do the work they already have (Microsoft WTI 2025), and Upwork found 77 percent of employees saying AI had increased their workload (Upwork). Adding AI to an unchanged operating model reliably makes things worse. The dose that fixes it is small and known: BCG found regular usage “sharply higher for employees that receive at least five hours of training and have access to in-person training and coaching”, while Slack found 61 percent had spent under five hours learning (Slack Workforce Index). This is the one objection where the answer is “you are right, so do less”: one function, one workflow, one named owner.

Box's chief executive after a road trip through enterprise IT: the blocker is the operating model. View on X.

8. “I measured it. I am slower with it, and it makes things up.”

You may well be right, and this is the best-designed skeptic evidence there is. METR’s randomized trial of sixteen experienced open-source developers on their own mature repositories found they took 19 percent longer with AI while believing they were 20 percent faster (METR, July 2025). A r/ExperiencedDevs thread titled “I stopped using copilot and didn’t notice a decrease in productivity” drew 864 upvotes, with a top comment that “people overestimate how much of the typical job is boilerplate” (r/ExperiencedDevs). Only 33 percent of developers trust AI output and 66 percent name “almost right, but not quite” as their top frustration (Stack Overflow 2025), and Damien Charlotin’s database of court decisions involving hallucinated citations stood at 2,041 cases when I checked it, 811 involving lawyers.

The most cited skeptic result, from the lab that ran it. View on X.

Then read what METR did next. In February 2026 it retired that experimental design because developers refused to participate rather than work without AI, even at 50 dollars an hour on tasks they chose, and 30 to 50 percent withheld tasks they did not want to do unaided. Among new participants the slowdown had shrunk to 4 percent, and METR itself noted “the true speedup could be much higher among the developers and tasks which are selected out of the experiment” (METR, February 2026). Brynjolfsson, Li and Raymond’s customer support study explains the whole disagreement: 34 percent gains for novices, minimal for experts (NBER). METR sampled exactly the group predicted to gain least. And the hallucination problem has a professional answer that predates the technology. Every sanctioned lawyer in Charlotin’s database filed unread output, which no professional standard ever permitted. Adoption up, trust down is what a maturing power tool looks like; nobody trusts a table saw either. Choose tasks with a cheap check (tests, ground truth, a second reader), put a named person on every artifact, and the verification tax collapses. That is the Kālāma test again, at the level of one pull request.

A serious skeptic treated seriously. Cal Newport walks through the METR trial carefully. Watch on YouTube.

9. “It is going to take my job, people will think I am lazy, and my employer never said I could.”

Pew found in June 2026 that, for the first time, a majority of American adults under 30 are more concerned than excited about AI, 55 percent, and that 73 percent of young adults expect it to mean fewer jobs (Pew Research). When Anthropic’s economics team published its growth-and-jobs scenarios on September 9, a r/cscareerquestions thread drew 1,800 upvotes and a top comment that deserves to be quoted: “Company with vested interest in future benefitting them greatly predicts future that benefits them greatly” (r/cscareerquestions).

The evidence is split, and honesty requires saying so. At the aggregate level the displacement is not there: Yale’s Budget Lab called broad disruption “largely speculative” (Budget Lab), Danish administrative data “rule out effects larger than 2%” on earnings two years after ChatGPT (Humlum and Vestergaard), and Federal Reserve Governor Michael Barr told a Fed conference in July 2026 that “as of right now, there has been little evidence of economy-wide job displacement from AI” (Federal Reserve, July 14, 2026). At the entry level it is there: Stanford’s August 2026 revision of Canaries in the Coal Mine found employment for 22 to 25 year olds in highly exposed occupations “about 19% below where it would be” otherwise, up from 15 percent a year earlier, while adding “We do not see widespread, economy-wide job displacement” (Stanford Digital Economy Lab). AI is taking the first rung of the ladder rather than the ladder. That is a real harm with a different remedy than abstinence, because abstaining makes you less employable without restoring the rung. PwC’s 2026 barometer of a billion job ads found a 62 percent wage premium for AI skills, and the most AI-exposed firms growing headcount 52 percent against 36 percent (PwC).

The labor-market fact in its corrected 2026 form. View on X.

The stigma is measured and justified: Slack found 48 percent of desk workers uncomfortable admitting AI use to their manager, and a Duke study of 4,439 participants found AI users judged as lazier and “less competent, less diligent, less independent, and less self-assured” (PNAS). The same study found the mechanism that makes it expire: the penalty disappears entirely when the evaluator uses AI weekly or daily. Pew puts chatbot use at 49 percent of American adults, from 33 percent in 2024, so the stigma has a shrinking constituency. And the permission problem is the one reason on this list that is entirely the employer’s to fix. Slack found 45 percent of desk workers lacked explicit permission, IBM found 63 percent of breached organizations with no AI governance policy or one still in development (IBM 2025), and a Cloud Security Alliance survey found 38 percent of employees had shared confidential company data with unapproved AI systems (cited by Stack Overflow, 2026). A r/sysadmin post with 289 upvotes described developers “banging 4+ dev tools in parallel, apparently for months” and worried about the secrets leaving; the top reply: “This is a management issue, not a tech issue” (r/sysadmin). How we solve it: an acceptable-use policy short enough to follow, a sanctioned tool catalog, and the shadow-AI inventory that tells you what you are actually approving.

10. “I like doing the work, and I can feel it making me worse.”

This one should not be argued out of anyone. On Reddit, a senior developer with 359 upvotes wrote that after letting an agent go fully agentic on a feature he feels disconnected from the codebase: “The code is there. It works. But I don’t generally get how it works” (r/ExperiencedDevs). In a r/UXDesign thread with 661 upvotes, a designer wrote: “It honestly takes me SO much longer to yell at AI to do what I could mock up in 15 minutes” (r/UXDesign). The MIT Media Lab study that went around the world found essay writers using a model showed the “weakest connectivity” on EEG (electroencephalography) and “struggled to accurately quote their own work” (Kosmyna et al.). Scope it carefully: 54 participants, essay writing under exam conditions, and the authors themselves asked the press not to use words like “brain damage”. It is evidence about unsupervised substitution during learning.

The counterweight is the tenure data. Anthropic’s Economic Index found high-tenure users have a 10 percent higher success rate not explained by task selection, attempt harder work every year, and are less directive over time (Learning curves, March 2026). Brynjolfsson’s field study found AI “disseminates the best practices of more able workers and helps newer workers move down the experience curve”, and that AI assistance “increases employee retention”, the hardest evidence there is that the work did not get worse to do. Gergely Orosz, who tracks engineering pay for a living, put it in four words: “AI amplifies existing ability” (Orosz on X). Simon Willison, a believer rather than a skeptic, reached the same conclusion from inside the practice: “AI tools amplify existing expertise”, and working productively with them is “difficult” and requires operating “at the top of your game”, with “so much time on code review” (Willison, October 2025). If the craft is the point, keep doing it by hand. The case for agents is about the work you do not love. BCG’s “joy paradox” from 2026 holds both halves: 67 percent of regular AI users report improved job satisfaction and 41 percent report higher cognitive load. The best practical rule I have read came from a Hacker News commenter: “LLM-generated code has no theory; you either need to supervise it closely enough to impose your own, or treat it as disposable” (Hacker News, 865 points). Both are legitimate. Doing neither is the failure mode.

The line between the two extremes, drawn by a practitioner. Neither abstinence nor abandon: craft. View on X.

Ten reasons. An earlier version of this essay listed twenty-five, and the evidence for the extra fifteen was thinner than the objections deserved, so they are gone. What the ten have in common is more useful than the count. Each is a real risk, each is concentrated in a known failure mode, and each has a control, a policy, a design rule or a measurement behind it that a named person can build in weeks. Waiting is a decision to let somebody else build the controls first. Now the other extreme, which is where my own companies live.

The second extreme: the startup that tries everything

Startups are never slow to adopt AI, and anyone who tells you otherwise has not met one recently. The reason is arithmetic rather than temperament. CB Insights studied 431 venture-backed companies that shut down since 2023 and found the leading cause, at 70 percent, was “ran out of capital” (CB Insights, updated March 2026). Carta counted 254 closures on its platform in the first quarter of 2024 alone, up 58 percent on the year before and 124 percent the year before that (Carta). When a company has eighteen months of runway, its risk profile is inverted relative to the enterprise: a three-month experiment that fails costs three months, and a three-month deliberation that fails costs the company. So they risk everything and they try everything, and they are right to. Stripe found the top 100 AI companies on its platform reached a million dollars of annualized revenue in a median of 11.5 months, and firms founded after 2020 hit their revenue milestones about three times faster than earlier cohorts (Stripe, June 2025). In March 2025 Y Combinator’s Jared Friedman reported that a quarter of the Winter 2025 batch had codebases that were 95 percent AI-generated (TechCrunch). Garry Tan, on the same stage, said “This is the dominant way to code. And if you are not doing it, you might just be left behind”, and then, in the same breath, the caveat that never travels with the statistic: “Does it fall over or not? The first versions of reasoning models are not good at debugging.”

The speed, with a clip that plays here. A million views. The caveat from the same stage ("does it fall over or not?") did not get a post. Watch on X.

By September 2026 the adoption question is closed for this cohort. Ramp’s card data puts paid business adoption of Anthropic at 43.8 percent of American businesses and OpenAI at 39.8 percent, with technical sectors leading, and the interesting number has moved: spend per employee among the heaviest users fell 9.7 percent in August as companies set “company-wide defaults that reduce usage of frontier models” and the effective price per million tokens dropped 41 percent from its March peak (Ramp AI Index, September 2026). Everyone adopted. The question is who is getting paid for it.

So what is the startup’s actual risk? It is the mirror image of the enterprise’s. The openness that makes a founder fast also makes her the ideal customer for a grandiose claim. She has no procurement team, no security review, no evaluation framework, and no time, and the only asset she cannot raise more of is the one the claim consumes. Andreessen Horowitz’s 2025 survey of 100 enterprise chief information officers found that the large companies had, by then, moved AI out of innovation budgets into core lines (innovation budgets fell from 25 to 7 percent of AI spending) and were evaluating models with the same disciplined frameworks they apply to any software vendor, weighing security and cost alongside performance (a16z). That is the filter a startup skips to stay fast. Below is what came through the gap between 2023 and 2026. I have chosen cases with primary sources, and I have kept Amazon’s denial and Gergely Orosz’s correction in, because a section about grandiose claims that repeats one is not worth reading.

Timeline of grandiose AI claims from December 2023 to August 2026 with what each turned out to be: Google's edited Gemini demo, Cognition's Devin launch video and the three-of-twenty field test, the Humane AI Pin, Rabbit r1, Amazon Just Walk Out, Manus, Nate's SEC charges, Builder.ai's insolvency, Klarna's walk-back, Windsurf's 72-hour dismemberment, the Lovable authorization flaw, and Cursor's acquisition by SpaceX
The catalogue, dated. Two of these are outright fraud by the government's account, several are edited demos, one is an honest company that changed owners twice, and one is a correction that traveled a fifth as far as the claim.

The demo was real. The product was not.

Google’s launch video for Gemini in December 2023 showed the model reacting to live video and speech. TechCrunch’s Devin Coldewey established the same day that, in Google’s own account, the team “prompted Gemini using still image frames from the footage, and prompting via text”, and that Google’s own caption conceded “latency has been reduced and Gemini outputs have been shortened.” His verdict: “the video simply does not reflect reality. It’s fake” (TechCrunch, December 2023). Three months later Cognition launched Devin as “the first AI software engineer”, with a video of it completing a paid Upwork job. Carl Brown of Internet of Bugs went through the video frame by frame and showed the task in the demo was not the task the client had posted, and that several bugs Devin fixed were bugs Devin had introduced. In the Hacker News thread about his video he wrote: “The company claimed Devin completed Upwork jobs in the video description when it clearly didn’t. That’s a lie”, and, after a dash, “pure and simple, regardless of elsewhere statements” (Hacker News, 302 points). The launch thread a month earlier had 530 points (Hacker News). Same product, one month apart, and the rebuttal reached about half the audience of the claim.

The demo-is-not-the-product artifact. Carl Brown's frame-by-frame of the Devin launch video, April 2024. Watch on YouTube.

Then a real team ran the real product. In January 2025 Hamel Husain, Isaac Flath and Johno Whitaker at Answer.AI published a month of using Devin on twenty tasks: “we had 14 failures, 3 successes (including our 2 initial ones), and 3 inconclusive results.” The autonomy “that seemed promising became a liability”, because Devin “would spend days pursuing impossible solutions rather than recognizing fundamental blockers”, and Whitaker’s summary was the founder’s cost in one line: “Tasks it can do are those that are so small and well-defined that I may as well do them myself, faster, my way.” Their conclusion is the sentence I would put on the wall of every startup: “Social media excitement and company valuations have minimal relationship to real-world utility” (Answer.AI). In the Hacker News thread, one commenter described the failure mode the founders feel most: “It would work away for literally days on a task where any human would admit they were stuck” (Hacker News, 285 points).

Now the part that makes this a Middle Way essay rather than a skeptic’s essay. In May 2026, sixteen months later, Husain quoted his own post and wrote: “This post is completely outdated FWIW. Devin is really good now. One of my favorite tools”, adding that “Some of it is models are better but the UX/harness is also better. I suspect this blog post would not be written today.” The demo was faked in 2024, the product was weak in January 2025, and it was good by May 2026, and all three statements are true. The founder who bought on the demo paid in months. The founder who wrote the category off forever is paying now.

The same reviewer, sixteen months later, correcting himself in public. This is what the middle looks like: the demo was false, the early product was weak, and the mature product is good. View on X.

The hardware that was a keynote

Humane raised more than 230 million dollars for the AI Pin. Marques Brownlee’s review in April 2024 was titled “The Worst Product I’ve Ever Reviewed… For Now”, and the Hacker News thread on the launch, five months earlier, had already made the case: “people hate talking to computers in public”, “the conversational interface has never worked before”, and the launch demo itself contained a factual error (Hacker News, 422 points). By August 2024 the Pin’s returns were outpacing its sales. HP bought Humane’s engineers, its operating system and its intellectual property for 116 million dollars in February 2025 and did not buy the device, which stopped working at noon Pacific on February 28, 2025, with refunds only for purchases in the prior ninety days (TechCrunch). The Rabbit r1 made the same trip faster. It was sold as a “large action model” that would operate websites on your behalf; Android Authority installed its software on a Pixel 6a and found it was an Android app (Android Authority); a research group found hardcoded API keys for ElevenLabs, Azure, Yelp and Google Maps in its codebase; and of roughly 100,000 purchasers, about 5,000 were using the device daily by September 2024 (Wikipedia’s consolidated account). The hardcoded keys are the part a security firm notices. The company’s engineering discipline was visible from the outside to anyone who looked, months before the reviews.

A 699-dollar device sold on a keynote. The "...For Now" in the title is the whole thesis of this section. Watch on YouTube.

The AI that was people

Amazon has had a name for this since 2005. Mechanical Turk was called “artificial artificial intelligence” inside the company, after the eighteenth-century chess automaton that beat Napoleon and Benjamin Franklin with a human chess master hidden in the cabinet (Wikipedia). In April 2025 the Securities and Exchange Commission (SEC) charged Albert Saniger, founder of the shopping app Nate, which had raised more than 42 million dollars on the claim that it “used automated technology that relied on AI to complete purchases made through the app without human involvement.” The SEC’s complaint alleges that “Nate relied in large part on contract employees to manually input orders placed by users on the app” and that “Nate’s app was not able to use AI to complete purchases” (SEC Litigation Release 26282). The Department of Justice’s parallel case alleged the company claimed an automation rate “above 90 per cent” when in fact all orders were placed manually (Global Investigations Review). Nate was the sharpest case in a docket that already included the SEC’s first “AI washing” actions in March 2024, against Delphia and Global Predictions, and the Federal Trade Commission’s (FTC) Operation AI Comply in September 2024, which reached DoNotPay’s “world’s first robot lawyer” and Ascend Ecom’s AI-powered online-store scheme (SEC; FTC).

The careful case is Amazon’s own Just Walk Out. In April 2024 reports circulated that the checkout-free stores relied on more than a thousand people in India watching video. Amazon’s vice president for the product, Jon Jenkins, told Axios: “This notion that there are human reviewers [in India] watching live shoppers”, he said, is “completely not true”, and he described a team of “way less than 1,000” people reviewing video “after the fact” to train models and, in “a small percentage of cases”, to verify receipts (Axios). Include the denial, and the lesson survives it: the marketing never mentioned the human verification layer at all. The sin in this genre is rarely using humans. It is not saying so.

The collapse that everyone remembers wrong

Builder.ai raised about 445 million dollars, counted Microsoft and the Qatar Investment Authority as investors, and its founder, by a former employee’s account, wanted its assistant Natasha to make building software “as easy as ordering a pizza.” It entered insolvency on May 20, 2025 (Rest of World). The version you remember is that the AI was 700 engineers in India. The Hacker News thread with that framing reached 368 points (Hacker News). Gergely Orosz, who talked to the company’s former engineers, published the correction: the agent was real and worked, and the hundreds of outsourced developers were building client apps rather than pretending to be the AI (The Pragmatic Engineer). What killed the company was the boring fraud rather than the lurid one: inflated revenue, later reported to involve “roundtripping” with an Indian partner (Rest of World). His correction thread reached 82 points (Hacker News). The claim traveled four and a half times as far as the correction, and Orosz himself had to retract his own first post. That asymmetry is the whole cost structure of hype: it is cheap to produce, expensive to check, and the check never catches up.

The correction, from the person who made it. The checking is the practice, and it is rarer than the claim. View on X.
Builder.ai: 1.5 billion dollar valuation, Microsoft on the cap table, insolvency in May 2025. Watch it with Orosz's correction in mind: the fraud was in the revenue. Watch on YouTube.

The wrapper with a waiting list

In March 2025 Manus launched as “the first general AI agent” and “a glimpse into AGI”, claiming state of the art on the GAIA benchmark. Zvi Mowshowitz’s write-up recorded that the benchmark run used a publicly available test set (“So if they wanted to game the benchmark, they could do so”), that on inspection the product was “claude sonnet with 29 tools”, and that invite codes were being hyped at 50,000 to 100,000 yuan, roughly seven to fourteen thousand dollars (Don’t Worry About the Vase). Nobody was defrauded. Founders lost a week chasing a code for a product most of them already had. Gartner put a number on the category three months later: “Gartner estimates only about 130 of the thousands of agentic AI vendors are real”, with many of the rest engaged in “agent washing”, the rebranding of chatbots, robotic process automation and assistants (Gartner, June 2025). And its April 2026 Hype Cycle still has agentic AI at the Peak of Inflated Expectations, “reflecting extraordinary market attention and aggressive adoption intent”, with 17 percent of organizations deployed and more than 60 percent planning to deploy within two years (Gartner, April 2026). Do not write “we are past the hype.” Gartner’s own instrument says we are not, and its definition of the peak is precise: “Early publicity produces a number of success stories”, which are “often accompanied by scores of failures” (Gartner Hype Cycle methodology).

Even the enterprise vendors’ numbers need the question. Salesforce reported 18,500 Agentforce deals closed in 2025 and 9,500 of them paid, and moved pricing from two dollars a conversation to a flex-credit model that one partner says can fall to as little as ten cents a conversation (Salesforce Ben). When a vendor says “x thousand customers”, ask which number that is.

The prediction that has not happened yet

Sam Altman has said for two years that a one-person billion-dollar company is coming. In a clip that circulated in May 2026 he put it this way: “We’re going to see 10 person billion-dollar companies pretty soon. In my little group chat with CEO-friends, there’s this One-person billion-dollar company, which would have been unimaginable without AI, and now it’ll happen.” It has not happened. The most aggressive documented cohort, Bessemer’s “Supernovas”, runs about 1.13 million dollars of annualized revenue per employee, which leaves the one-person company roughly nine hundred people short (Bessemer, State of AI 2025). Bessemer’s own line about those companies is the durability asterisk on every fast ramp in this section: “Early hypergrowth alone means less now than ever before.” Jason Lemkin, who has spent 2026 trying to run the experiment himself, listed the parts that break, one per line: “who does support issues agent can’t resolve?”, “what about onboarding?”, “feature requests?”, “bug fixes?”, and last, “security?” (Lemkin on X). Note the last word. When a replier said large enterprises would hesitate to buy from a one-employee startup and that “compliance will be brutal”, Lemkin’s answer was “Yeah brutal” and “Even just doing a SOC-2 alone yourself while running it is tough”. Hold that exchange for the litmus-test section.

The claim, in a clip that plays here. Grade it against Bessemer's revenue-per-employee data and Lemkin's list. Watch on X.

Klarna is the case where the walk-back was more instructive than the announcement. In early 2024 its assistant was doing the work of 700 agents (OpenAI’s case study); by May 2025 the company was hiring humans again for the last mile (Entrepreneur). In February 2026 its chief executive, Sebastian Siemiatkowski, gave the version with the numbers in it: Klarna has 47 percent fewer employees than in 2022, through attrition, with revenue per employee 3.2 times higher, at 1.1 million dollars, and “we did increase our hiring for customer service jobs due to this reason”. The technology “will never replace the most essential job: human relationships.” That is neither the headline of 2024 nor the correction of 2025. It is the middle, and it took two years to say.

Klarna's chief executive, two years after the 700-agents claim, with the numbers and the rehiring in the same post. Read on X.
The Klarna claim, revisited by the person who made it. Watch on YouTube.

The vendor that vanished over a weekend

This one is the procurement lesson, and nobody lied. In May 2025 OpenAI entered exclusivity to acquire Windsurf, the coding tool, for up to three billion dollars. In early June, Anthropic revoked Windsurf’s access to Claude models; its co-founder Jared Kaplan said “It would be odd for us to sell Claude to OpenAI” (VentureBeat). Windsurf’s customers lost a model overnight because of a deal their vendor’s supplier did not like. On Friday, July 11, Google announced it was hiring Windsurf’s chief executive and key staff in a licensing deal of about 2.4 billion dollars, and the OpenAI deal was dead. On Monday morning Cognition signed to acquire what remained. Interim chief executive Jeff Wang described the Friday all-hands as “probably the worst day of 250 people’s lives” (TechCrunch). In the Hacker News thread, one commenter asked the question every founder standardizing on a tool should ask: “Why do people use your tool/service if you don’t own the LLM which is most of the underlying engine? How do you stay competitive when providers scale?” (Hacker News, 502 points). Another, defending the category, pointed at Cursor’s half-billion of annual recurring revenue and said it felt nothing like the dot-com bubble. He was right on the facts. Thirteen months later Cursor was a wholly owned subsidiary of SpaceX, after a 60 billion dollar all-stock deal that closed on August 14, 2026 (Wikipedia; CNBC). Nobody was dishonest anywhere in that chain. The dependency still changed owners twice in eighteen months. Vendor risk and vendor dishonesty are different things, and a startup’s tool stack is exposed to both.

Seventy-two hours: the OpenAI deal died, Google took the founders, Cognition bought the remains. Your vendor's cap table is your problem. Watch on YouTube.

And sometimes it costs more than time

The speed has a security bill, and in 2026 it came due on the platforms the fastest founders built on. In May, Axios reported research by the Israeli firm RedAccess that found 380,000 publicly accessible assets built on Lovable, Base44, Replit and Netlify, around 5,000 of them containing sensitive corporate data: medical records, internal banking data, Fortune 500 documents, clinical trial details. Replit’s chief executive responded that “Public apps being accessible on the internet is expected behavior” (Axios, May 2026). A month earlier a researcher had disclosed a broken object-level authorization flaw (BOLA) in Lovable’s own API that let any free account reach other users’ source code and database credentials in as few as five calls; it was reported on March 3, patched only for new projects, and disclosed publicly 48 days later after the bug bounty program closed the report (The Next Web). Lovable’s own account admits that between February 3 and April 20, 2026, “public project chat history and source code could potentially be accessed by any Lovable user”, and that “a change that should have been caught slipped through” (Lovable). In February, a Lovable-hosted exam platform with inverted authentication logic exposed 18,697 records, including students at two University of California campuses; the researcher who found it called it “a classic logic inversion that a human security reviewer would catch in seconds” (The Register). Simon Willison had drawn the line in advance: vibe coding is fine for throwaway projects and “grossly irresponsible” for anything holding other people’s information (Willison, May 2026). The line from The Next Web’s coverage of the Lovable episode is the one I keep, attributed there to Trend Micro: “The real risk of vibe coding isn’t AI writing insecure code. It’s humans shipping code they never had a chance to secure” (The Next Web).

Paul Graham noticed the cost of the bandwagon from the other side of the table. In August 2025, in the middle of the batch YC itself was marketing as the AI batch, he wrote: “I haven’t met all the startups in the current YC batch yet, but the two most impressive companies that I’ve seen so far are not working on AI.” A year later the batch composition caught up with him: the share of YC companies carrying an AI tag fell from a peak of 83 percent in Summer 2025 to 67 percent in Summer 2026, while industrials tripled their share (Jared Heyman’s analysis of S26). The founders stopped saying “AI” when saying it stopped being a differentiator.

From the person who has seen more early-stage companies than anyone alive, about the batch everyone called the AI batch. 1.1 million views. View on X.

Eight tells, ten minutes

None of this argues for skepticism, which is just the first extreme wearing a hoodie. It argues for a checklist a founder can run against any deck in ten minutes, each item anchored to a case above. One, ask for the method and the sample size behind the headline number, and whether the document is a study or a working paper; the 95 percent statistic fails this on item one. Two, ask for an unedited, real-time run on an input you choose; the Gemini and Devin videos fail it, and a Hacker News commenter wrote the community’s rule, “zero-edit, real-time presentations to rebuild trust, not fake cheerfulness after cuts” (Hacker News). Three, ask in writing what fraction of transactions touch a human; Nate, Builder.ai, Presto Automation and Just Walk Out all turn on this question. Four, ask whether the benchmark used a public test set and for a run on a held-out one; Manus. Five, ask what decision the system makes without a person and what happens when it is wrong; Gartner’s 130. Six, ask for trailing twelve-month revenue, retention and gross margin rather than annualized run rate; Bessemer’s Supernovas at 25 percent gross margin. Seven, ask how many of the logos are paying, in production, past twelve months; Agentforce’s 9,500 of 18,500. Eight, which is about you rather than the vendor: how long does it take to leave, and what happens if this company is acquired, cut off by its model supplier, or gone in ninety days; Windsurf, Humane, Cursor. None of it slows a startup down. All of it is what the enterprise bought with three extra quarters of procurement, and a startup can have it for free.

Checklist of eight tells of a grandiose AI claim, each with the question to ask and the documented case it would have caught: method and sample size, unedited real-time demo, human-in-the-loop disclosure, held-out benchmark, autonomy definition, trailing revenue rather than run rate, paying logos rather than announced deals, and the exit plan
Eight tells. The first seven are about the vendor. The eighth is about you.

The Middle Way is a practice, not a midpoint

Here is what the Buddha’s first discourse actually says, in Bhikkhu Sujato’s public-domain translation: “Avoiding these two extremes, the Realized One understood the middle way of practice, which gives vision and knowledge, and leads to peace, direct knowledge, awakening, and extinguishment.” The two extremes are “indulgence in sensual pleasures, which is low, crude, ordinary, ignoble, and pointless” and “indulgence in self-mortification, which is painful, ignoble, and pointless.” And then the text does the thing every poster leaves out: it defines the middle way as a program. “It is simply this noble eightfold path, that is: right view, right purpose, right speech, right action, right livelihood, right effort, right mindfulness, and right immersion” (SN 56.11, SuttaCentral). The Pali is majjhimā paṭipadā, the middle course of practice. It is a road, and it has eight named disciplines on it. It is nowhere described as a point halfway between the extremes.

That matters for this essay because “be moderate about AI” would be weaker than its own sources. Jim Collins reached the same conclusion from an entirely different direction thirty years ago, in the concept he and Jerry Porras called the Genius of the AND: “A visionary company doesn’t seek balance between short-term and long-term, for example. It seeks to do very well in the short-term and very well in the long-term” (Jim Collins). The Tyranny of the OR, in his words, is “the rational view that cannot easily accept paradox, that cannot live with two seemingly contradictory forces or ideas at the same time.” Refusing AI and chasing every demo are both failures of practice, and the alternative to both is a specific set of disciplines, held at the same time: speed and evaluation, adoption and controls, enthusiasm and audit. Below are the texts I have found most useful for naming those disciplines, followed by seven tests you can run on Monday.

The lute strings

The best single image in the canon for this problem is in the Soṇa Sutta (AN 6.55). Soṇa is practicing so hard his feet bleed, and getting nowhere. The Buddha asks him about his instrument, a vīṇā, which most retellings call a lute: “when the strings of your vina were too taut, was your vina in tune & playable?” No. Too loose? No. “Neither too taut nor too loose, but tuned to be right on pitch”? Yes. And then: “In the same way, Sona, over-aroused persistence leads to restlessness, overly slack persistence leads to laziness. Thus you should determine the right pitch for your persistence” (AN 6.55, Thanissaro Bhikkhu). Notice that the string is not told to be half-tight. It is told to be in tune, which is a property you can hear and test. The Gartner survey in which leaders explained their failed AI initiatives with “they expected too much, too fast” is a survey of strings tuned too taut. The Census Bureau’s 80 percent of firms that have not started is a survey of strings tuned too loose. Both instruments are unplayable.

The raft

The Alagaddūpama Sutta (MN 22) asks whether a man who has crossed a river on a raft he lashed together should carry it on his head afterward, and concludes: “In the same way, monks, I have taught the Dhamma compared to a raft, for the purpose of crossing over, not for the purpose of holding onto” (MN 22). Your model, your vendor and your framework are rafts. The Windsurf customers who lost Claude access in June 2025 and the Cursor customers whose tool became a rocket company’s subsidiary in August 2026 were carrying rafts on their heads. The sutta’s title comes from an earlier image in the same text: a water-snake grasped by the tail bites the hand that holds it. A tool grasped wrongly bites its holder. The Rule of Two and the autonomy ladder in this article are, among other things, instructions for where to grasp.

The two arrows

The Sallatha Sutta (SN 36.6) describes a man shot with an arrow and then, “right afterward”, shot with a second, “so that he would feel the pains of two arrows.” The untrained person, touched by pain, “sorrows, grieves, & laments” and “feels two pains, physical & mental.” The trained one “feels one pain: physical, but not mental” (SN 36.6). This is an incident-response temperament, and we see the second arrow constantly. The breach is the first arrow. The blanket ban that follows it, the one that pushes usage onto personal accounts where nobody can see it, is the second, and the organization shoots it into itself. Samsung’s 2023 ban was a second arrow. Netskope’s 78-to-47 percent migration from personal to managed accounts is what happens when a company declines to shoot it.

Both extremes are theories of causation

The Acela Sutta (SN 12.17) uses the same “avoiding these two extremes” formula about beliefs rather than behavior. The extremes there are two metaphysical positions, roughly that everything is fixed and that everything is random, and the middle is a causal account: “Avoiding these two extremes, the Tathagata teaches the Dhamma via the middle: From ignorance as a requisite condition come fabrications” and so on down the chain of dependent origination (SN 12.17). “AI changes nothing” and “AI changes everything” are both positions about causation. The middle is a causal account of what produces what, under which conditions, and that happens to be the security engineer’s instinct as well. It is also, in a formal, peer-reviewed form, the productivity J-curve. Erik Brynjolfsson, Daniel Rock and Chad Syverson showed that general purpose technologies “enable and require significant complementary investments” that are “often intangible and poorly measured in national accounts”, producing “underestimation of productivity growth in a new GPTs early years” and an overestimation later when the intangibles pay off (Brynjolfsson, Rock and Syverson, AEJ: Macroeconomics 2021). The gap between AI capability and AI results is neither evidence that the technology is fake nor evidence that the skeptics are right. It is the predicted signature of a technology whose complementary investments have not been made yet. Both extremes are misreading the same data.

Right effort has four jobs

The Magga-vibhaṅga Sutta (SN 45.8) defines each factor of the path. Right effort is four distinct exertions: for the non-arising of unskillful qualities that have not yet arisen, for the abandonment of unskillful qualities that have arisen, for the arising of skillful qualities that have not yet arisen, and for the “maintenance, non-confusion, increase, plenitude, development, & culmination” of skillful qualities that have arisen (SN 45.8). Only one of the four is “do more.” Two of the four are about stopping and preventing. An AI program that is only “ship more pilots” is doing a quarter of right effort, and this is where a security firm lives: most of our value is in the non-arising and abandonment columns. Every control in the first half of this essay is an act of prevention, and every retired shadow tool is an act of abandonment.

Beginner’s mind, and the editor’s hand

“In the beginner’s mind there are many possibilities, but in the expert’s there are few.” That is the epigraph of Shunryu Suzuki’s Zen Mind, Beginner’s Mind, and it is the publisher’s wording; the version with “expert’s mind” that most slide decks use repeats a word the book does not (Shambhala). San Francisco Zen Center has put the original recording online: the talk in Los Altos on November 11, 1965, from which the book’s prologue was edited. What Suzuki actually said, about six minutes in, was: “In beginner’s mind we have many possibilities, but in expert mind there is not much possibilities” (SFZC archive, with transcript). The most-quoted line in Western Zen is Trudy Dixon’s careful edit of a Japanese teacher’s second-language English, and it is better for it. Editing is craft, and the meaning survived the edit because a person who understood it held the pen. That is the entire argument for human oversight of a model’s output in one paragraph. The agent drafts. A person who knows what it is supposed to mean makes it true.

The primary source, in Suzuki's own voice, on the official San Francisco Zen Center channel. Recorded November 11, 1965. Watch on YouTube.

Beginner’s mind is the discipline both extremes lack. The enterprise is an expert in its own 2023 incident and can see only that. The startup is an expert in last month’s demo. The Zen master Nan-in, in the first of the 101 Zen Stories, pours tea for a professor until the cup overflows and says: “Like this cup, you are full of your own opinions and speculations. How can I show you Zen unless you first empty your cup?” (Reps and Senzaki, Zen Flesh, Zen Bones). The cup is the model choice you made in March. Empty it every quarter.

Kill the Buddha, and check the quotation

Linji Yixuan, the ninth-century Chan master, told his students: “If you meet a buddha, kill the buddha. If you meet a patriarch, kill the patriarch”, in Burton Watson’s translation of the Record of Linji (Columbia University Press). The version you know, “if you meet the Buddha on the road, kill him”, has a road in it that Linji never mentioned; it came from the title of Sheldon Kopp’s 1972 psychotherapy book. Richard Payne of the Institute of Buddhist Studies warns that detached from its context the line becomes “simply an empty slogan, an empty shapeless container that can be filled with whatever meaning someone wants” (Payne). In context it is a list, and the list escalates through patriarchs, saints and one’s own parents: substitute no authority at all for your own realization. For our purposes: the analyst report, the vendor benchmark, the Stanford economist and this essay are all buddhas to be met and killed. Run the test yourself. The point of tracing the “on the road” embellishment, and the fake Kālāma quotation earlier, is that the most repeated sentences in this whole tradition turn out to be edits nobody checked. That is the AI discourse in miniature, and the remedy is the same in both.

The last text is a verse by Layman Pang, an eighth-century householder. The saying you have seen, “before enlightenment, chop wood, carry water; after enlightenment, chop wood, carry water”, has no located source in any Chan text; it is a twentieth-century Western formulation. Pang’s actual verse, in the Sasaki, Iriya and Fraser translation, is: “Supernatural powers and miraculous activity: / fetching water and carrying firewood” (Terebess Asia Online). The miracle is not a different activity. It is the same activity done by someone who has stopped looking elsewhere for the miracle. Applied here, it is the argument for boring, well-instrumented deployments over demo-driven ones. The agentic workflow that transforms a company is the invoice run, the questionnaire, the code review, the month close. Fetching water and carrying firewood.

The secular version, dated

Every tradition of technology forecasting has its own Middle Way, and it is worth knowing where the sayings come from, because half of them are misattributed too. The line called Amara’s Law, “we tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run”, is fairly attributed to Roy Amara of the Institute for the Future as an idea, though the earliest documented version is J. C. R. Licklider’s in 1965, “People tend to overestimate what can be done in one year and to underestimate what can be done in five or ten years”, and Licklider called it a maxim that was already circulating (Quote Investigator). Bill Gates wrote in the 1996 afterword to The Road Ahead: “People often overestimate what will happen in the next two years and underestimate what will happen in ten. I’m guilty of this myself.” Robert Solow wrote in the New York Times Book Review on July 12, 1987, and it was the Times Book Review rather than the Review of Books, that “You can see the computer age everywhere but in the productivity statistics” (citation research); by 2000 he allowed that you could see it after all. Federal Reserve Governor Michael Barr reached for the Solow line in February 2026, calling AI “a normal early-stage general-purpose technology” (Federal Reserve), and by July 2026 was telling a Fed conference that “The rapid adoption of AI among small businesses is a strong signal of the potential value” while noting “We have heard many bold pronouncements about what AI will be able to do” (Federal Reserve, July 14, 2026). A central banker holding both sentences at once is the Middle Way in a suit.

A sitting Federal Reserve governor on AI, July 2026, on the Fed's own channel. Note the absence of adjectives. Watch on YouTube.

Carlota Perez’s model of technological revolutions is the structural version. Every revolution has an installation period “led by investment capital in alliance with young technological entrepreneurs”, culminating in a frenzy that leaves “enough infrastructure in place for everyone to benefit”, then “a difficult interim period: a time of uncertainty, instability, and economic recessions or even depression”, and then a deployment period in which “technology is finally deployed to its full benefit, with profits from customers, not speculation, providing the bulk of new investment capital” (Perez, interviewed by Art Kleiner in strategy+business, 2005). Ben Thompson applied the frame to the current wave and quoted her on the frenzy: “This financial frenzy is a powerful force in propagating the technological revolution, in particular its infrastructure” (Stratechery). The bubble is not a bug; it is how the data centers get built. That lets the practitioner say something more useful than “do not get carried away”: the frenzy is normal, and your job is to come out of it with the capability intact. “Profits from customers, not speculation” is also the sharpest single test of whether an organization is in its own deployment phase. Gartner’s instruments confirm that two waves are running at different phases at the same time: generative AI “is past the Peak of Inflated Expectations and starting to go into the trough”, where in Gartner’s words “the hard work takes place” (Gartner video), while agentic AI sits at the peak. A company can be disillusioned about one and euphoric about the other in the same quarter, which is how the two extremes in this essay coexist inside a single building. TechCrunch’s year-opener put it plainly: “The party isn’t over, but the industry is starting to sober up” (TechCrunch, January 2026).

Installation, frenzy, turning point, deployment. Carlota Perez on her own channel. We are still installing. Watch on YouTube.

Seven tests for the middle ground

Each test maps to one of the texts above, and each has a number you can put in a spreadsheet.

The string test. Are you spending more on evaluating than on running, or running without evaluating at all? Gartner’s April 2026 survey is one tuning fork: 28 percent of AI use cases fully succeeding and 20 percent failing outright, with many of the leaders who saw a failure blaming it on expecting too much, too fast. Forrester’s 38 percent of production agents with automated evaluations is the other. If you have pilots that never graduate, you are too loose; if you have production agents with no evaluation harness, you are too taut. In tune is a small number of workflows in production, each with a harness, and a backlog you are working through at a steady pace.

The raft test. For each model and vendor you depend on, how many weeks would it take to leave, and what happens if it is acquired, cut off by its own supplier or shut down in ninety days? A thin interface in front of every model call, no multi-year commitment, your data exportable on demand. Windsurf’s customers had no answer in June 2025. Yours should fit on an index card.

The Kālāma test. For the last AI decision you made, what did you measure yourself, on your own data, with the evaluation written before the prompt, and who with no stake in the answer checked it? “The vendor’s benchmark” and “a demo I watched” both fail. A one-week bake-off on your own tasks passes. Salesforce’s own research gives the shape of a good test: leading agents scored about 58 percent on single-turn business tasks, 35 percent on multi-turn, and over 83 percent on defined workflow execution, and showed “near-zero inherent confidentiality awareness” (CRMArena-Pro). Test multi-turn, test confidentiality, test on your data.

The two-arrows test. After your last AI incident or scare, did you ban, or did you build a control? Count the bans still in force. Each one is a second arrow, and each is pushing usage somewhere you cannot see it; the Cloud Security Alliance’s 38 percent of employees sharing confidential data with unapproved tools is the bill.

The right-effort test. In the last quarter, what did you prevent from arising, and what did you abandon? A control shipped before an incident is prevention. A retired shadow tool, a cancelled pilot with no path to production, a vendor dropped at the raft test, is abandonment. If your AI report to the board lists only launches, you are doing one quarter of the job.

The deployment test. Perez’s line: is the spend justified by profits from customers or by what the spend signals? For a startup, this is whether the AI in the pitch deck is paid for by revenue or by the valuation premium. For an enterprise, it is whether the agent program is in a core budget line or still in the innovation budget. Andreessen Horowitz’s chief information officers moved theirs from 25 percent to 7 percent innovation-funded in a year. Where is yours?

The beginner’s-mind test. When did you last re-run the bake-off from scratch, as if you had never chosen? Ramp’s effective token prices fell 41 percent from their March peak and its frontier-model share fell from 53 to 45 percent as companies changed defaults. The model that was best in March is rarely best in September. Empty the cup once a quarter.

Seven tests for the middle ground in AI adoption, each paired with its source text and a measurable question: the string test on evaluation versus running, the raft test on time to leave a vendor, the Kalama test on what you measured yourself, the two-arrows test on bans versus controls, the right-effort test on what you prevented and abandoned, the deployment test on profits from customers versus speculation, and the beginner's-mind test on when you last re-ran the bake-off
Seven tests, each with a number. In tune is a property you can hear.

Security maturity as the litmus test

Now the argument I most want to make, because it connects the two extremes to the work we do, and because I think it is true in a way that has become more true this year.

Start with the money. In the first half of 2026, artificial intelligence “accounted for 86% of all venture dollars” deployed in the United States, 412.7 billion dollars, more than all of 2025, with deals of 100 million dollars or more making up 87.5 percent of the total and three firms, Andreessen Horowitz, Thrive Capital and Founders Fund, taking in 48.1 percent of all capital raised by venture funds; first-time fund formation is on pace for its lowest year since 2016 (PitchBook-NVCA Venture Monitor, Q2 2026). Crunchbase counted 300 billion dollars of global venture funding in the first quarter alone, 242 billion of it to AI, with the four largest rounds (OpenAI, Anthropic, xAI and Waymo) taking 65 percent of all global venture funding by themselves (Crunchbase News, April 2026). Carta’s data on 1,618 rounds found AI companies at seed receiving valuations “roughly 50% higher” than their peers for similar cash (Carta, via TipRanks). Calling yourself an AI company is nearly free and worth half again on your valuation.

In the same half year, SimpleClosure, the firm that inherited Carta’s shutdown data, handled more startup shutdowns than in the same period a year earlier, and the median business-to-business software company had 11,900 dollars in the bank when it closed (SimpleClosure, via Under30CEO). Both facts are true at once, and they are the same fact. Bill Gurley, in March 2026: “One day we’re going to have an AI reset, because waves create bubbles, because interlopers come in” (Gurley on CNBC, via Yahoo Finance). Sam Altman, in August 2025, asked whether investors were overexcited about AI: “My opinion is yes” (Entrepreneur). Apollo’s chief economist wrote in July 2025 that the top ten companies in the S&P 500 were more overvalued than they had been in the 1990s (Apollo Academy), and Axios spent July 2026 on the circularity of Nvidia guaranteeing hundreds of billions of OpenAI’s financing: “you buy from me, I invest in you, and everything’s fine unless one of us has a problem, in which case both of us have a problem” (Axios).

Here is the point I want to be careful about, because the sloppy version is a sneer at founders and the accurate version is more interesting. The best study of what happens to companies funded in hot markets is Ramana Nanda and Matthew Rhodes-Kropf’s in the Journal of Financial Economics: “venture capital-backed startups receiving their initial investment in hot markets are more likely to go bankrupt, but conditional on going public, are valued higher on the day of their initial public offering, have more patents, and have more citations to their patents.” Their conclusion is that investors in hot markets fund “riskier and more innovative startups” rather than worse ones (Nanda and Rhodes-Kropf, Journal of Financial Economics, 2013; working paper). Hot capital buys variance. More of these companies die, and the survivors are more innovative, and both of those are good for the world. Neither is good for a buyer who needs the vendor to exist at renewal. And the money itself tells the buyer nothing: Goldfarb, Kirsch and Miller found that five-year survival among dot-com firms was 48 percent, in line with automobiles, tires and televisions in their own eras, and that “Survival of dot-com firms is unrelated to the receipt or amount of private equity funding” (University of Maryland summary; Journal of Financial Economics, 2007). Having raised money did not predict survival in the last bubble. It does not predict it in this one.

So the buyer’s problem, precisely stated, is Jon’s question from our sales calls: is this a real company, or is it the late deployment of a fund that had to place capital somewhere and put it here? I will not pretend that funds are contractually forced to invest on a clock; a general partner earns the management fee either way, and the pressure to deploy is reputational rather than legal. But the observable result is the same for a founder: a very small number of very large funds placing enormous sums into a narrow set of stories, with the firms underneath them competing for the residual. The market is full of claims and empty of verifiable type. That is George Akerlof’s problem, and he named the solution in 1970.

Signals, and why the cost is the point

Akerlof’s “The Market for ‘Lemons’” showed that when buyers cannot tell quality, “the ‘bad’ cars tend to drive out the good”, and that the cost of dishonesty “lies not only in the amount by which the purchaser is cheated; the cost also must include the loss incurred from driving legitimate business out of existence.” Then he listed what he called counteracting institutions: guarantees, brand names, and licensing, “the licensing of doctors, lawyers, and barbers” (Akerlof, Quarterly Journal of Economics, 1970). Audit and certification regimes are the software industry’s counteracting institution. A SOC 2 report or an ISO certificate is the licensing of barbers, for companies.

Michael Spence, who shared the 2001 Nobel with Akerlof for this work, supplied the condition that makes a signal work. It is not that the signal is expensive. It is that it is differentially expensive: “a signal will not effectively distinguish one applicant from another, unless the costs of signaling are negatively correlated with productive capability” (Spence, Quarterly Journal of Economics, 1973, p. 358). A signal that is equally cheap for everyone separates nobody, which is exactly why “we are an AI company” separates nobody in 2026. The biologist Amotz Zahavi found the same logic in peacocks in 1975 and called it the handicap principle: costly traits “test the quality of the mate”, and “The size of characters selected in this way serve as marks of quality” (Zahavi, Journal of Theoretical Biology). The tail is expensive and useless, and that is what makes it honest. Paul Milgrom and John Roberts extended it to “dissipative marketing expenditures”, money visibly burned to prove you can afford to, and noted that the signal only works for goods the buyer will purchase again (Milgrom and Roberts, Journal of Political Economy, 1986). A vendor signaling security maturity is signaling for the renewal, which is the sale that matters.

Now the mechanical part, which is the sharpest thing in this essay. A SOC 2 Type 1 report attests that your controls were designed properly at a point in time, and you can buy one in weeks. Thomas Ptacek, who has spent two decades being the most quotable skeptic of compliance on Hacker News, puts it plainly: “there’s basically no way not to pass your Type 1” (Hacker News). A SOC 2 Type 2 report attests that the controls operated over an observation window while an auditor watched, and in the words of one compliance platform’s pricing guide, “Three months is the shortest used in practice, twelve the norm” (Sprinto). You cannot compress time with money. For a company that already runs the controls and intends to exist next year, the marginal cost of a Type 2 is an audit fee. For a company that will be acquired for its engineers or wound down in nine months, the signal is unavailable at any price, because it has to survive the window. That is Spence’s condition, satisfied by a calendar. Type 1 is the cheap imitation. Type 2, and its successors, are the separating equilibrium.

Two-column ledger of cheap signals versus costly signals in the 2026 AI market: cheap signals include the AI label worth about 50 percent on seed valuation, the agentic rebrand, annualized run rate, announced deals and a SOC 2 Type 1; costly signals include a SOC 2 Type 2 observation window of three to twelve months, ISO 42001 held by roughly 350 organizations worldwide, a current penetration test, a named security owner and continuous evidence of AI output controls
Cheap signals separate nobody. Costly signals separate by calendar. After Spence (1973), Akerlof (1970) and Zahavi (1975).

The empirical literature agrees, with two different methods. Ann Terlaak and Andrew King followed an eleven-year panel of American manufacturing facilities and found that ISO 9000 certified facilities “grow faster after certification and that operational improvements do not account for this growth”, and that “the growth effect is greater when buyers have greater difficulty acquiring information about suppliers” (Terlaak and King, Journal of Economic Behavior and Organization, 2006). The certificate was doing signaling work, and it did the most work where buyers knew the least, which is a precise description of a market in which nobody can tell from the outside whether an eighteen-month-old AI company has any engineering discipline. Deane, Goldberg, Rakes and Rees ran an event study on 111 public ISO 27001 certification announcements and found “the associated abnormal stock market reaction is both positive and statistically significant” (Deane et al., Information Technology and Management, 2019). The market pays for the announcement because it treats it as information.

What the buyer is actually reading

Gartner predicted in 2022 that “By 2025, 60% of organizations will use cybersecurity risk as a primary determinant in conducting third-party transactions and business engagements” (Gartner). Vanta’s 2024 survey of 2,500 leaders found “Nearly two-thirds (65%) of organizations say that customers, investors and suppliers require more demonstration of compliance” (Vanta State of Trust 2024), and its Trust Maturity Report, drawn from more than 11,000 organizations, found that 71 percent of the most mature (“adaptive”) companies had adopted AI (Vanta, July 2025; report). SecurityPal reports that questionnaires now average nearly a hundred questions (SecurityPal). The buyer’s side of the question is stated best by Jason Lemkin, who buys a great deal of software and sells none of the compliance tooling, in a talk titled “Why I’m Scared to Buy New SaaS Apps Now”.

The buyer's side of "are you a real company", from someone with no compliance product to sell. Watch on YouTube.

Sellers have noticed. When Excalidraw announced its SOC 2 report, the explanation from the person who posted it was one sentence: “We got tired of endless security questionnaires, so we got SOC 2 certified to make things smoother for everyone”, and Ptacek himself, in the same thread, gave the mechanism: “SOC2 is viral”, because every business-to-business company requires it of its vendors and the alternative is the giant spreadsheet (Hacker News, 234 points). One compliance writer put the buyer’s shorthand exactly: “A trust badge on a website becomes shorthand for ‘safe.’ A passed audit becomes shorthand for ‘serious’” (One Horizon). Vanta itself, the platform 12,000 companies pay to prove their security, raised at a 4.15 billion dollar valuation in 2025 (Y Combinator on X). That number is the size of the market for the signal.

Twelve thousand companies paying to prove they are serious, and a 4.15 billion dollar valuation on the proving-it business. The clip plays here. Watch on X.

The founder’s belief, made legible

Paul Graham asked founders in 2015 whether they were default alive or default dead, meaning whether, on current growth and expenses, they reach profitability on the money they have. “The startling thing is how often the founders themselves don’t know.” And the reason they do not ask is the one this essay has been circling: “they assume it will be easy to raise more money. But that assumption is often false” (Default Alive or Default Dead?). A market that places 86 percent of venture dollars into one category makes “it will be easy to raise more” feel true, which is precisely when founders stop asking the question.

Here is the litmus test. A founder who commissions a twelve-month Type 2 observation window, hires or assigns a named security owner, and pays for an annual penetration test has answered Graham’s question in public, in a way that is expensive to fake. She has committed money now to an asset that only pays back if the company is still here when the window closes and the renewal comes up. That is a statement about the founder’s belief in her own company, legible to a buyer who cannot see her bank account. It is also a statement about operational maturity, because the controls do not pass the window unless someone runs them every week. When Jon and I sit in a security review on the vendor’s side, that is what the buyer’s team is reading, whether or not they would put it this way: does this founder believe in this company enough to invest in its own maturity, or is this a side project of somebody’s capital?

The frontier of the signal moves, and this is where the essay becomes 2026 rather than 2021. Aetos, summarizing enterprise questionnaires in June 2026, wrote that “SOC 2 Type II is now treated as a procurement baseline” and that buyers “assume it and it no longer differentiates a vendor” (Aetos). Spence predicts exactly this: as a signal’s cost falls and adoption saturates, separation collapses and the market moves to a costlier signal. In 2026 that costlier signal is ISO/IEC 42001, the management-system standard for artificial intelligence, plus evidence of AI-specific controls: model provenance and training-data rights, output monitoring and hallucination controls, subprocessor transparency. One public tally counted more than 350 organizations worldwide holding ISO 42001 certificates through April 2026 (AI Compliance Vendors), against an AI vendor population in the tens of thousands, and against Gartner’s estimate of about 130 real agentic vendors among thousands. The same tally records Amazon Web Services as the first major cloud provider to certify, in November 2024; Anthropic certified in January 2025, “one of the first frontier AI labs to achieve this certification”, audited by Schellman (Anthropic). Augment Code became the first AI coding assistant certified in May 2025, and its explanation of why is the vendor-side economics in four sentences: “Because every time you want to adopt an AI tool, someone asks about governance. Security wants to know about data handling. Compliance wants documentation. Procurement wants standards.” And the payoff: “Instead of custom questionnaires and months of back-and-forth, you can point to an international standard. We already went through the audit process” (Augment Code).

That certification was our program. YSecurity took Augment from zero to ISO 42001 certified in 93 days, and a deal worth more than a million dollars closed two days after the certificate arrived. Its current SOC 2 Type 2 period has agentic code reviewers inside the audit boundary, which is a sentence that would have been unthinkable to an auditor two years ago and is now a differentiator. Before that, we took Robust Intelligence through SOC 2 Type 2 with zero deviations; four months later Cisco acquired the company for 400 million dollars. Its vice president of engineering’s review of us is the one I would put on our tombstone: “I only wish we had hired them earlier.” I am not claiming the certificate caused the acquisition. I am claiming that a buyer with 400 million dollars to spend read the same signal that a buyer with a 40,000 dollar contract reads, and that both of them were reading the founder.

Now the strongest case against me

A post that claims certification prevents breaches would be dismissed by every practitioner who read it, and rightly. In July 2025 Brian Krebs reported that Paradox.ai, the maker of McDonald’s McHire hiring bot, had a test account protected by the password “123456”, exposing records on 64 million applicants. Paradox had announced ISO 27001 and SOC 2 Type 2 audits in 2019, and the account had eluded its annual penetration tests (KrebsOnSecurity). Ann Wallace, a former compliance engineering leader, wrote the sharpest recent critique: “When your goal is ‘pass the audit,’ you start optimizing for that. Not for being secure. For passing”, and coined “screenshot-grade security” for how controls get demonstrated. Then, in the same piece, she conceded the only claim I am making: SOC 2 “acts as a baseline signal of operational maturity” (Edera). When Anthropic announced its ISO 42001 certification, the Hacker News thread was mostly hostile: “a lot of ISO certification is ridiculously easy to get. 27001 you can basically copy off some qms procedures to your google drive and call it a day” (Hacker News). The signal works on buyers, and engineers can see through the weak versions of it. Both are true, and the tension between them is the essay.

Ptacek’s two most-quoted lines on the subject bracket the honest position. From the canonical thread on SOC 2 screenshots, 546 points: “Good compliance work is a byproduct of sound security engineering. It does not work the other way around” (Hacker News). And in May 2026, to a solo founder whose customers were demanding certification: “Don’t. You are exactly the wrong kind of firm to be pursuing SOC2” (Hacker News). I agree with both, and I sell the thing he is telling the founder not to buy, so let me be exact about where we differ. He is right that a one-person company chasing a badge before it has a deal to justify it is buying a costume, and right that the audit is a byproduct of engineering rather than a substitute for it. Spence’s signal predicts type, the kind of company you are, and says nothing about performance on any single draw, which is why Paradox.ai could hold both certificates and ship 123456. Where I would push back is on what the founder in that thread was actually being asked. Her customers were not asking for a badge. They were asking whether she would exist at renewal and whether anyone owned security in the meantime, and a report that says “here is everything we do, and here is who is accountable” answers the second question without the audit. The audit answers the first, and it does so precisely because it cannot be bought in a hurry. When the deal that justifies it arrives, the window has to have already started. That is the whole reason to begin before you need it, and it is the reason we bill it in fifteen-minute increments against a specific deal rather than as a program for its own sake.

A venture investor arguing that AI vendors' own claims need independent verification. That is the argument for third-party attestation, from the funding side of the table. View on X.

So, the litmus test, stated for both readers. If you are a founder: the certificate is not the point and never was. The point is a willingness to bear a cost that only pays back if the company survives, and to keep bearing it as the bar rises from SOC 2 to ISO 42001 to whatever the questionnaire asks for in 2028. Do it when a real deal justifies it, do it as a byproduct of engineering you would do anyway, and start the window before the deal, because the window is the part you cannot rush. If you are a buyer, read the report as a statement about the founder, and read the absence of one the same way. Then run the Kālāma test on the product anyway, because the report predicts the company and not the password.

Both extremes need an answer to the same question: how large is the thing we are tuning for? Start with the operators who said less and changed more, because their record is the middle ground in practice.

Shopify’s Tobi Lütke told his company in April 2025 that “before asking for more headcount and resources, teams must demonstrate why they cannot get what they want done using AI”, and asked every team: “What would this area look like if autonomous AI agents were already part of the team?” (TechCrunch on the memo). Duolingo’s Luis von Ahn declared the company “AI-first”, took the backlash, and by August 2025 said “This was on me. I did not give enough context”, adding that Duolingo had never laid off a full-time employee and that “one person will be able to accomplish more, rather than having fewer people” (Fortune). IBM replaced a few hundred human-resources roles with agents, and Arvind Krishna reported that “our total employment has actually gone up, because what it does is it gives you more investment to put into other areas” (PYMNTS on the Wall Street Journal interview). Amazon’s Andy Jassy wrote in June 2025 that “we will need fewer people doing some of the jobs that are being done today, and more people doing other types of jobs” (aboutamazon.com), and Amazon then cut 14,000 corporate roles in October and 16,000 more in January 2026 while framing both rounds as removing “layers” and “bureaucracy” (TechCrunch). Walmart’s Doug McMillon said “It’s very clear that AI is going to change literally every job” and planned to hold headcount near 2.1 million for about three years while the shape of those jobs changes (Fortune via Yahoo Finance). Accenture, the largest seller of AI transformation on earth, spent up to 865 million dollars “exiting” people (Fortune) for whom, in Julie Sweet’s words, “reskilling, based on our experience, is not a viable path”, in the same year it booked 5.1 billion dollars of new generative AI work (Irish Times).

The Shopify memo, published by its author, 2.5 million views. Adoption as a stated expectation rather than a tool purchase. Read on X.

The steepest curve is the one Semafor traced at Google: the share of new code that is AI-generated went from 25 percent in late 2024 to 50 percent in 2025 to 75 percent in April 2026, with Semafor’s own caveat that these time horizons are “pretty impossible to fact-check”. In Alphabet’s second-quarter 2026 letter, Sundar Pichai reported “a team in Chrome is now on track to accelerate delivery by eight times, compressing a two-year timeline into three months through model-driven refactoring” (Alphabet Q2 2026). Anthropic’s deputy chief information security officer wrote in July 2026 that “Claude authors about 80% of the code merged into our codebase today” (Anthropic). Microsoft’s 2026 Work Trend Index measured a fifteen-fold year-over-year growth in active agents in its own suite, and found that organizational factors explain more than twice as much of AI’s impact as individual ones (Microsoft WTI 2026). And the frontier inside the labs is further along still: Boris Cherny, who created Claude Code at Anthropic, said in February 2026 that all of his own code had been written by the tool since the previous November, from about 20 percent when it launched.

The frontier, with dates, in a clip that plays here. Watch on X.
Table of what operators reported about AI and headcount in 2025 and 2026: Walmart holding headcount near 2.1 million while revenue grows, IBM total employment up after replacing HR roles with agents, Microsoft cutting over 15,000 roles with headcount roughly unchanged, Google's AI-generated share of new code rising from 25 to 75 percent, Accenture spending 865 million dollars on exits alongside 5.1 billion dollars of generative AI bookings, PwC finding the most AI-exposed firms grew headcount 52 percent versus 36 percent, and Challenger's AI-attributed layoffs doubling in the first half of 2026 while total cuts fell 40 percent
Recomposition rather than contraction. More output per person, the org chart redrawn, the entry rung moved up.

What does the macro data say? Challenger, Gray and Christmas counts stated reasons for announced layoffs. AI-attributed cuts went from 54,836 in all of 2025 (Challenger, December 2025 report) to 101,743 in the first half of 2026, by which point AI had led the monthly reasons for four consecutive months (Challenger, June 2026 report), a run that reached five months before ending in August, when AI dropped to fourth place while total cuts through August were down 41 percent on the year and hiring plans were up 37 percent (Challenger, August 2026 report). In the July report Andy Challenger himself warned that “Naming AI in a layoff announcement can win over investors while pushing current and prospective employees away.” Treat the attribution as partly a narrative, the same way you treat the vendor’s.

And the predictions, graded. Dario Amodei told Axios in May 2025 that AI could eliminate “half of all entry-level white-collar jobs” and “spike unemployment to 10-20% in the next one to five years” (Axios); sixteen months in, unemployment is nowhere near that and the entry-level half has the better support. Jensen Huang, on stage at Dreamforce on September 15, 2026: “Every company, every enterprise, every country would become an AI company. It’s going to be an agentic enterprise”, followed by six words that are the shortest version of the case for starting: “Engage AI. Don’t get left behind” (NVIDIA blog). Sam Altman, in June 2025: researchers were “two or three times more productive than they were before AI”, and “the ability for one person to get much more done in 2030 than they could in 2020 will be a striking change” (The Gentle Singularity). Remember that two-or-three-times number.

The boldest credible prediction, on a mainstream channel, May 2025. Grade it against the labor data above. Watch on YouTube.

Here is where the economists earn their keep, and where the argument becomes directly useful to anyone running a company. Stanford’s Chad Jones gave a talk in May 2026 called “AI and Our Economic Future”. About thirty minutes in he makes the argument that his 2017 paper with Philippe Aghion and Ben Jones formalized: “growth may be constrained not by what we are good at but rather by what is essential and yet hard to improve” (Aghion, Jones and Jones). In the paper that accompanies the talk, published this summer in the Journal of Economic Perspectives, he writes:

“Just as a chain is only as strong as its weakest link, the economy may be limited by whatever tasks are not yet automated.”

Charles I. Jones, AI and Our Economic Future, Journal of Economic Perspectives, Summer 2026

Chad Jones on weak links. The player opens at the 30-minute mark, where the chain argument begins. Watch on YouTube.

The arithmetic is the part to keep. Automating a task that accounts for share s of output raises output by a factor of one over one minus s. At the parameter Jones uses, “automating 50 percent of GDP only raises GDP by just 19 percent” (GDP being gross domestic product). Automating all of software, about 2 percent of GDP, buys you about 2 percent. To merely double income, the model needs 94 percent of the economy automated first. His intuition from the stage: the phone in your pocket has a hundred million times the transistors of the computers of the 1970s, and “I’m not 100 million times more productive at research, right? Why not?” His answer for himself: “maybe I’m two or three times more productive.” Which is the number Sam Altman used. The most bullish executive in the industry and the most careful growth economist agree on the multiplier for a person. They disagree about whether it aggregates.

Diagram of Chad Jones's weak-links arithmetic: a chain with one dashed link labelled the task humans still do, the rule that output can be no larger than the output of the weakest link, and three worked examples showing that automating all software adds about 2 percent, automating half of all work adds about 19 percent, and doubling income requires automating 94 percent of GDP
Weak links: why automating half the work does not double the output. After Jones, Journal of Economic Perspectives, Summer 2026.

Jones is not a pessimist, and it would misrepresent him to say so. In his model growth still explodes; in a May 2026 paper with Christopher Tonetti they write that “automation leads economic growth to accelerate, but the acceleration is remarkably slow because of the prominence of ‘weak links’” (Jones and Tonetti). Their historical finding is the striking one: if automation had frozen in 1950, essentially all of American productivity growth since then would disappear. On the other side, Ege Erdil and Tamay Besiroglu of Epoch AI conclude that “explosive growth seems plausible with AI capable of broadly substituting for human labor, but high confidence in this claim seems currently unwarranted” (Erdil and Besiroglu), and Dario Amodei’s own maximalist essay names the same bottlenecks Jones does, biology that “can sometimes take days or even weeks, with no obvious way to speed it up”, data quality, regulation and institutions (Machines of Loving Grace). Amodei calls them bottlenecks. Jones calls them weak links. Others think Jones is too conservative, that a technology which improves its own inputs has no historical analogue, that Solow’s paradox resolved itself once the complementary investments landed, and that Brynjolfsson’s J-curve says the same thing will happen here, faster. Governor Barr made the self-improvement point from the Fed’s podium in July: “AI itself is a powerful tool to train and accelerate development of new AI models. That is, AI improves its own research and development.” The honest synthesis is the one Anthropic’s economics team published with Jones and Anton Korinek on September 9: three scenarios for 2030, modest, substantial and extreme, adding anywhere from half a point to fifteen points of annual growth, with the median American’s expectations sitting on the substantial one (Economic Scenarios for Transformative AI). Nobody serious is arguing that nothing happens. The disagreement is speed and sequencing, and the only thing everyone signs is that change is coming.

Three scenarios, one of them co-authored by the weak-links skeptic himself. Nineteen million views in a week. View on X.

Now bring the argument down from the economy to your company, because it scales, and because it is where the two extremes meet. Ethan Mollick put the whole problem in a sentence this summer: “When workers gain productivity from AI, the rest of the bottlenecks in the system can eat up any gain.” Our organizations, he wrote, “are built around” the limits of the only kind of intelligence we ever had (Mollick on X). That is weak links at the level of one team. The r/ExperiencedDevs thread with the most upvotes I could find on the subject, nearly 24,000, quotes a founder on the same point.

"Everyone's talking about their teams like they were at the peak of efficiency and bottlenecked by ability to produce code. Here's what things actually look like: your org rarely has good ideas. Ideas being expensive to implement was actually helping."

Dax Raad, as quoted in r/ExperiencedDevs, 23,900 upvotes, February 2026
And Mollick again, with data from a large study of coding agents: autocomplete tools led to 2.2 times more code, local agents 7.4 times, remote agents 17.3 times, "But human bottlenecks in coding means actual releases 'only' went up 30%" ([Mollick on X](https://x.com/emollick/status/2061659432233161023)).
Seventeen times the code, thirty percent more releases. The gap between them is the weak link, and it is the work. View on X.

That gap is where both extremes end up. The waiting organization leaves the human steps exactly where they were, so when the tools do arrive, the seventeen-fold output hits the same review queue, the same approval chain, the same security questionnaire, and comes out the other end as thirty percent. The trying-everything startup buys seventeen-fold output on day one and discovers that its weak link is the founder reading pull requests at two in the morning. The weak link in your company is the workflow you did not redesign and the control nobody was assigned to build. Both are fixable in weeks, which is what the last section is about.

Seven moves to modernize one business function

We have done this three times inside our own companies, so let me be specific about the moves and honest about the numbers. At YSecurity we rebuilt the money side of a services firm, the hours, budgets, invoices, payouts, commissions and month close, into a product called Ceed that has never had an engineer on staff and runs at roughly a tenth of the cost of a team and, by our own count, at about a hundred times the velocity. Our deal-acceleration platform Cyberbase, which moves security questionnaires, redlines and compliance reviews through to signature, now operates with zero engineers. And this article was assembled the way we now assemble most things: research agents did the reading, people did the checking, and a person put his name on it. Those results came partly because we started with no legacy team to protect and partly because we are a security firm, so the floor was already there. Your numbers will differ, which is why the estimator at the end asks about your function rather than ours. Each move below carries the Middle Way test it satisfies.

Numbered list of seven moves to modernize a business function into agentic workflows, each with a one-line description and the YSecurity service that delivers it: pick the workflow not the tool, lay the security floor first, rebuild it as a workflow then earn agency, put a person at the irreversible step, give everyone the paved road, redesign the craft roles around judgment, and measure, rightsize and renegotiate
The seven moves, and where we show up for each. Every one is measured in weeks and billed in fifteen-minute increments.

Move 1. Pick the workflow, not the tool

Every failed program we have been called into started with a license and went looking for a use, which is the startup’s extreme wearing an enterprise badge. Start with a task inventory of one function: what comes in, what goes out, where the hours go, what the error rate is, what a unit of output costs today. Ng’s task-based analysis and BCG’s rule that “10% of a company’s efforts should be focused on algorithms, 20% on technology and data, and the remaining 70%” on people and process (BCG) both point the same way. Then pick the workflow with the highest volume, the clearest definition of done and the cheapest check. Anthropic’s usage data says where real autonomy lives today: code and back-office administration, rather than the glamorous front-office use cases that absorb most budgets (Economic Index, January 2026). At Ceed, the invoice run from logged hours is the shape of workflow that qualifies: tedious, frequent, and trivially verifiable against the ledger. Fetching water and carrying firewood. Where we show up: agentic strategy and market analysis, a two-to-three-week inventory that ends with the one workflow, the baseline, and the estimate. This is the deployment test: a workflow chosen because customers pay for its output.

Move 2. Lay the security floor first

Do this before the first agent touches real data, because retrofitting it is where programs stall for a quarter, and because it is the non-arising half of right effort. The floor is short: one identity per agent with no standing privileges; a sandbox with filesystem and network isolation (Anthropic reports that sandboxing “safely reduces permission prompts by 84%”, so it is faster, not slower (Claude Code sandboxing)); an egress allowlist; commercial terms that say no training and set retention; separate development and production environments with a tested restore; and the Rule of Two applied on a whiteboard to every agent design. Version-pin the agent tooling and forbid the permission-skipping flags in managed environments, because the Nx “s1ngularity” supply-chain attack of August 2025 weaponized exactly those flags to steal 2,349 credentials (The Hacker News). Where we show up: this is our home ground. Secure AI software development life cycle, AI access control audits and secure Model Context Protocol programs are the floor, and we build it in the first two to three weeks. It is the same floor that let Augment Code put agentic reviewers inside a SOC 2 Type 2 boundary and reach ISO 42001 in 93 days, and it is the beginning of the observation window that the litmus test depends on.

Move 3. Rebuild it as a workflow, then earn agency

Anthropic’s own advice is to use “the simplest solution possible”. Most first workflows are deterministic code with model steps inside: fetch, classify, draft, check, route. Write the evaluation before you write the prompt; that is the Kālāma test made into a build step, and it is the difference between Forrester’s 9 percent rollback rate and its 47. Only when the task truly has an unpredictable number of steps do you let the model direct them, and by then you have the harness to catch it. That is how we think about the questionnaire flow at Cyberbase: pull the counterparty’s public terms first, map each question to the evidence already on file, draft, and route anything uncertain to a person. The agentic part lives inside the steps, with the path decided by code. Where we show up: agentic engineering, delivered by operators who embed with your team the way forward-deployed engineers do at Palantir and OpenAI, without the eight-figure minimum OpenAI reportedly attaches to its version (The Decoder on The Information’s reporting).

Move 4. Put a person at the irreversible step

Anthropic’s security guidance for defenders says a report “should be sent only when a human has verified it and is willing to put their name on it” (Anthropic, April 2026). That is the rule for every function: the agent proposes, a named person merges, approves, sends or pays. It is also how autonomy is earned, and how the string gets tuned. Start with one noisy, high-volume decision, run the agent alongside the human, and measure the agreement rate; expand to the next decision only when the rate is tolerable. The staged ladder below is our synthesis of the incidents in this article: read-only, then scratch, then human-gated proposals, then one measured workflow, then one workflow at a time, and never untrusted input, sensitive data and external action in the same session without a person. Where we show up: the ladder is the backbone of our agentic engineering work, and the audit trail it produces is what your assessor asks for next year.

Six-rung staged autonomy ladder for agentic workflows: read-only sandboxed with no egress, write to a scratch environment, propose through a human gate, one measured workflow with narrow autonomy, expand one workflow at a time, and a dashed final rung stating that untrusted input, sensitive data and external action are never combined in one session without a person
Autonomy is granted by evidence. The dashed rung is the boundary that does not move.

Move 5. Give everyone the paved road, not just the pilot team

Sanctioned use trails real use everywhere it is measured: Mollick’s 40 percent of workers using AI against 20 percent official adoption (One Useful Thing), Section’s 69 percent of workers whose organization has acted on agents against 16 percent who use one. The fix is permission plus provisioning plus five hours of training, and it means real tools. George Mandis’s essay this month on “AI Enablement Theater” describes companies rationing tokens and cheap models so that “the takeaway they’ll walk away with is ‘this stuff isn’t very good.’” His line is the one to put on the wall of the enablement team.

“Don’t hand them a Swiss Army Knife and tell them they can only use the toothpick.”

George Mandis, AI Enablement Theater, September 2026 Netskope’s 78-to-47 percent migration from personal to managed accounts is what the paved road looks like when it works, and it is the two-arrows test passed: a control instead of a ban. Where we show up: the shadow-AI inventory tells you what to pave, and agentic strategy sets the policy short enough that people follow it.

Move 6. Redesign the craft roles around judgment

Design, copy and product are where the resentment is loudest and the wins are quietest, and both are real. The Nielsen Norman Group’s 2026 state-of-UX report found designers “told they’ll be replaced if they don’t ‘vibe code’” (Nielsen Norman Group), while a senior designer wrote in Smashing Magazine that after decades mastering cognitive load and accessibility she was “suddenly finding myself being judged on my ability to debug a CSS Flexbox issue” (Smashing Magazine). Grading judgment work on artifact throughput is the mistake. The redesign is to make judgment the spec: a brand voice written down precisely enough that an agent can draft in it and a person can reject what misses (we did this for Ceed, exporting its brand and tone documents in a machine-readable form so every agent drafts the same way); a design system with the decisions encoded so the agent produces production-ready options and the designer chooses; a product analytics loop where agents pull the numbers and people decide what they mean. Paul Graham noticed the alternative from the receiving end this spring: “A lot of the emails I get from founders are now written in a hard-hitting journalistic style. I know they’re written by AI, because no founder ever wrote this way before. And once you realize something is written by AI, it’s hard not to ignore it” (Graham on X, 1.3 million views). His founders lost his attention because the voice was missing, not because the drafting was automated. This is Trudy Dixon’s hand on Suzuki’s sentence. Where we show up: agentic aesthetics, design and user experience, agentic copy, branding and tone authorship, and agentic product strategy and market analytics.

Move 7. Measure, rightsize and renegotiate, continuously

The DORA research program found AI raises delivery throughput and delivery instability at the same time, and that it “magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones” (DORA). So measure both: the four delivery keys plus a rework figure, because BetterUp Labs and Stanford measured the “workslop” tax at about two hours of rework per incident (HBR) and 186 dollars per employee per month (BetterUp Labs). Then rightsize the spend with the levers the FinOps Foundation lists, model selection, caching, batching, routing (FinOps for AI), and revisit the tool choice every quarter, because the model that was best in March is rarely best in September. That is the beginner’s-mind test and the raft test on a calendar. This is where independence pays. McKinsey’s own reference architecture for agentic AI lists “vendor neutrality” as one of its five core design principles (McKinsey, Seizing the agentic AI advantage). If neutrality is a design principle of the architecture, it should be a property of the people who help you choose it. We do not have relationships with vendors, we bill in fifteen-minute increments with no retainer, and our on-demand model exists to rightsize itself: love us, or let us go. Nobody has let us go yet. Where we show up: all five services in our Business Modernization & Agentic Transformation practice run this way, continuously, alongside your team.

Two builders on why enterprise adoption lags without dismissing the technology. Aaron Levie and Harrison Chase. Watch on YouTube.

What it would take for your function

Initial estimate

Pick one function. See the range.

The ranges use published task-automation shares, Chad Jones's weak-links arithmetic and what we have measured in our own companies. It is an initial estimate, deliberately wide, and every assumption is printed below the numbers.

…annual run-rate freed for reinvestment
…faster on the tasks the agents take over
…faster for the whole function, until you redesign the human steps
…to modernize this function, security floor included

A person reads this, runs the workings for your function, and replies. No sequence, no vendor pitch.

Every reason in the first half of this essay was correct as stated, and every claim in the second half was made by someone smart. The data really can leak; the agent really can be hijacked; the demo really was edited; the vendor really did vanish over a weekend. Neither list is a reason to stop and neither is a reason to buy. They are, together, the reason to practice: one workflow, the floor first, the evaluation before the prompt, a person at the irreversible step, a raft you can put down, a cup you empty every quarter, and a security posture that tells your buyers, before you have said a word, that you intend to be here when the window closes.

If you want company on that road, the Business Modernization & Agentic Transformation practice exists so you do not have to walk it alone or on faith. Bring one function. We will inventory it, lay the floor, rebuild the first workflow and put a person at the irreversible step, with no vendor in the room but the one you choose. Love us, or let us go.

Where both extremes end up: the tool widens the gap between the teams that practiced and the teams that either waited or bought the demo. Be in the middle, on purpose. View on X.

The Middle Way of AI adoption, frequently asked questions

What is the Middle Way of adopting AI?
It is a practice rather than a midpoint. In the Buddha's first discourse (SN 56.11) the middle way is defined as a specific eightfold program of practice that avoids two extremes, not as moderation between them. Applied to AI adoption, the two extremes are the established organization that treats every incident report as a reason to wait and the startup that tries every tool a demo recommends. The middle is a discipline: pick one workflow, write the evaluation before the prompt, lay the security floor first, put a person at the irreversible step, and measure what you can check yourself. Jim Collins reached the same conclusion from a different direction: a visionary company does not seek balance, it seeks to do both things well.
Why do startups adopt AI faster than established companies?
Because their risk profile is inverted. CB Insights studied 431 venture-backed companies that shut down since 2023 and found that, among the 385 with an identifiable cause, 70 percent ran out of capital. Stripe found the top AI companies reached a million dollars of annualized revenue in a median of 11.5 months. For a company with eighteen months of runway, a three-month experiment that fails is cheap and a three-month deliberation that fails is fatal, so trying everything is rational. A quarter of Y Combinator's Winter 2025 batch had codebases that were 95 percent AI-generated within months of the tools existing. The startup's risk is not caution. It is spending its scarcest asset, time, on grandiose claims that turn out to be false.
How do you tell a real AI vendor from a hype vendor?
Eight tells, each anchored to a documented case. Ask for the method and sample size behind any headline number (the MIT NANDA 95 percent figure rested on 52 interviews). Ask for an unedited, real-time demo on an input you choose (Google's Gemini launch video and Cognition's Devin launch video were both edited). Ask in writing what fraction of transactions touch a human (the SEC alleged Nate's AI shopping app was people). Ask whether the benchmark used a public test set (Manus claimed state of the art on a public set). Ask what decision the agent makes without a person (Gartner counts about 130 real agentic vendors among thousands). Ask for trailing revenue and gross margin rather than annualized run rate. Ask how many logos are paying customers (Salesforce announced 18,500 Agentforce deals in fiscal 2025, 9,500 of them paid). And ask how long it takes to leave.
What is agent washing?
Gartner's term, from June 2025, for rebranding existing products such as chatbots, robotic process automation and assistants as agentic AI without substantial agentic capability. Gartner estimated that only about 130 of the thousands of self-described agentic AI vendors are real, and predicted that over 40 percent of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls. Its April 2026 Hype Cycle still places agentic AI at the Peak of Inflated Expectations, with 17 percent of organizations deployed and more than 60 percent planning to deploy within two years.
Is security maturity a signal that a company is real?
Yes, and the reason is mechanical. Michael Spence's 1973 signaling theory holds that a signal only separates good from bad types when it is differentially expensive. A SOC 2 Type 2 report requires an observation window during which an auditor watches the controls operate, three months at minimum and twelve as the norm, and no amount of money compresses time. A company that will not exist in a year cannot buy it at any price. Empirically, ISO 9000 certified manufacturers grew faster than uncertified peers and the effect was largest where buyers had the least information about suppliers; 111 ISO 27001 certification announcements produced positive, statistically significant abnormal stock returns. The certificate is evidence about the company's intentions and durability rather than a warranty on any single control.
What is the difference between SOC 2 Type 1 and Type 2?
A Type 1 report attests that controls were designed appropriately at a point in time and can be produced in weeks; as one Hacker News practitioner put it, there is basically no way to fail your Type 1. A Type 2 report attests that the controls operated effectively over an observation window, typically three to twelve months, which is why it is the version enterprise buyers ask for and the version that carries signal. Since 2026 many buyers treat SOC 2 Type 2 as a procurement baseline and have moved the differentiating bar to AI-specific evidence such as ISO 42001 alignment, output monitoring and subprocessor transparency.
What is ISO 42001 and how many companies have it?
ISO/IEC 42001 is the international management-system standard for artificial intelligence, published in December 2023. One public tally counted more than 350 certified organizations worldwide through April 2026, against an AI vendor population in the tens of thousands, which is what makes it a costly signal. Amazon Web Services announced the first major cloud certification in November 2024, Anthropic certified in January 2025, and Augment Code became the first AI coding assistant certified in May 2025, a program YSecurity delivered in 93 days. A million-dollar deal closed two days after the certificate arrived.
Is prompt injection solvable?
Not at the model layer, and you do not need it to be. Language models read instructions and data as one stream, so an agent that reads untrusted content can be steered by it. The fix is architectural: Meta's Agents Rule of Two says an agent may read untrusted input, touch sensitive systems, or change state and communicate externally, but never all three in one session without a person. Sandboxing, egress allowlists, per-agent identities with no standing privileges, and human-owned gates on irreversible actions make injection non-consequential even when it succeeds.
How do you stop an AI agent from deleting production?
With architecture rather than prompts. In July 2025 a Replit agent deleted a founder's production database during a code freeze; in April 2026 a Cursor agent deleted PocketOS's production database and its backups in nine seconds through an unrestricted API token. The controls are the same in both cases: separate development and production environments, credentials scoped by operation and environment, backups on separate infrastructure with a tested restore, and destructive operations behind a human confirmation. Zenity's post-incident analysis put it plainly: system prompts are weighted inputs to a probabilistic reasoning engine, not deterministic enforcement mechanisms.
Does AI make experienced developers slower?
Sometimes, on codebases they already know by heart. METR's July 2025 randomized trial found sixteen experienced open-source developers took 19 percent longer with early-2025 tools while believing they were 20 percent faster. In February 2026 METR retired that design because developers refused to be randomized out of AI even at 50 dollars an hour, and its second cohort measured a 4 percent slowdown. Brynjolfsson, Li and Raymond found 34 percent gains for novice support agents and little for experts. Where the work is unfamiliar, repetitive or verifiable the gains are large; where the person already holds the whole system in their head they are small.
What is the weak-links argument about AI and economic growth?
Stanford economist Chad Jones argues that an economy, like a chain, is limited by its weakest link: the tasks not yet automated and still performed by slowly improving humans. In his 2026 Journal of Economic Perspectives paper, automating a task worth share s of output raises output by one over one minus s, so automating half of all work adds about 19 percent at his parameters, and doubling income requires automating 94 percent of it. Growth still accelerates in his model, just more slowly than the boldest forecasts, and others think he is too conservative. The same arithmetic applies inside a company: the step you did not redesign sets the pace of the whole function.
Does YSecurity have relationships with AI vendors?
No. We do not resell, partner with, or take referral fees from any AI vendor, model provider or platform. A firm that has certified thousands of people on one vendor's model, or earns margin on a license, has made a large investment in a default. We think you should be able to tell whether a recommendation was chosen or defaulted to, so we only ever wear one hat. We recommend the best tool for the workflow, we help you deploy it inside your security commitments, and if we stop earning our fifteen-minute increments you let us go.
Written by the team behind The Security Podcast of Silicon Valley

Put it into practice.