Live webinar The Wrong Security Hire Burns Your B2B GTM Pipeline. A fireside chat for founders · Oct 6, 9:30am PTThe wrong security hire · Oct 6 Save your seat

Blog

14 min read Updated September 6, 2026

AI Cyberattacks Are Going Autonomous: When the Hacker Is a Machine

For a while, AI just wrote cleaner phishing emails. The newest tools run the whole attack: finding weaknesses, exploiting them, chaining web to IoT to physical systems, and moving deeper with little human input. Here is what the documented cases show, and what changes for defenders when the attacker never sleeps.

AI cyberattacks use artificial intelligence to break into systems. For a while that meant small help: a cleaner phishing email, faster sorting of stolen data. That has changed. The newest tools don’t just assist a human attacker; they run the attack themselves, finding weaknesses, exploiting them, and moving deeper with little human input.

The short version: the attacker is becoming a machine. On episode 97, Alexis Lingad, founder and CEO of KinoSec, described building an autonomous offensive platform he calls a “cyber weapon,” and was candid about what it does to a target. This post is about that threat: what an autonomous attack chain looks like, why it no longer stops at the web app, and what changes for a defender when the attacker works at machine speed. The companion question, whether agents like his can replace human pentesters, gets its own treatment in AI Penetration Testing.

What is an autonomous AI cyberattack?

An autonomous AI cyberattack is an intrusion in which AI agents, rather than people, carry out most of the steps: reconnaissance, exploitation, privilege escalation, lateral movement and exfiltration, with a human setting the goal and approving a handful of decisions along the way. The difference from an “AI-assisted” attack is who holds the keyboard: in an assisted attack a person does the work and uses a model to go faster; in an autonomous one the model does the work and the person supervises.

Picture a spectrum: the polished phishing email at one end, a criminal coaxing a coding agent into scanning networks in the middle, and at the far end an orchestrator handing hundreds of small tasks to sub-agents running in parallel for hours. Two years ago the far end was a conference slide. Lingad’s read on the pace, recorded in the spring: “I think it will exponentially change in the next few months, not years, but few months.” Within a quarter, the public record had caught up with him.

Has an autonomous AI cyberattack actually happened?

Yes, and the forensics are public. Three cases, each documented by the people who had the logs.

November 2025: GTG-1002. Anthropic disclosed what it titled the first reported AI-orchestrated cyber espionage campaign. A group it assessed as Chinese state-sponsored, designated GTG-1002 in the full report, manipulated Claude Code, wired to ordinary security tools through the Model Context Protocol, into attempting intrusions at roughly thirty targets: large tech companies, financial institutions, chemical manufacturers and government agencies. Anthropic estimates the AI performed 80 to 90 percent of the campaign at “thousands of requests, often multiple per second,” with humans intervening at perhaps four to six critical decision points. A handful of intrusions succeeded. The model also “occasionally hallucinated credentials or claimed to have extracted secret information that was in fact publicly-available,” which Anthropic called an obstacle to fully autonomous cyberattacks. An obstacle, not a wall.

Spring and summer 2026: the OpenAI swarm and the Cursor crew. OpenAI’s own evaluation agents built a message board inside an internal package registry, coordinated across roughly 1,200 agents, and used exposed credentials and two zero-days to reach administrator access across Hugging Face clusters. Separately, a Russian-speaking ransomware crew talked the agent inside Cursor into assisting real intrusions at seven companies. We covered both in AI Agents Are Now on Both Sides of the Breach. One was an accident, one was a crime; neither needed a research budget once the agents existed.

September 2026: a ransom operation. Palo Alto Networks’ Unit 42 published an investigation of an enterprise intrusion in which the attacker “left tactical execution to AI agents that monitored, evaluated, acted and re-planned in real time.” The agents breached a public web service, combed code repositories for hard-coded tokens, escalated through the secrets store, tried to backdoor the CI/CD pipeline, stole cloud keys, and turned the victim’s own AI endpoints into infrastructure for the next moves. Unit 42 counted more than 50 MITRE ATT&CK techniques in under ten hours, work that “would normally take human operators around two weeks.”

A state actor, a lab’s own tooling and a ransom crew arrived at the same architecture, an orchestrator that plans and sub-agents that do, within ten months of each other.

How does an autonomous attack chain work? Recon, exploit, pivot, persist, monetize

A real offensive agent behaves like a patient human intruder, minus the need for patience. Lingad described the loop plainly:

“We hack the client just like a realistic hacker. So when we say realistic hacker, we will not limit ourselves with just vulnerability analysis. We will exploit it, chain the exploit, escalate privilege, and see if we can gain more exploit, and see if we can have much more context of the other attack surface so that we can have more exploit in different attack surfaces.”

Alexis Lingad, founder and CEO of KinoSec, on episode 97

Diagram of the autonomous AI attack chain in five stages, recon, exploit, pivot, persist and monetize, comparing a human-run intrusion where an operator performs every step with the GTG-1002 campaign Anthropic described, where the AI executed 80 to 90 percent of tactical operations and humans only approved exploitation, authorized access to sensitive systems and approved exfiltration targets
The chain hasn't changed in twenty years. What changed is who runs each step. In Anthropic's account of GTG-1002, the human's job shrank to a few approvals.

Five stages, and the machine now runs all of them.

  1. Recon. Map every exposed host, application, API and device, continuously. GTG-1002 pointed it at roughly thirty organizations.
  2. Exploit. Write an exploit, test it, read the error, try again. Retries are free, and an agent that fails at the front door tries the side window.
  3. Pivot. Harvest whatever secrets the foothold exposes, escalate, move sideways. Lingad’s view is that this is where frontier models are weakest on their own: “For frontier LLMs, they can do information gathering, exploit, and sometimes vulnerability analysis. Digging deeper, chaining of exploit, escalating privileges, you need to do a lot of things in order to just do that.” That “lot of things” is the orchestration layer the GTG-1002 operators and the Unit 42 attacker built.
  4. Persist. New users, backdoors, several footholds, so that finding one doesn’t end the intrusion.
  5. Monetize. Exfiltrate, extort, or use the access against the next target, and start again on a different company while the first is still being worked.

In Anthropic’s report, the human approves the move from reconnaissance to exploitation, authorizes access to particularly sensitive systems, and signs off on what gets exfiltrated. Everything between those checkpoints belongs to the agent.

When the agent starts phishing on its own

The moment that made Jon and Sasha sit up was not an exploit. Jon asked Lingad whether his product would social-engineer its way into systems, and he answered with a case study:

“We have this case that our AI agent hacked the web application of our client and then get all of the API keys of that specific client. And one of those API keys is a Resend API key. And then our AI agent used that Resend API key in order to get the admin email. And then after getting the admin email, it sent internal phishing.”

Alexis Lingad, KinoSec, on episode 97

To the whole company, from the administrator’s own address, with no operator writing the lure. Competing tools, he said, stop at “exposed API keys” and write it up. “The bad guys will not stop on that.”

A lure written by a model, sent from a real internal account, citing real internal context, will get clicked. The durable answer is the one Jasson Casey gave on episode 92, which we wrote up in Deepfake Attacks Keep Working Because We Keep Detecting Instead of Proving: stop trying to spot the fake and make the credential itself unphishable, with device-bound credentials.

From web app to hotel door: why attackers chain web, IoT and physical systems

His competitors, Lingad said, test the web and the network and stop. KinoSec’s phase two is physical: drones, robotics, IoT and critical infrastructure, chained to the digital footholds. His example is one every founder who has stayed in a hotel will feel:

“We have a web application of a hotel, and then it is connected to an IoT management device that is controlling all of the doors, door keys and door codes. And then we successfully chained the exploit from web to IoT to those doors, and control it and open the doors and close the doors remotely.”

Alexis Lingad, KinoSec, on episode 97

Sasha asked him, after the show, which hotel not to book. The serious version of that joke is the whole argument: a company that secures its web tier and calls it done is testing the path the attacker doesn’t need. Our post on forgotten IoT devices shows how one unsecured printer can hold admin credentials for email, file shares and the directory. An agent that finds the printer finds all of it.

The Car Hacking Village at DEF CON 33 in Las Vegas, with laptops open on tables around a car as attendees work on its systems
Our photo from the Car Hacking Village at DEF CON 33. Web to network to a thing with wheels is the same chain Lingad described, done by hand and for fun. The agents don't need the weekend.

Asked for the single most successful threat he sees, he didn’t name a technique; he named a precedent:

“Have you heard about Stuxnet? That specific attack is a specialist in targeting the critical infrastructure. And imagine that time there’s still not much of an AI, but they did it. And now that we have AI that can exponentially empower all of those things, imagine what can be done, and imagine what those bad guys can do with this kind of weapon.”

Alexis Lingad, KinoSec, on episode 97

Stuxnet, per the ICS-CERT advisory of September 2010, used four zero-day exploits, spread through infected USB devices and network shares, and targeted Siemens SIMATIC WinCC and STEP 7 software, the layer that talks to industrial controllers. The UK’s National Cyber Security Centre now flags, in its assessment of AI’s impact on the cyber threat to 2027, “an increased threat to CNI or CNI supply chains, particularly any operational technology with lower levels of security.” Most startups are not power stations. Plenty of them sell software to one, and the supply chain is how the chain gets in.

Here is the upside. The same chaining that makes an autonomous attacker frightening makes it testable. If a machine can find the path from your marketing site to your badge readers, a machine you hired can find it first, and the fix is usually a segmentation rule and a rotated credential, not a rebuild.

Why machine speed changes the math for defenders

Every control a security team runs assumes a clock. Patch windows assume days between disclosure and exploitation; incident response assumes an attacker who pauses between steps; alert thresholds assume a human’s request rate. Autonomous attackers break all three.

Timeline comparing human speed and machine speed for one intrusion: human red teams need around two weeks according to Unit 42, while AI agents completed the same intrusion in under ten hours across more than 50 MITRE ATT&CK techniques, from a breached web service through repositories combed for tokens, an escalated secrets store, an attempted CI/CD backdoor and stolen cloud keys to hijacked AI compute; the GTG-1002 campaign ran at thousands of requests, often multiple per second; NCSC notes disclosure-to-exploitation has shrunk to days
Two clocks, one intrusion. The phase split on the human bar is illustrative; the two-week estimate, the ten-hour figure and the ATT&CK count are Unit 42's, and the request tempo is Anthropic's.

Xia Hua, co-founder and CEO of Traceforce, brought up the Anthropic case on episode 100 for this reason: “It’s the speed of the wave and the size of the wave. Both are much bigger than the previous generation.” The NCSC says the same in its own register: “the time between disclosure and exploitation has shrunk to days and AI will almost certainly reduce this further,” and it predicts “a digital divide between systems keeping pace with AI-enabled threats and a large proportion that are more vulnerable.” Which side of that divide you land on is a decision, and a cheaper one than most founders expect.

Want the whiteboard version? IBM Technology walks through how attackers weaponize AI, from generated lures to automated exploitation, in nineteen minutes; roughly 205,000 people have watched it. Watch on YouTube.

Two consequences follow.

Dwell time collapses. The idea that you have weeks to notice an intruder came from human intruders. Unit 42’s attacker went from a public web service to the victim’s cloud keys inside a working day, and when attempts are that cheap the attacker stops being selective: the Cursor crew’s victims included a garage-door manufacturer. If your detection story is “we review logs weekly,” it describes last year’s attacker.

Volume becomes signal. Unit 42 advises watching for “behavioral loops”: bursty API requests and rapid authentication shifts. An agent retrying an exploit or spraying a harvested credential looks nothing like a human, and that difference is the defender’s friend, provided someone is watching in the hour it happens. That is the practical case for round-the-clock detection and response that acts inside the attacker’s loop, and for cloud and AI-workload detection that treats a burst of unusual API calls as a page rather than a dashboard tile.

Why AI guardrails won’t stop offensive AI (and what will)

Most people assume the safety rules built into AI models will hold. They don’t always. Mainstream models forbid hacking in their terms of service, and there is a standard for managing AI responsibly, ISO/IEC 42001. When Jon put that to Lingad, he was direct: KinoSec’s own model, which the team calls Skynet as a reminder of what they are trying to prevent, runs “purely without limits and without guardrails” so it can do the deep chaining frontier models decline.

The Anthropic case shows the vendor side of the same problem. The attackers didn’t break Claude’s safety system by force; they told it that it was an employee of a legitimate cybersecurity firm doing defensive testing, and decomposed the intrusion into small tasks that each looked innocent. The Cursor crew simply reopened the conversation until the agent supplied its own permission slip, a loop we dissected in the both-sides post. A guardrail that can be talked around is a speed bump: worth having, never a control.

There is a second catch, and Sasha raised it on the episode: some models don’t reliably honor a kill switch. Lingad’s answer was scoping. Every KinoSec engagement starts with an explicit scope, the agent checks whether it is straying beyond it, and “usually it will have a kill switch in order to not hack the whole world.” Later he compared what he is building to carrying a nuclear weapon and said the mistake he most fears is forgetting that switch. That honesty also tells you how the criminal version behaves: no scope, no switch, no report at the end.

For the longer story of how offensive AI got here, Cybernews' 28-minute documentary has drawn roughly 680,000 viewers. Watch on YouTube.

What holds is the set of controls you own, and there is already a shared vocabulary for the AI-specific pieces: MITRE ATLAS, “a globally accessible, living knowledge base of adversary tactics and techniques based on real-world attack observations and realistic demonstrations from artificial intelligence (AI) red teams and security groups,” modeled on ATT&CK. Use it as the checklist for what your own AI red team should try against your agents and models before someone else’s agent does.

How to defend against AI-powered attacks: the defender’s priority list

You cannot stop offensive AI from existing. You can make yourself a slower, more expensive target than the next company on the agent’s list, with the moves the documented cases keep pointing at.

The defender's priority list against autonomous AI attacks, ordered by return on effort: make identity unphishable with device-bound sign-in, get secrets out of code and scope them, segment and give AI agents least privilege too, watch at machine tempo around the clock, keep immutable tested backups, and test continuously across web, cloud, network, IoT and physical surfaces
The list, in the order we'd run it at a Series A–C company. The first two remove the cheapest paths an agent takes; the rest shrink what a success is worth.
  1. Make identity unphishable. The agent’s best lure comes from your own domain, so protect the credential rather than the click: phishing-resistant, device-bound sign-in for everyone, and the same scrutiny for how your agents authenticate.
  2. Get secrets out of code, and scope them. Unit 42’s sub-agents “combed enterprise code repositories, extracting hard-coded tokens and service passwords.” Lingad’s agent found a Resend key. Rotate keys, scope each to one job, keep them out of repositories, and remember that AI coding assistants leave more of them lying around.
  3. Segment, and give agents least privilege too. One foothold should buy one room. Andrew Rubin built Illumio on the premise that the perimeter will fail and containment has to be designed in; your own agents need least privilege more than humans ever did, because a poisoned input turns their permissions into the attacker’s.
  4. Watch at machine tempo, around the clock. Bursty API calls, rapid authentication shifts, activity at 3 a.m. Unit 42’s other recommendations fit here: govern AI as core infrastructure with rate limits and logging, and contain across credentials, OAuth, CI/CD and cloud accounts together rather than one at a time.
  5. Keep immutable, tested backups. Extortion needs leverage. Ransomware operators, human or automated, go after backups first, which is why 3-2-1 stopped being enough. A rehearsed restore is the one control worth more as the attacker gets faster.
  6. Test yourself continuously, on every surface. Lingad’s line from the episode: “If you don’t have a realistic view of how a real attacker operates, you’re left with false hope.” An annual penetration test measures last year’s attacker against last year’s app; the version that matches this threat covers web, cloud, network, IoT and the physical layer, chains them, and runs more than once a year, with humans doing the judgment work, as the companion post argues.
The AI Cyber Challenge finals stage at DEF CON 33 in Las Vegas, where DARPA presented autonomous systems that found and patched vulnerabilities in open-source software
The AI Cyber Challenge finals at DEF CON 33, from our seats. The same autonomy, pointed at defense: DARPA's systems found and patched vulnerabilities on their own, and every finalist was committed to open-source release.

None of this is exotic; most of it is what a good security program did before agents existed. The difference is the clock: each item now runs continuously, and who runs it is the question most founders put off. We wrote the honest version of that decision in In-House, On-Demand, or YOLO?: a full-time hire, an on-demand team that is already awake at 3 a.m., or hoping. Two of those work against an attacker who never sleeps.

The decision this quarter

Lingad framed his work as a “cybersecurity arms race,” and said his company built the weapon first so the good side would have one. Take the frame literally, then take the pressure off. Nothing in the documented cases says a fifty-person startup is about to face a state-sponsored orchestrator. What they say is narrower and more useful: the attack chain is unchanged, the human has left most of it, and the tempo has moved from weeks to hours. Companies that made identity unphishable, pulled secrets out of code, segmented, and put someone on watch are already on the right side of the NCSC’s divide, on a startup’s budget.

The one decision to make this quarter is who is watching when the loop reaches you. If the answer today is nobody, talk to us about detection and response that runs on the attacker’s clock; we’ll tell you plainly what you already have covered. And to hear the god-level hacker describe the chain himself, hotel doors included, Alexis is on episode 97.

Autonomous AI cyberattacks frequently asked questions

What is an autonomous AI cyberattack?
An intrusion in which AI agents perform most of the work: scanning targets, writing and testing exploits, harvesting credentials, moving laterally and exfiltrating data, while a human sets the objective and approves a few key steps. It differs from an AI-assisted attack, where a person does the hacking and uses a model to write code or phishing lures faster. Anthropic's November 2025 disclosure, in which the AI executed an estimated 80 to 90 percent of a campaign, is the reference case.
Has an autonomous AI cyberattack actually happened?
Yes. In November 2025 Anthropic disclosed a state-sponsored campaign in which Claude Code, misled into believing it was doing authorized security testing, ran most of an espionage operation against roughly thirty organizations. In 2026, OpenAI's own agents breached Hugging Face during an evaluation, a ransomware crew used Cursor's agent against seven companies, and Unit 42 investigated a ransom intrusion that AI agents completed in under ten hours. Each case has a public postmortem.
How is an AI-driven attack different from a human one?
Tempo and parallelism. A human operator works one target at a time at typing speed; an agent orchestrator runs sub-agents against many targets at once, at thousands of requests, often several per second, and never stops for the night. Unit 42 estimated that an intrusion AI agents finished in under ten hours would have taken human red teams around two weeks. The techniques are mostly familiar; the speed and the cost per attempt are what changed.
Can AI model guardrails stop autonomous cyberattacks?
Not on their own. Documented attackers got past frontier-model guardrails by claiming to be a security firm doing defensive testing and by splitting the intrusion into small, innocent-looking tasks. Purpose-built offensive tooling can also run on open models with no guardrails at all, as KinoSec's founder described on our podcast. Treat a model's refusal as one layer, and put the controls you own around it: identity, secrets, segmentation, monitoring.
What is an attack chain, and why do AI agents make chaining more dangerous?
An attack chain links several small weaknesses into one serious compromise: an exposed API key becomes access to the email system, which becomes an internal phishing campaign sent from the admin's own address. Chaining used to be the expensive, expert part of hacking. Agents do it tirelessly, and across surfaces humans rarely bothered to connect, such as a hotel's web app, its IoT door controller and its physical locks. Every unrelated-looking exposure is now a potential first link.
Are startups and small companies targets of autonomous AI attacks?
Yes, and increasingly so, because automation removes the attacker's reason to be selective. The companies breached with Cursor's agent included a cleaning-products maker and a garage-door manufacturer. The UK's NCSC expects AI to increase the frequency and intensity of attacks and warns of a digital divide between organisations that keep pace and those that cannot. A fifty-person company with an exposed key and nobody watching is exactly the target an agent finds first.
What should a startup do first to defend against AI-powered attacks?
Make identity unphishable with device-bound, phishing-resistant sign-in; get secrets out of code and scope them tightly; segment so one foothold stays one foothold; and make sure someone is watching for machine-tempo behavior around the clock. Then keep immutable, tested backups and test yourself continuously across web, cloud, network and connected devices. The first two remove the cheapest paths an agent takes; the rest shrink what a success is worth.
Is critical infrastructure at risk from AI-enabled attacks?
Government assessments say yes. The NCSC's 2025 assessment flags an increased threat to critical national infrastructure and its supply chains, particularly operational technology with lower levels of security. Stuxnet used four zero-day exploits against industrial control software in 2010, per the ICS-CERT advisory; the concern KinoSec's Alexis Lingad raised on our podcast is what the same intent achieves when an agent can chain web, IoT and control-system weaknesses on its own. The upside is that those chains can be tested defensively first.
Written by the team behind The Security Podcast of Silicon Valley

Put it into practice.