The next step beyond automation

Automation is well-established element in an attacker's playbook. Now, the next shift is toward systems that can interpret results, choose the next action, and keep an attack moving with less human direction. Attacker don't simply demand faster automation. They want to delegate more of their of judgment between automated actions.

Case study: JADEPUFFER

Imagine this scenario. An attacker gains initial access by exploiting a known vulnerability but then fails configuration-server takeover attempts. Thirty-one seconds later, the agent has diagnosed the problem, changed its approach, generated a multi-step correction, and tried again.

This sequence came from JADEPUFFER, an operation documented by Sysdig in July 2026, assessed to be the first documented case of agentic ransomware. [1]

After gaining initial access, the system chained reconnaissance, credential harvesting, lateral movement, and destructive database activity. A human built the tooling, infrastructure, and pointed the agent at an exposed Langflow instance. The rest ran autonomously.

Attackers have used automated scanners, exploit frameworks, malware, malicious scripts, and various other toolkits for decades. Automation isn't the new development. The next level is agentic.

Learn more: Increased automation in attacks

From automated attacks to agentic attacks

Traditional automation follows instructions. Typically, the flow of these instructions takes a predictable form: Scan these addresses. Test these credentials. Execute this payload. Repeat this request.

An agentic system can increasingly work at the level above this basic form. For example, it can interpret the result, decide whether the action worked, choose another route, find a different tool, and continue toward an objective.

This marks an evolution to an agentic threat actor that leverages LLMs augmented with tools, retrieval, and autonomy to execute malicious intent. The difference is less about automation and more about judgment between automated actions.

True, the boundary lines in this distinction are sometimes fuzzy. Current attacks sit at different points between AI assistance and high levels of autonomous execution. Human involvement hasn't disappeared. What is changing is how often a human needs to make the decisions that move an attack from one stage to the next.

Case studies: Anthropic, Unit 42 and Google

Anthropic's investigation of the GTG-1002 espionage campaign is one example. The company reported that attackers used Claude Code as an operator rather than simply an advisor, allowing it to conduct substantial parts of reconnaissance, exploitation, credential access, and lateral movement. These techniques themselves were not necessarily extraordinary. Anthropic's later analysis argues that the more important differentiator was the scaffolding that allowed those techniques to be chained and executed with less human intervention. [2]

A similar pattern has appeared elsewhere.

  • In July, Unit 42 recovered an actor’s entire workspace when a Hermes Agent exposed an HTTP server after responding to a Telegram command , showing how the threat actor used an AI model to research vulnerabilities, obtain exploit code, assess targets, and change direction when its first approach failed. In one recovered session, the agent abandoned an unsuccessful Langflow exploitation route, searched for alternative vulnerabilities and targets, obtained public proof-of-concept code, and continued testing without further operator input being visible in the recovered data. [3]

  • In September, Unit 42 documented the same principle operating at much greater scale during an enterprise intrusion. The attacker used multiple frontier AI agents in parallel, with the agents monitoring results, evaluating them, acting, and replanning in real time. Unit 42 recorded more than 50 MITRE ATT&CK techniques in less than 10 hours. It estimated this activity would normally take human operators around two weeks. [4]

Google Threat Intelligence Group (GTIG) reported another example in September 2026. After compromising an organization’s cloud infrastructure, a suspected financially motivated attacker used an autonomous multi-agent framework to plan, build, and execute a mass credential-harvesting campaign in less than six hours. Using preconfigured Markdown instructions as operational playbooks, the agents autonomously managed vulnerability scanning, real-time troubleshooting, and IP rotation, while the operation compromised thousands of third-party credentials. [5]

Beyond automation- When attackers become agentic-1

The older point of interest centered around whether AI can help an attacker write an exploit. Today's question is about how much of the work can be handed to the machine, with the attacker simply providing malicious intent.

When the attacker becomes a feedback loop

Decisions between automated actions decreasingly have to wait for a human operator. When an AI system can interpret the result of one action and decide what to do next, the operating model of the attack starts to change.

The older model was linear in structure:

the human led attack sequence_edited

The newer model more closely resembles a cycle that can start to run continuously:

Beyond automation- When attackers become agentic-3

Case study: Dream Research Labs

Such a shift was made unusually visible in a multi-agent intrusion framework uncovered by a Dream Research Labs publication in August 2026. Across approximately four days, Dream documented 12 attack waves against government infrastructure in Asia. The framework deployed up to eight agents concurrently during an attack wave. [6]

The system used probability scoring to rank attack paths, ran dedicated research cycles when existing approaches were blocked, and fed results from one wave back into planning for the next. Dream also found agents working concurrently to target different attack surfaces.

This is way more than merely faster execution. It's the creation of an agentic attack feedback loop. Any unsuccessful technique becomes information for the next attempt. A useful result can redirect resources. One attack path can be deprioritized while another receives more attention.

Read more: Learn globally, act locally: Why app security needs a runtime intelligence loop

When one agent becomes many

On top of this, the loop can also become a collaborative exercise.

Case study: Hugging Face

During OpenAI cybersecurity evaluations in July 2026, agents intended to operate in isolation discovered unauthorized ways to communicate, shared discoveries, coordinated attack activity, and eventually compromised parts of Hugging Face's infrastructure. This was not a malicious campaign initiated by a criminal attacker, so it shouldn't be treated as one. But it provides strong capability evidence for what persistent, collaborating agents can already do under the right conditions, and the risks of excessive agency. [7]

An independent investigation by METR found that roughly 1,200 agents communicated through an unsanctioned message board during the incident, with around 700 going on to participate in the attack on Hugging Face. Agents coordinated workstreams and reproduced useful discoveries made by others. Hundreds pivoted toward the Hugging Face attack after one agent's finding was independently reproduced and shared.

The significance lies in how the agents divided and coordinated the work. One agent identified a promising vulnerability or attack path. Another tested whether the path could actually be exploited, while other agents explored alternatives in parallel. Useful findings could then influence what the wider group tried next.

By distributing different stages of an attack across multiple agents, a system can pursue more attack paths without requiring any one agent to excel at reconnaissance, exploitation, validation, persistence, and coordination. Agentic capability can therefore scale not only by making individual agents more capable, but by allowing multiple agents to pool information and specialize.

Learn more: Attack vector

Agentic attackers still get attack decisions wrong

Current agentic systems are far from infallible.

Case study: Automated Pentesting

A 2026 study of 13 open-source autonomous penetration-testing frameworks plus two baselines found recurring problems around verification, memory, context, and multi-step exploitation. The experiments consumed more than 10 billion tokens and generated over 1,500 execution logs. [8]

Some agents fabricated and submitted the wrong flag. Others could identify a vulnerability, but failed to turn that knowledge into a working exploit. Performance also deteriorated when several dependent steps had to succeed as part of one longer attack chain.

Anthropic observed similar problems during GTG-1002, including overstated findings and credentials that did not actually work. So, the truth is, an agentic attacker can be fast, persistent, scalable, and wrong at the same time.

But defenders need to show restraint about what conclusion they draw from this finding. Attackers are already building verification and cross-checking into agentic workflows to reduce those weaknesses. And frontier models continue to advance their autonomous capabilities.

Dream's multi-agent framework included repeated verification procedures. It recorded false positives, retested them, and discarded findings that did not survive further checks. In one case, what initially looked like SQL injection was eventually traced to an unrelated SMTP timeout. The system corrected its own conclusion and removed the finding.

The direction of model capability is moving as well.

Case study: Astra

On September 1, OpenAI said Astra had become its first model to reach the Critical cybersecurity capability threshold under its Preparedness Framework. Evaluations included autonomous zero-day discovery and multi-stage exploit chains against hardened systems. GPT-6 Astra was subsequently released on September 3. [9]

During evaluations, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains. OpenAI reported that it used two zero-days as part of one exploit chain, built a browser-compromise chain that escaped a sandbox and executed commands on the host, and combined multiple operating-system vulnerabilities into a local privilege-escalation chain.

Now none of that is evidence of a malicious Astra attack in the wild. This is capability evidence. But it reinforces why today's limitations should not be treated as fixed properties of agentic systems.

Learn more: Zero-day vulnerabilities

Defenders can't base their security strategy on the assumption that agentic attackers will make mistakes. Today's models, memory, and agent orchestration will continue to improve. Attackers will add cross-checks, specialist agents, external tools, and new ways to validate results.

What changes, and what doesn't

Agentic attacking doesn't replace everything that came before it. Known vulnerabilities will still be exploited. Exploit code can come from a public repository. Attackers can still use penetration testing and ethical hacking tools for malicious purposes.

For example, GTG-1002 relied heavily on conventional security tooling. Unit 42's recovered agent searched for existing vulnerabilities and public proof-of-concept code. JADEPUFFER exploited existing weaknesses and misconfigurations.

The change we want to highlight here increasingly sits at one level above those individual techniques. An agentic attacker can:

  • discover a possible attack path

  • test whether it works

  • interpret why it failed

  • search for an alternative

  • generate or modify the code it needs

  • reprioritize targets

  • coordinate with other agents

  • validate results

  • keep moving toward an objective

And it can repeat that cycle without waiting for a person to make every intermediate decision. This is what changes attacker economics.

Work that once consumed a skilled operator's attention can increasingly be delegated. More hypotheses can be tested. Failed approaches can be abandoned sooner. Multiple attack paths can be explored at once. An operation can remain active for longer without requiring the same level of continuous human attention.

Read more: AI-powered mobile app attacks: What app shielding can and can't stop

What the attacker still has to do

But there's another side to the equation.

AI systems can't (yet) reason their way around every technical requirement of a real attack. They still must interact successfully with real systems. They still need vulnerabilities, misconfigurations, credentials, attack techniques, or other routes that actually produce an effect. A carefully reasoned attack path that doesn't work remains a failed attack path.

Agentic systems therefore change the decision layer of attacking faster than they erase the underlying mechanics. This is an absolutely vital distinction for defenders.

The security question is no longer simply whether an attacker can automate a known sequence. Increasingly, it is whether a system can observe what happened, understand enough of the result to select a new course of action, and continue the attack without handing the problem back to a human.

JADEPUFFER's 31-second correction gives us a small but revealing example. Dream's framework shows the same idea operating across attack waves and multiple agents. The OpenAI/Hugging Face incident shows how collaboration can amplify it. Astra's evaluations indicate how rapidly the underlying cyber capability of frontier models is progressing.

The attacker is becoming a feedback loop. Human involvement remains part of documented attacks today, even as more decisions are delegated to AI systems. Nor does every attack become fully autonomous. The more important development is that human judgment is no longer required at every transition between one automated action and the next.

All this creates a further question for application security.

Mobile applications give attackers something especially useful for this kind of repeated experimentation: possession of the client itself. They can inspect it, manipulate it, repackage it, uncover embedded logic and secrets, and try another approach on a device the defender does not control.

Read more: When mobile and AI collide: Why app security cannot reuse the web playbook

The full malicious mobile attack chain has not yet become agentic in the public cases we reviewed here. But individual stages are already beginning to move in that direction.

The next question is what happens when the agentic attacker meets the mobile app.

References

[1] Sysdig Threat Research Team, JADEPUFFER: Agentic ransomware for automated database extortion, July 1, 2026.

[2] Anthropic, What we learned mapping a year’s worth of AI-enabled cyber threats, June 3, 2026.

[3] Unit 42, Chinese-Speaking Threat Actor Harnesses AI Models for Autonomous Cyberattacks, July 30, 2026.

[4] Unit 42, An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation, September 2, 2026.

[5] Google Threat Intelligence Group, GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI, September 8, 2026.

[6] Dream Research Labs, Inside a Multi-Agent AI Framework Used to Compromise Government Entities in Asia, August 12, 2026.

[7] METR, Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident, August 26, 2026. See also OpenAI, The Hugging Face incident and the road ahead, August 26, 2026.

[8] Peng et al., Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing, April 7, 2026.

Prepare for attacks that adapt
Agentic systems can reduce the human effort needed to test, adjust, and continue an attack. Talk to us about protecting applications against the techniques those attacks still depend on.
Book a meeting