Much of the speculation around the capabilities of AI-powered models such as Mythos and Fable has focused on the future.

There have been widespread, and wide-ranging, predictions about the speed with which models capable of autonomously discovering zero-day attacks will take shape. It’s not just businesses that are worried. Governments and regulators everywhere are racing to build guardrails and policies, and the uncertainty around future performance only makes that job harder. Tracking a fast-moving target makes estimates on timeframes and capabilities subject to constant review.

With the power of these models potentially posing risks to public security, in May, the NCSC issued a warning to businesses, urging them to prepare now for when a ‘patch wave’ of urgent software updates arrives and to address the ‘technical debt’ of unresolved issues.

But is the hype around the future of ‘superhuman’ AI capabilities justified? And just how much of a step change do systems like Anthropic’s Mythos actually represent?  

The automated research pipeline

Much of the conversation is still theoretical, but we wanted to understand what is possible now, and how far current AI models can take us in finding real, exploitable vulnerabilities in production software.  To understand this, we created our own automated vulnerability research pipeline, with no human in the loop. Our pipeline took a codebase and ran it through a traditional code scanning engine which, for this experiment was Joern,  generating slices of code relevant to each finding. 

We used the LLM to triage and exploit the issue. This is because, whilst LLMs are great at taking small segments of code, or a description of a specific problem, if they’re asked to find security problems in a large codebase, they will ingest every file in the repository trying to find everything.

We then ran it against the top 200 WordPress plugins. These are already picked over extensively by bug bounty researchers. Finding an impactful bug would show our process can hold its own against human researchers.

The codebase is run through Joern with a set of rules that were designed to look for ‘interesting things’.  This broad prompt was used intentionally to avoid creating rules that are too specific and might miss bugs. The triage agent would filter any results, so in this experiment, erring toward false positives was an acceptable trade-off.

What our experiment revealed

The first finding that stood out was a SQL injection vulnerability in the Creative Mail plugin. It was significant for a number of reasons; it was high impact and it required multiple steps to exploit which made it less likely that a ‘classic’ tool would detect it. 

Our finding aligned with an independent discovery by another researcher, which was assigned CVE‑2026‑3985 and publicly disclosed.

In short, a remote, multi-stage SQL Injection zero-day we discovered in a WordPress plugin with more than 300k users, was fully automated from discovery through to exploitation, with no human in the loop.

What we discovered shows us not only that today’s pre-Mythos LLMs are already pretty capable but also underlines the implications and importance of attack surface management, here and now.

Ultimately, it tells us that the new normal is already here and many of the concerns being discussed as future risks are already relevant today.

Attack surface management and patching speed

It’s clear that AI has a big part to play in speeding up and streamlining vulnerability research.  But what we found is one example of how current LLMs can already assist in discovering and exploiting real zero-day vulnerabilities.   

The lesson for businesses is to assume that attackers are already using AI‑enabled tooling to find complex vulnerabilities quickly. 

And in an era in which attackers can move at lightning speed from discovery to exploitation, there are very real consequences to delays in patching and remediation of vulnerabilities. Research reveals that organizations are struggling to close security gaps, with mid-market businesses facing the longest remediation times, averaging 56 days to remove exposures. This leaves them dangerously at risk.

The answer lies in processes and technology that can automate discovery and patching and provide continuous attack surface management. Detection capabilities need to keep pace with the speed of AI‑accelerated discovery and, to this end, continuous exposure management platforms can instantly flag exploits powered by AI models.

Hype and reality

The hype around the capabilities of AI-powered models will continue to intensify; one recent report from the UK government’s AI Security Institute (AISI) claims that capabilities are improving faster than expected. 

How this rate of progress on the latest frontier models will evolve continues to fuel debate but the message for security leaders is clear: the time to take action to reduce and remove long patch windows for exposed services, is now. Current LLM models are already capable of discovering and exploiting subtle, multi-stage vulnerabilities.   

The teams that recognize this and take the steps to close these gaps with continuous threat exposure management will be best placed to meet the challenges ahead, in whatever form they take.

Sam Pizzey

Sam Pizzey is a security engineer at Intruder

Personalized Feed
Personalized Feed