Guide · AI and security
AI hacking agents:what is real and what changes.
AI is already reducing the effort needed to find vulnerabilities: security flaws that can expose data or allow actions without permission. Google and Mozilla have documented real cases. That progress matters to people protecting software and to those trying to attack it, and the capabilities are reaching more people.
What an agent adds to a security review
An AI model can read code and point out a possible mistake. An agent can also use tools, check what it has found and continue the review using that result. This ability to carry work through several steps helps explain why better coding models also make better security tools.
The UK National Cyber Security Centre warns that AI will make finding and exploiting flaws easier, faster and cheaper. For a business, that leaves less room to put off fixes. Sources and model versions reviewed on 30 September 2026.
For the distinction between fixed workflows and agent autonomy, read where AI agents help and what should stay under rules.
Three cases that go beyond a demo
In July 2025, Google reported that its Big Sleep agent had found a flaw in SQLite, a database component used in many applications. The investigation started with information from its threat intelligence team: a capable agent was working with context supplied by people.
In March 2026, Mozilla confirmed 22 security flaws, including 14 rated high severity, found through its collaboration with Anthropic. Its engineers verified them and shipped fixes in Firefox 148.
In April, Mozilla reported another 271 flaws fixed in Firefox 150 after evaluating Claude Mythos Preview. The two figures come from different investigations; they cannot tell us how many times better the model had become.
The useful evidence is that the software maintainers confirmed the problems and released fixes. That carries more weight than a convincing chatbot response.
Finding a flaw is not the same as breaking into a system
Different flaws allow different things. In an online shop, a bug might expose order details to the wrong person without giving them access to the rest of the business. That is still serious, but its reach differs from gaining control of the server.
A useful review therefore needs to explain what is wrong, who is affected and how to check the fix. An alert from an agent is the start of that work.
- 01 Possible flaw The agent flags something worth reviewing.
- 02 Confirmed flaw The problem is verified.
- 03 Known impact The affected data or functions are identified.
- 04 Verified fix The change is checked to confirm it solves the problem.
Why some providers limit these capabilities
Anthropic provides a concrete example. As of September 2026, Mythos 5.1 is reserved for vetted organisations. Fable 5.1 shares the same underlying model and can find flaws in source code, but has safeguards restricting penetration testing and the creation of code that exploits vulnerabilities.
The capability exists even when the service limits its use. This helps explain why a tool may agree to review an application and refuse other tasks. Conditions depend on the provider and version; US models do not all work the same way.
GLM-5.3: what we know about progress in Chinese models
GLM-5.3, from Z.ai, can now be downloaded and run outside the developer’s service. Its model card reports improvements over GLM-5.2 in both coding and security tests. These are the developer’s results, which should be considered alongside external evaluations.
On 29 September, Anthropic described GLM-5.3 finding previously unknown flaws and combining them into a working attack against a browser during isolated tests with researchers. It also found inadequate safeguards against misuse. This is another developer’s evaluation, adding evidence beyond Z.ai’s own results.
On 17 September 2026, CAISI, the US AI evaluation centre within NIST, assessed it as the most capable open-weight model for cybersecurity evaluated up to that point. It nevertheless estimated a gap of about four months behind leading US models on its tests.
The four-month gap compares released models; it is not a forecast. CAISI also tested US models with cybersecurity safeguards disabled where applicable. For a business, who can access these capabilities matters alongside what a model can do.
What changes when a model can be downloaded
“Open weights” means the data needed to run a model can be downloaded and used on your own infrastructure, subject to its licence. It does not require publication of the entire training process or mean it will run on any laptop: large models still need powerful hardware.
Control changes too. As the UK AI Security Institute, AISI, explains, a provider can block accounts on its service but has much less control over copies running elsewhere.
This widens access for researchers and businesses, as well as the scope for misuse. Access no longer depends entirely on a few platforms, although using the tools well still requires resources and expertise.
What I would check in a website or application today
For someone maintaining a website, the immediate decision is to review which flaws could affect customers, data or day-to-day operations. I would start here:
- Know which services are accessible from the internet and retire those no longer in use.
- Give someone responsibility for applying security updates, with a clear deadline.
- Check that each user can only access their own data and that private keys stay on the server.
- Give each account, integration or agent only the permissions it needs.
- Keep logs for investigating incidents and test that backups can be restored.
AI can help with this review too. I would use it on my own code, in a test environment and with limited permissions. A person needs to confirm the findings and review the fixes. A resolved problem is a useful result; an agent finding nothing does not guarantee that everything is sound.
To examine where an application keeps credentials and authority, start with mobile app secrets, API keys and backend security.
Sources and further reading
- Google · Big Sleep, July 2025SQLite finding informed by threat intelligence.↗
- Mozilla · March 2026Validation and fixes in Firefox 148.↗
- Mozilla · April 2026Results from the Mythos Preview evaluation.↗
- Anthropic · access and safeguardsTerms checked on 30 September 2026.↗
- Z.ai · GLM-5.3 model cardAvailable weights and results published by the developer.↗
- Anthropic · GLM-5.3 study, 29 September 2026Another developer’s assessment of capabilities and safeguards in controlled environments.↗
- CAISI/NIST · GLM-5.3, September 2026External assessment published on 17 September and its test conditions.↗
- AISI · open-weight risksControl differences after distributing a model.↗
- NCSC · defence as AI advancesBusiness guidance, 15 April 2026.↗
Questions before we start
Does AI make hacking easier?
Yes, it reduces the time and knowledge needed for some tasks, including finding certain flaws. This helps both people reviewing software and those trying to attack it. Whether an attack succeeds still depends on the system and its defences.
Does GLM-5.3 match the best closed models in cybersecurity?
Not according to CAISI’s September 2026 evaluation. It stands out among open-weight models but remained behind leading US models on those tests. Results also depend on the tools and safeguards used in each evaluation.
Can an AI review tell me whether my website is secure?
It can uncover problems and help fix them, but cannot certify security on its own. Findings need validation, and permissions, configuration and external components need review too.