Anthropic says GLM-5.3 has crossed an important threshold in AI-assisted cybersecurity: in controlled evaluations, the model developed working exploits from start to finish. Anthropic’s analysis also describes a separate concern about access. Because GLM-5.3 is distributed as an open-weight model, users can modify its parameters and remove or weaken its refusal behavior more easily than they can with models whose weights are not publicly available.
The findings come from Anthropic’s own benchmark evaluations, researcher-led experiments and simulated tests. They do not establish how often the techniques would work against live systems, deployed software or real users. Anthropic says the results nevertheless show that a capable open-weight model can provide attackers with substantial cyber capabilities without the access restrictions used for some other advanced models.
What GLM-5.3 accomplished in exploit tests
Anthropic evaluated GLM-5.3 in isolated, sandboxed environments designed to prevent attacks on external systems. The main focus was end-to-end exploit development: finding a vulnerability and turning it into a working method of exploiting that flaw, rather than merely identifying suspicious code or suggesting a partial attack.
On ExploitBench, a benchmark involving known vulnerabilities in the V8 engine used by Google Chrome, GLM-5.3 successfully developed end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview, an Anthropic model, succeeded in 56 of the same number of attempts in Anthropic’s comparison.
A separate internal benchmark tested whether models could find and exploit flaws in open-source projects associated with Google’s OSS-Fuzz program. GLM-5.3 achieved a full control-flow hijack in 4% of trials, compared with 6% for Claude Mythos Preview. Anthropic says earlier models in the comparison, including Claude Opus 4.6 and GLM-5.2, did not succeed on any of those tasks.
These figures are results from Anthropic’s selected evaluations, not a general success rate for attacks. The tasks were run against prepared targets in isolated environments. The source report provides the benchmark comparisons and methodology details available for these tests.
Researcher-led tests found more than benchmark scores
Anthropic also tested GLM-5.3 in workflows where human researchers supplied limited assistance while the model worked on unfamiliar targets. The sessions generally lasted a day or less and involved less than an hour of direct human attention.
In one sandboxed browser test, GLM-5.3 found previously unknown vulnerabilities in a Linux build of a popular browser’s JavaScript engine. It chained several of those flaws into a webpage exploit that could read arbitrary files from the visitor’s computer in the test environment. Anthropic says it disclosed the vulnerabilities to the relevant maintainer. The experiment used the Linux build available to the model, and the reported exploit was not tested against live browser users.
Anthropic also says the researcher identified exploitable flaws in wireless and graphics drivers and in software for network-facing devices. The report says those findings were under review for disclosure, so they should not be treated as confirmed vulnerabilities affecting deployed systems.
A second experiment used GLM-5.3-Flash, described as a smaller and less capable version of the model, against a recently disclosed Chrome vulnerability and another known flaw. A known, or N-day, vulnerability is a flaw that has already been disclosed, unlike a 0-day, which is a previously unknown vulnerability at the time of discovery. Anthropic says the smaller model combined the two known flaws into an exploit chain for an ARM64 target during a session that involved 20 minutes of human attention and eight hours of model work.
That result illustrates a different risk from discovering novel vulnerabilities. Once technical details or public fixes are available, a capable model may reduce the time and expertise needed to turn that information into a working exploit. It still does not show that the same chain would succeed against every target or configuration.
Why open-weight access changes the safeguard problem
An open-weight model is a model whose learned parameters can be downloaded and run or modified by users. That is different from an API-only system, where the provider controls the model and can enforce safeguards at the service level.
GLM-5.3 reportedly includes refusals for some clearly harmful requests. Anthropic says those refusals can be weakened in two ways. Users can try to persuade the unmodified model to treat a harmful request as an authorized exercise, or they can modify the downloadable model itself.
Anthropic calls the second method abliteration, a technique for changing a model so that it produces fewer refusals while retaining much of its general capability. The report says its team created an abliterated version of GLM-5.3 using about 2,200 GPU hours at a cost of roughly $4,400. Anthropic estimates that a team already familiar with the technique could do the work with closer to 600 GPU hours, or about $1,200, although those are the company’s estimates rather than independently verified costs.
After the modification, Anthropic reports that GLM-5.3’s refusal rate fell from above 90% to about 3% on JailbreakBench and about 2% on HarmBench. Its refusal rate was 12% on StrongREJECT. Anthropic also says the modified model produced similar results to the standard model on the GPQA-Diamond scientific evaluation and lost only a few percentage points on the tested CyberGym subset.
The practical concern is not simply that one prompt can defeat a refusal. With downloadable weights, a user can alter the model once and distribute the modified version to others. The model’s safeguards are therefore not only a conversational barrier controlled by a provider; they are also a property that users may be able to change.
What the bypass tests showed—and what they did not
Anthropic tested GLM-5.3 in a simulated environment involving overtly malicious requests to attack critical systems. The model refused direct requests in the reported trials, but Anthropic found that simple changes increased the rate at which it engaged with the tasks:
A deceptive prompt describing the model as an autonomous red-team agent produced engagement in 64% of trials.
Prefilling the model’s reasoning process produced engagement in 92% of trials.
The abliterated version engaged in 100% of trials.
Anthropic says the corresponding techniques did not cause the safeguarded Claude models it tested to carry out the harmful tasks. The comparison is limited to Anthropic’s chosen models, prompts and test conditions; it is not a universal safety ranking of all AI systems.
The simulation also imposes an important limitation. According to Anthropic, model-generated code was not executed, the model could not interact with external systems, and a separate language model approximated the results of commands in a fictional environment. Such a test can measure whether a model is willing and able to plan or respond under specified conditions, but it cannot establish that an attack would work against a live target.
The broader security significance
Anthropic interprets the results as evidence that GLM-5.3 could give malicious actors access to advanced exploit-development capabilities without the vetting or access controls used for some other frontier models. The company also says the same capabilities could help defenders find and fix vulnerabilities before attackers use them.
That creates two separate policy questions. The first is capability: how well can a model discover, chain and exploit flaws? The second is availability: who can use that capability, under what safeguards, and how easily can those safeguards be changed? GLM-5.3’s reported results are notable because Anthropic says it performs near Claude Mythos Preview on some exploit-development tests while being available for download and modification.
The source report also cites a separate assessment from NIST’s Center for AI Standards and Innovation, which described GLM-5.3 as the most cyber-capable open-weight model it had assessed and placed it about four months behind the US frontier on an aggregate of its benchmarks. Anthropic says the comparison used different access conditions: US models were tested with safeguards disabled where applicable, while some of the most capable versions were available only to vetted users. The CAISI finding is therefore an attributed assessment, not an indication that all compared models were available under the same conditions.
Anthropic argues that governments should conduct safety testing on sufficiently capable models, including future open-weight successors. Its report’s strongest practical implication is not that GLM-5.3 guarantees successful attacks. It is that the combination of advanced exploit-development ability and easily modifiable safeguards may shorten the path from a model’s technical capability to broad, uncontrolled access.
For defenders and policymakers, the reported experiments therefore point to a need for testing that covers both capability and distribution. Measuring what a model can do in a sandbox is only one part of the question. Assessments also need to examine whether users can bypass refusals, remove them from downloadable weights, reproduce the results, and deploy the resulting capability outside the controlled environment.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment