A new community-hosted GLM-5.3 UNCENSORED EXL3 3.0bpw quant is available on Hugging Face. In a post on X, @k2sbhai points readers to Infatoshi’s repository, which describes the model as an EXL3 conversion of a weight-edited GLM-5.3 FP8 checkpoint. The release is listed under the MIT license and uses Safetensors, but it is not presented as an official Z.ai release.

Image credit: @k2sbhai on X
The repository says it is not affiliated with either dealignai or Z.ai. That distinction matters because the model’s lineage includes the upstream GLM-5.3 model, a separate weight-edited FP8 variant, and Infatoshi’s subsequent EXL3 conversion.
What was released
The release is Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw. EXL3 is a quantization format used by ExLlamaV3 that stores model weights at a lower average number of bits per weight. Lower-bit storage can reduce the memory required for inference, although the resulting footprint still depends on the model architecture, unquantized components, runtime overhead and context settings.
The model is described as a mixture-of-experts, or MoE, model. In an MoE architecture, the model contains many expert subnetworks but routes each token through only a subset of them. The repository identifies the architecture as GlmMoeDsaForCausalLM, with 753 billion total parameters, 256 routed experts, eight active routed experts and one shared expert. It also lists 78 layers plus one MTP layer, with MLA attention and a DSA sparse indexer.
The repository reports an average bitrate of 3.04 bits per weight. Its conversion command used ExLlamaV3’s -b 3.0 --hq settings, with attention and shared experts at 5 bits per weight, dense MLPs at 4 bits, routed experts at 3 bits and the language-model head at 6 bits. The repository lists the resulting model size as 273 GiB.
How the checkpoint was assembled
The release is not described as a fine-tune. Its documented lineage is:
Z.ai’s GLM-5.3 is the upstream base model.
dealignai’s GLM-5.3 UNCENSORED FP8 is a weight-edited variant of that model. Its documentation describes a permanent edit to selected residual-writer weights rather than a fine-tune, LoRA or runtime modification.
Infatoshi converted that FP8 checkpoint to EXL3 using ExLlamaV3, reading directly from the FP8 source.
The Infatoshi page says the weight-edit instructions are documented in CRACK_SURGERY.json, copied unchanged from the FP8 source repository. Because the final quant is based on the edited FP8 checkpoint, comparisons with stock GLM-5.3 do not isolate the effect of EXL3 quantization alone.
The related JANGQ-AI GLM-5.3 FP8 page documents another FP8 quantization of the upstream model. It is part of the repository’s model tree, but the Infatoshi release specifically identifies dealignai’s weight-edited FP8 checkpoint as its conversion source.
What the reported tests show
The repository reports a limited fidelity comparison against its FP8 source using 20 rows of the WikiText-2 test set, with sequences of 2,048 tokens. It reports KL divergences of 0.089 for quantized-versus-FP8 probabilities and 0.097 in the reverse direction. Reported perplexity was 3.440 for the quant and 3.302 for the FP8 source.
The page also says that the median KL divergence was 0.0011 on tokens where the FP8 model’s top probability was at least 0.95, a condition covering 44% of the tested tokens. These figures describe the supplied comparison; they are not an independent evaluation of the model’s general quality.
The agent evaluation used tau2-bench’s airline and retail tasks. The quant was compared with stock GLM-5.3 served at FP8 by Z.ai through OpenRouter, rather than with the dealignai FP8 checkpoint. The repository therefore notes that the comparison combines the dealignai weight edit and EXL3 quantization.
On the reported setup, the quant scored 0.640 ± 0.048 on 50 airline tasks across two trials, compared with 0.710 ± 0.045 for the reference. On 114 retail tasks, it scored 0.482 ± 0.047 in one trial, compared with 0.504 ± 0.033 across two baseline trials. The repository describes neither difference as statistically significant at those sample sizes, so the results should not be treated as proof of a broad performance regression or improvement.
Reported serving configuration
The repository says the quant was tested with TabbyAPI and ExLlamaV3 on eight A100 GPUs with 40GB of memory each. Its example configuration includes:
model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
backend: exllamav3
max_seq_len: 65536
cache_size: 98304
gpu_split_auto: true
tool_format: glm4_7
reasoning: true
draft_model:
draft_mode: mtpThe MTP layer is a multi-token-prediction layer that can be used as a speculative drafting mechanism. In the supplied TabbyAPI configuration, it is enabled with draft_mode: mtp. The reported setup uses a maximum sequence length of 65,536 tokens and a shared cache of 98,304 tokens.
These are the repository’s reported test conditions, not a hardware compatibility guarantee. The X discussion included an estimate that roughly 4–6GB of VRAM might be enough for the 3.0bpw quant, but that estimate is not supported by the repository’s reported 273 GiB model size or its eight-A100 test configuration. Users should not treat it as a verified recommendation for consumer GPUs.
The repository also displays a configuration parsing warning: quantization_config.bits must be an integer. The page does not specify whether this warning prevents loading or how it should be corrected, so users should check the configuration before attempting deployment.
The tool-calling problem for agents
The most consequential warning is about structured tool calls. The model reportedly writes tool arguments as raw text, using forms such as <arg_value>9523456873</arg_value>. The repository says TabbyAPI’s glm4_5 parser, at the cited commit, JSON-decodes every value without checking the tool schema.
That behavior can turn string-like values into integers. Examples include order IDs, product IDs and ZIP codes, where the digits may look numeric but must remain strings for the receiving tool. The repository reports that this caused about 70% of tool calls to fail on its tau2-bench retail testing.
Its recommended fix is to make the parser preserve parameters whose schema type is string as raw text before decoding other values. The agent evaluation results above were run with that fix applied. Anyone testing the model in an agent workflow should therefore validate argument types at the parser or tool boundary instead of assuming that a successful text response means the tool call is safe to execute.
What remains unestablished
The reported testing covers the documented TabbyAPI and A100 setup; it does not establish compatibility with ordinary consumer GPUs, performance in other serving environments, or production readiness for tool-using agents. The release also does not include an independent evaluation of its results.
The reported agent comparison is useful as an initial signal, but it does not separate the effects of the dealignai weight edit from the EXL3 conversion. The fidelity test is also small, using 20 rows of WikiText-2 rather than a broad benchmark suite.
The practical takeaway is narrower: Infatoshi has published a 3.0bpw EXL3 conversion in the GLM-5.3 UNCENSORED lineage, with documented ExLlamaV3 and TabbyAPI settings and a specific parser fix for tool calls. The release is best approached as a community-hosted technical experiment whose hardware requirements and agent behavior still need to be tested in each target environment.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment