QLean model quantization tool (v0.1.0 post1)
Updated at:
Copy as MD
Installation
The tool is a wheel package that can be installed using pip.
# Install dependencies
pip install triton_kernel==1.0.0+ppu2.0.0.oe
# Install qlean
pip install qlean==0.1.0+ppu2.0.0post1Usage
1. W8A8-INT8 quantization for Qwen3.5
Create Qwen3.5-recipe.yaml.
--- quant_stage: quant_modifiers: qwen35Day0Modifier: ignore: ["re:.*lm_head", "re:.*embed_tokens", "re:visual.*", "re:model.visual.*", "re:.*mlp.shared_expert_gate$", "re:.*mlp.gate$", "re:.*conv1d$", "re:.*in_proj_a$", "re:.*in_proj_b$", "re:.*fc$", "re:.*pre_fc_norm_embedding$", "re:.*pre_fc_norm_hidden$"] scheme: W8A8 ...To perform W8A8-INT8 quantization, specify the path to the original model and the save path for the quantized model.
qlean --model_name Qwen/Qwen3.5 --model_path /path/to/Qwen3.5/ --save_path /path/to/Qwen3.5-INT8/ --recipe /path/to/Qwen3.5-recipe.yaml
2. W8A8-INT8 quantization for GLM-5
Create GLM-5-recipe.yaml.
--- quant_stage: quant_modifiers: generalDay0Modifier: ignore: ["re:.*lm_head", "re:.*embed_tokens", "re:.*mlp.gate$", "re:.*mlp.shared_expert_gate$", "re:.*linear_attn.*", "re:.*self_attn.*"] scheme: W8A8 ...To perform W8A8-INT8 quantization, specify the path to the original model and the save path for the quantized model.
qlean --model_name zai-org/GLM-5 --model_path /path/to/GLM-5/ --save_path /path/to/GLM-5-INT8/ --recipe /path/to/GLM-5-recipe.yaml
Is this page helpful?