QLean model quantization tool (v0.1.0 post1)

Updated at:
Copy as MD

Installation

The tool is a wheel package that can be installed using pip.

# Install dependencies
pip install triton_kernel==1.0.0+ppu2.0.0.oe
# Install qlean
pip install qlean==0.1.0+ppu2.0.0post1

Usage

1. W8A8-INT8 quantization for Qwen3.5

  1. Create Qwen3.5-recipe.yaml.

    ---
    quant_stage:
      quant_modifiers:
        qwen35Day0Modifier:
          ignore: ["re:.*lm_head", "re:.*embed_tokens", "re:visual.*", "re:model.visual.*", "re:.*mlp.shared_expert_gate$", "re:.*mlp.gate$", "re:.*conv1d$", "re:.*in_proj_a$", "re:.*in_proj_b$", "re:.*fc$", "re:.*pre_fc_norm_embedding$", "re:.*pre_fc_norm_hidden$"]
          scheme: W8A8
    ...
  2. To perform W8A8-INT8 quantization, specify the path to the original model and the save path for the quantized model.

    qlean --model_name Qwen/Qwen3.5 --model_path /path/to/Qwen3.5/ --save_path /path/to/Qwen3.5-INT8/ --recipe /path/to/Qwen3.5-recipe.yaml

2. W8A8-INT8 quantization for GLM-5

  1. Create GLM-5-recipe.yaml.

    ---
    quant_stage:
      quant_modifiers:
        generalDay0Modifier:
          ignore: ["re:.*lm_head", "re:.*embed_tokens", "re:.*mlp.gate$", "re:.*mlp.shared_expert_gate$", "re:.*linear_attn.*", "re:.*self_attn.*"]
          scheme: W8A8
    ...
  2. To perform W8A8-INT8 quantization, specify the path to the original model and the save path for the quantized model.

    qlean --model_name zai-org/GLM-5 --model_path /path/to/GLM-5/ --save_path /path/to/GLM-5-INT8/ --recipe /path/to/GLM-5-recipe.yaml