From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.5-27B-GPTQ-Int4 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Downloads (HF counter at last fetch)
- 354,659
- Likes (HF counter at last fetch)
- 55
- Model card
- https://huggingface.co/Qwen/Qwen3.5-27B-GPTQ-Int4
Use cases
- Multimodal image+text inference at reduced VRAM with official GPTQ
- Self-hosted vision-language assistant on RTX 4090 or A100
- vLLM serving with GPTQ INT4 quantization
- Comparing GPTQ vs AWQ accuracy tradeoffs for Qwen3.5-27B
Pros
- Apache-2.0 license
- Official Alibaba quantization — highest confidence in calibration quality
- GPTQ INT4 brings 27B into RTX 4090 range
- Transformers-compatible
Cons
- INT4 quantization introduces accuracy regression on vision-language tasks
- 27B GPTQ INT4 still needs ~15 GB VRAM — not entry-level
- GPTQ is generally slower than AWQ at inference for equivalent bit-width
- Vision quality in quantized multimodal models degrades faster than text-only tasks
Tags
transformerssafetensorsqwen3_5image-text-to-textconversationallicense:apache-2.0endpoints_compatible4-bitgptqdeploy:azureregion:us