From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.6-27B-FP8 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model Qwen/Qwen3.6-27B
- quantized from Qwen/Qwen3.6-27B
- Downloads (HF counter at last fetch)
- 7,537,356
- Likes (HF counter at last fetch)
- 353
- Model card
- https://huggingface.co/Qwen/Qwen3.6-27B-FP8
Use cases
- Serving Qwen3.6-27B on a single 40GB A100 or H100
- Throughput-optimized batch inference on FP8-capable hardware
- Production deployment where BF16 27B doesn't fit single GPU
- Benchmarking FP8 vs BF16 quality trade-offs
Pros
- Halves VRAM requirement vs BF16 on FP8 hardware
- Minimal benchmark regression on standard evaluations
- Apache-2.0 licensed
- Compatible with vLLM FP8 serving
Cons
- FP8 requires Hopper (H100/H200) or Ada Lovelace GPU
- Accuracy degradation on precision-sensitive arithmetic tasks
- Less tested than GGUF quantization for general community use
- Not a substitute for BF16 in fine-tuning scenarios
Tags
transformerssafetensorsqwen3_5image-text-to-textconversationalbase_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Blicense:apache-2.0endpoints_compatiblefp8region:usdeploy:azuredeploy:sagemaker