From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.6-35B-A3B-FP8 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model Qwen/Qwen3.6-35B-A3B
- quantized from Qwen/Qwen3.6-35B-A3B
- Downloads (HF counter at last fetch)
- 13,251,463
- Likes (HF counter at last fetch)
- 372
- Model card
- https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
Use cases
- Production serving on H100/H200 where FP8 hardware acceleration is available
- Fitting the 35B MoE model into fewer GPU cards
- Throughput-optimized batch inference workloads
- Memory-constrained deployments needing the full 35B parameter count
Pros
- Roughly 2x memory savings vs BF16 on FP8-capable hardware
- Maintained by Qwen team with official support
- Apache-2.0 licensed
- Compatible with vLLM's FP8 serving backend
Cons
- FP8 inference requires Hopper or newer GPU architecture
- Quality degradation measurable on tasks sensitive to numerical precision
- Narrower compatibility than GGUF for local deployment
- Less community vetting than standard BF16 checkpoints
Tags
transformerssafetensorsqwen3_5_moeimage-text-to-textconversationalbase_model:Qwen/Qwen3.6-35B-A3Bbase_model:quantized:Qwen/Qwen3.6-35B-A3Blicense:apache-2.0endpoints_compatiblefp8region:usdeploy:azure