AI Tools.

Search

image text to text by Qwen

Qwen3.6-35B-A3B-FP8

FP8-quantized version of Qwen3.6-35B-A3B for deployment on hardware with FP8 support (H100/H200). Reduces memory footprint and inference latency compared to BF16 with minimal quality degradation on most benchmarks.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.6-35B-A3B-FP8 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Downloads (HF counter at last fetch)
13,251,463
Likes (HF counter at last fetch)
372
Model card
https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8

Use cases

  • Production serving on H100/H200 where FP8 hardware acceleration is available
  • Fitting the 35B MoE model into fewer GPU cards
  • Throughput-optimized batch inference workloads
  • Memory-constrained deployments needing the full 35B parameter count

Pros

  • Roughly 2x memory savings vs BF16 on FP8-capable hardware
  • Maintained by Qwen team with official support
  • Apache-2.0 licensed
  • Compatible with vLLM's FP8 serving backend

Cons

  • FP8 inference requires Hopper or newer GPU architecture
  • Quality degradation measurable on tasks sensitive to numerical precision
  • Narrower compatibility than GGUF for local deployment
  • Less community vetting than standard BF16 checkpoints

Tags

transformerssafetensorsqwen3_5_moeimage-text-to-textconversationalbase_model:Qwen/Qwen3.6-35B-A3Bbase_model:quantized:Qwen/Qwen3.6-35B-A3Blicense:apache-2.0endpoints_compatiblefp8region:usdeploy:azure