AI Tools.

Search

image text to text by Qwen

Qwen3.6-27B-FP8

FP8-quantized version of Qwen 3.6 27B for H100/H200 serving. Reduces memory from ~54GB (BF16) to approximately 27GB while maintaining near-BF16 quality on most benchmarks for a dense multimodal model.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.6-27B-FP8 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Downloads (HF counter at last fetch)
7,537,356
Likes (HF counter at last fetch)
353
Model card
https://huggingface.co/Qwen/Qwen3.6-27B-FP8

Use cases

  • Serving Qwen3.6-27B on a single 40GB A100 or H100
  • Throughput-optimized batch inference on FP8-capable hardware
  • Production deployment where BF16 27B doesn't fit single GPU
  • Benchmarking FP8 vs BF16 quality trade-offs

Pros

  • Halves VRAM requirement vs BF16 on FP8 hardware
  • Minimal benchmark regression on standard evaluations
  • Apache-2.0 licensed
  • Compatible with vLLM FP8 serving

Cons

  • FP8 requires Hopper (H100/H200) or Ada Lovelace GPU
  • Accuracy degradation on precision-sensitive arithmetic tasks
  • Less tested than GGUF quantization for general community use
  • Not a substitute for BF16 in fine-tuning scenarios

Tags

transformerssafetensorsqwen3_5image-text-to-textconversationalbase_model:Qwen/Qwen3.6-27Bbase_model:quantized:Qwen/Qwen3.6-27Blicense:apache-2.0endpoints_compatiblefp8region:usdeploy:azuredeploy:sagemaker