AI Tools.

Search

image text to text by cyankiwi

Qwen3.5-9B-AWQ-4bit

AWQ 4-bit quantization of Qwen3.5-9B, a dense image-text-to-text model. At 9B parameters with AWQ INT4, inference requires roughly 6-8 GB VRAM, placing it within reach of RTX 3080/4070-class cards. compressed-tensors format is vLLM-native.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository cyankiwi/Qwen3.5-9B-AWQ-4bit at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
cyankiwi
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Downloads (HF counter at last fetch)
363,295
Likes (HF counter at last fetch)
36
Model card
https://huggingface.co/cyankiwi/Qwen3.5-9B-AWQ-4bit

Use cases

  • Multimodal image-text inference on mid-range consumer GPUs
  • vLLM-hosted vision-language assistant
  • Visual question answering without cloud dependency
  • Comparative evaluation of dense vs MoE Qwen3.5 variants

Pros

  • 9B AWQ INT4 fits in ~7 GB VRAM — broad consumer GPU compatibility
  • Multimodal image-text capability in a self-hostable package
  • Apache-2.0 license
  • compressed-tensors for clean vLLM integration

Cons

  • Community quantization — no linked accuracy regression report
  • Vision tasks are more sensitive to INT4 quantization than pure text
  • Dense 9B has higher active FLOP count than MoE alternatives at same param count
  • No GGUF variant from this uploader for llama.cpp users

Tags

transformerssafetensorsqwen3_5image-text-to-textconversationalbase_model:Qwen/Qwen3.5-9Bbase_model:quantized:Qwen/Qwen3.5-9Blicense:apache-2.0endpoints_compatiblecompressed-tensorsregion:us