AI Tools.

Search

image text to text by Qwen

Qwen3.5-27B

Qwen 3.5 27B is a dense image-text-to-text model from Alibaba, positioned between the 14B and 72B variants for users who need more capacity than 14B but can't serve 72B. It handles both vision and language instructions.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.5-27B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
Qwen
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Downloads (HF counter at last fetch)
2,369,643
Likes (HF counter at last fetch)
1,040
Model card
https://huggingface.co/Qwen/Qwen3.5-27B

Use cases

  • Multimodal document and chart analysis
  • Complex multi-step reasoning with image context
  • High-quality bilingual Chinese-English content generation
  • Fine-tuning base for domain-specific vision-language tasks

Pros

  • Apache-2.0 licensed
  • Supports both image and text input natively
  • Larger capacity than 14B class models for knowledge-intensive tasks
  • Strong multilingual performance from Alibaba's training data

Cons

  • 27B in bfloat16 requires ~54GB VRAM — needs multi-GPU or quantization
  • Instruction following can be verbose without explicit length constraints
  • Benchmark results show quality gaps vs GPT-4o on spatial reasoning
  • Limited third-party quantization support compared to Llama/Mistral

Tags

transformerssafetensorsqwen3_5image-text-to-textconversationallicense:apache-2.0eval-resultsendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure