From the model card
Fields below are copied from the tags and counters on the HuggingFace repository Qwen/Qwen3.5-35B-A3B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Qwen
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model Qwen/Qwen3.5-35B-A3B-Base
- fine-tune of Qwen/Qwen3.5-35B-A3B-Base
- Downloads (HF counter at last fetch)
- 2,362,260
- Likes (HF counter at last fetch)
- 1,498
- Model card
- https://huggingface.co/Qwen/Qwen3.5-35B-A3B
Use cases
- Multimodal document processing combining text and image understanding
- Visual question answering with large effective model capacity
- OCR and chart interpretation at MoE-reduced inference cost
- Cost-efficient deployment where a dense 35B model would be impractical
Pros
- ~3B active parameters per token reduces actual inference compute significantly
- Apache 2.0 license permits commercial use without restrictions
- Multimodal capability spans both text and image input modalities
Cons
- MoE router complexity increases memory bandwidth requirements at inference
- 35B total weights require substantial storage and host RAM for loading
- Less community tooling and fine-tuning coverage than dense Qwen2.5 variants
Tags
transformerssafetensorsqwen3_5_moeimage-text-to-textconversationalbase_model:Qwen/Qwen3.5-35B-A3B-Basebase_model:finetune:Qwen/Qwen3.5-35B-A3B-Baselicense:apache-2.0eval-resultsendpoints_compatibledeploy:azureregion:us