From the model card
Fields below are copied from the tags and counters on the HuggingFace repository stelterlab/Mistral-Small-24B-Instruct-2501-AWQ at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- stelterlab
- Pipeline tag
- text-generation
- Library
- vLLM, Transformers
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Lineage
-
- base model mistralai/Mistral-Small-24B-Instruct-2501
- quantized from mistralai/Mistral-Small-24B-Instruct-2501
- Language tags
- English (en), French (fr), German (de), Spanish (es), Italian (it), Portuguese (pt), Chinese (zh), Japanese (ja), Russian (ru), Korean (ko)
- Downloads (HF counter at last fetch)
- 333,422
- Likes (HF counter at last fetch)
- 29
- Model card
- https://huggingface.co/stelterlab/Mistral-Small-24B-Instruct-2501-AWQ
Use cases
- Production serving of Mistral Small on 24 GB VRAM GPUs
- API latency reduction vs BF16 via quantized inference
- Cost-effective batch processing of instruction-following tasks
- Edge server deployment of a capable 24B model
Pros
- AWQ is one of the better-quality 4-bit quantization schemes for instruction models
- 24B scale provides solid reasoning at a reasonable memory budget
- Mistral models have well-documented evals and commercial track record
- Compatible with vLLM and other AWQ-aware inference servers
Cons
- AWQ quantization can cause accuracy drops on math and code vs BF16 baseline
- 24 GB VRAM requirement still rules out consumer single-GPU setups
- stelterlab conversion — not official Mistral release; weight fidelity unverified
- No quantization-specific benchmark comparison provided in model card
Tags
vllmsafetensorsmistraltext-generationtransformersconversationalenfrdeesitptzhjarukobase_model:mistralai/Mistral-Small-24B-Instruct-2501base_model:quantized:mistralai/Mistral-Small-24B-Instruct-2501license:apache-2.0text-generation-inference