AI Tools.

Search

text generation by stelterlab

Mistral-Small-24B-Instruct-2501-AWQ

An AWQ (Activation-aware Weight Quantization) conversion of Mistral Small 24B Instruct (January 2025), offering 4-bit quantized inference at reduced memory while preserving most of the original model's instruction-following quality.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository stelterlab/Mistral-Small-24B-Instruct-2501-AWQ at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
stelterlab
Pipeline tag
text-generation
Library
vLLM, Transformers
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Language tags
English (en), French (fr), German (de), Spanish (es), Italian (it), Portuguese (pt), Chinese (zh), Japanese (ja), Russian (ru), Korean (ko)
Downloads (HF counter at last fetch)
333,422
Likes (HF counter at last fetch)
29
Model card
https://huggingface.co/stelterlab/Mistral-Small-24B-Instruct-2501-AWQ

Use cases

  • Production serving of Mistral Small on 24 GB VRAM GPUs
  • API latency reduction vs BF16 via quantized inference
  • Cost-effective batch processing of instruction-following tasks
  • Edge server deployment of a capable 24B model

Pros

  • AWQ is one of the better-quality 4-bit quantization schemes for instruction models
  • 24B scale provides solid reasoning at a reasonable memory budget
  • Mistral models have well-documented evals and commercial track record
  • Compatible with vLLM and other AWQ-aware inference servers

Cons

  • AWQ quantization can cause accuracy drops on math and code vs BF16 baseline
  • 24 GB VRAM requirement still rules out consumer single-GPU setups
  • stelterlab conversion — not official Mistral release; weight fidelity unverified
  • No quantization-specific benchmark comparison provided in model card

Tags

vllmsafetensorsmistraltext-generationtransformersconversationalenfrdeesitptzhjarukobase_model:mistralai/Mistral-Small-24B-Instruct-2501base_model:quantized:mistralai/Mistral-Small-24B-Instruct-2501license:apache-2.0text-generation-inference