From the model card
Fields below are copied from the tags and counters on the HuggingFace repository lmstudio-community/LFM2-24B-A2B-MLX-5bit at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- lmstudio-community
- Pipeline tag
- text-generation
- Library
- Transformers, MLX
- Weight formats
- safetensors
- License tag
other— read the license file in the repo before relying on it- Lineage
-
- base model LiquidAI/LFM2-24B-A2B
- quantized from LiquidAI/LFM2-24B-A2B
- Language tags
- English (en), Arabic (ar), Chinese (zh), French (fr), German (de), Japanese (ja), Korean (ko), Spanish (es), Portuguese (pt)
- Downloads (HF counter at last fetch)
- 317,034
- Likes (HF counter at last fetch)
- 1
- Model card
- https://huggingface.co/lmstudio-community/LFM2-24B-A2B-MLX-5bit
Use cases
- Balanced local inference of LFM2-24B where memory is limited but quality matters
- Comparing quantization levels for optimal quality-memory tradeoff
- On-device AI workloads targeting M2 Max or M3 Max class hardware
Pros
- 5-bit quantization noticeably recovers accuracy over 4-bit on nuanced tasks
- Lower memory than 8-bit while retaining most quality gains
- MLX native acceleration on Apple Silicon
- MoE architecture means active compute remains low per token
Cons
- 5-bit is a less standard quantization level — fewer community resources
- ~22 GB unified memory needed — still a high-end Mac requirement
- MLX-only; no cross-platform use
- Community conversion with no formal accuracy delta benchmark
Tags
transformerssafetensorslfm2_moetext-generationliquidlfm2edgemlxconversationalenarzhfrdejakoesptbase_model:LiquidAI/LFM2-24B-A2Bbase_model:quantized:LiquidAI/LFM2-24B-A2B