From the model card
Fields below are copied from the tags and counters on the HuggingFace repository deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- deepseek-ai
- Pipeline tag
- text-generation
- Library
- Transformers
- Weight formats
- safetensors
- License tag
mit— read the license file in the repo before relying on it- Papers cited
- arXiv:2501.12948
- Downloads (HF counter at last fetch)
- 453,242
- Likes (HF counter at last fetch)
- 1,571
- Model card
- https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Use cases
- Chain-of-thought reasoning at minimal compute cost
- Edge deployment where reasoning traces are needed
- Benchmarking reasoning distillation at the 1.5B scale
- Teaching environments demonstrating step-by-step problem solving
Pros
- MIT license — full commercial freedom
- 1.5B size enables CPU inference or very low-VRAM GPUs
- Reasoning trace format is consistent with R1's training signal
- Transformers and TGI compatible
Cons
- 1.5B heavily constrains the depth of reasoning — short or simple problems only
- Distilled reasoning can be verbose without being correct on harder tasks
- Math and coding capability significantly below 7B+ R1 distillations
- Generates long reasoning traces that inflate token costs
Tags
transformerssafetensorsqwen2text-generationconversationalarxiv:2501.12948license:mittext-generation-inferenceendpoints_compatibleregion:usdeploy:sagemaker