AI Tools.

Search

text generation by datajuicer

LLaMA-1B-dj-refine-150B

LLaMA-1B fine-tuned on 150B tokens of RedPajama data filtered and refined by Data-Juicer, a data-cleaning toolkit from Alibaba DAMO. The training corpus was pruned using quality heuristics across Wikipedia, arXiv, Books, and Common Crawl slices. At 1B parameters it trades capability for low inference cost.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository datajuicer/LLaMA-1B-dj-refine-150B at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
datajuicer
Pipeline tag
text-generation
Library
Transformers
Framework tags
PyTorch
License tag
apache-2.0 — read the license file in the repo before relying on it
Papers cited
arXiv:2309.02033
Datasets declared
datajuicer/redpajama-wiki-refined-by-data-juicer, datajuicer/redpajama-arxiv-refined-by-data-juicer, datajuicer/redpajama-c4-refined-by-data-juicer, datajuicer/redpajama-book-refined-by-data-juicer, datajuicer/redpajama-cc-2019-30-refined-by-data-juicer, datajuicer/redpajama-cc-2020-05-refined-by-data-juicer, datajuicer/redpajama-cc-2021-04-refined-by-data-juicer, datajuicer/redpajama-cc-2022-05-refined-by-data-juicer and 11 more on the model card
Downloads (HF counter at last fetch)
1,297,632
Likes (HF counter at last fetch)
3
Model card
https://huggingface.co/datajuicer/LLaMA-1B-dj-refine-150B

Use cases

  • Edge inference on CPU or low-VRAM devices
  • Benchmarking data-cleaning pipelines for LLM pretraining
  • Studying effect of data quality on small-model perplexity
  • Prototype text-generation features before scaling up

Pros

  • Apache-2.0 license — no usage restrictions
  • Tiny footprint fits in under 2 GB RAM at fp16
  • Training data lineage is documented via Data-Juicer repo
  • Compatible with standard Transformers text-generation pipeline

Cons

  • 1B parameters produces noticeably weaker reasoning than 7B+ models
  • Instruction-following is absent — requires fine-tuning for chat use
  • Outperformed on most benchmarks by Qwen2-0.5B despite smaller size
  • No GGUF or quantized variants from the original author

Tags

transformerspytorchllamatext-generationdataset:datajuicer/redpajama-wiki-refined-by-data-juicerdataset:datajuicer/redpajama-arxiv-refined-by-data-juicerdataset:datajuicer/redpajama-c4-refined-by-data-juicerdataset:datajuicer/redpajama-book-refined-by-data-juicerdataset:datajuicer/redpajama-cc-2019-30-refined-by-data-juicerdataset:datajuicer/redpajama-cc-2020-05-refined-by-data-juicerdataset:datajuicer/redpajama-cc-2021-04-refined-by-data-juicerdataset:datajuicer/redpajama-cc-2022-05-refined-by-data-juicerdataset:datajuicer/redpajama-cc-2023-06-refined-by-data-juicerdataset:datajuicer/redpajama-pile-stackexchange-refined-by-data-juicerdataset:datajuicer/redpajama-stack-code-refined-by-data-juicerdataset:datajuicer/the-pile-nih-refined-by-data-juicerdataset:datajuicer/the-pile-europarl-refined-by-data-juicerdataset:datajuicer/the-pile-philpaper-refined-by-data-juicerdataset:datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicerdataset:datajuicer/the-pile-pubmed-central-refined-by-data-juicer