From the model card
Fields below are copied from the tags and counters on the HuggingFace repository EleutherAI/gpt-neo-125m at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- EleutherAI
- Pipeline tag
- text-generation
- Library
- Transformers
- Framework tags
- PyTorch, JAX, Rust (candle)
- Weight formats
- safetensors
- License tag
mit— read the license file in the repo before relying on it- Language tags
- English (en)
- Papers cited
- arXiv:2101.00027
- Datasets declared
- EleutherAI/pile
- Downloads (HF counter at last fetch)
- 502,268
- Likes (HF counter at last fetch)
- 229
- Model card
- https://huggingface.co/EleutherAI/gpt-neo-125m
Use cases
- Teaching GPT-style language model concepts
- Baseline comparison for sub-200M autoregressive models
- Pretraining recipe research using Pile dataset
- Fast iteration on fine-tuning pipelines before scaling
Pros
- MIT license
- Multiple framework exports: PyTorch, JAX, Rust, safetensors
- Well-documented training and architecture details from EleutherAI
- Pile training data gives broad English coverage
Cons
- 125M parameters — not competitive with modern models on any practical task
- Trained on an older corpus (Pile v1) with known quality and bias issues
- Outperformed on most tasks by Qwen2-0.5B at a smaller parameter count
- Causal LM format requires careful prompt engineering for task-specific use
Tags
transformerspytorchjaxrustsafetensorsgpt_neotext-generationtext generationcausal-lmendataset:EleutherAI/pilearxiv:2101.00027license:mitendpoints_compatibleregion:usdeploy:sagemakerdeploy:azure