From the model card
Fields below are copied from the tags and counters on the HuggingFace repository google/flan-t5-base at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- Library
- Transformers
- Framework tags
- PyTorch, TensorFlow, JAX
- Weight formats
- safetensors
- License tag
apache-2.0— read the license file in the repo before relying on it- Language tags
- multilingual; English (en), French (fr), Romanian (ro), German (de)
- Papers cited
- arXiv:2210.11416, arXiv:1910.09700
- Datasets declared
- svakulenk0/qrecc, taskmaster2, djaym7/wiki_dialog, deepmind/code_contests, lambada, gsm8k, aqua_rat, esnli and 2 more on the model card
- Downloads (HF counter at last fetch)
- 1,464,241
- Likes (HF counter at last fetch)
- 1,090
- Model card
- https://huggingface.co/google/flan-t5-base
Use cases
- Zero-shot text classification and question answering
- Lightweight instruction-following generation
- Summarization and translation with simple prompts
- Legacy systems already using T5 that want instruction-following capability
Pros
- Strong zero-shot performance for a 220M-parameter model
- Apache-2.0 licensed
- Instruction-tuned on 1800+ tasks — broad task coverage
- Efficient encoder-decoder architecture for seq2seq tasks
Cons
- Outperformed by decoder-only models of similar size on generation tasks
- Limited context window (512 encoder, 512 decoder tokens)
- Not suitable for complex multi-step reasoning
- Flan-T5-large provides meaningful improvement if compute allows
Tags
transformerspytorchtfjaxsafetensorst5text2text-generationenfrrodemultilingualdataset:svakulenk0/qreccdataset:taskmaster2dataset:djaym7/wiki_dialogdataset:deepmind/code_contestsdataset:lambadadataset:gsm8kdataset:aqua_ratdataset:esnli