From the model card
Fields below are copied from the tags and counters on the HuggingFace repository llava-hf/llava-1.5-7b-hf at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- llava-hf
- Pipeline tag
- image-text-to-text
- Library
- Transformers
- Weight formats
- safetensors
- License tag
llama2— read the license file in the repo before relying on it- Language tags
- English (en)
- Datasets declared
- liuhaotian/LLaVA-Instruct-150K
- Downloads (HF counter at last fetch)
- 1,927,882
- Likes (HF counter at last fetch)
- 372
- Model card
- https://huggingface.co/llava-hf/llava-1.5-7b-hf
Use cases
- Visual question answering on natural images
- Image captioning and description generation
- Multimodal chat prototyping and experimentation
- Baseline for evaluating newer vision-language models
Pros
- Simple architecture makes it easy to understand and modify
- Strong performance on standard VQA benchmarks for its size
- Converted to HuggingFace Transformers format for easy loading
- Apache-2.0 licensed
Cons
- Superseded by LLaVA-1.6, Qwen2-VL, and InternVL at the same scale
- Single image input only — no video or multi-image context
- 336px crop resolution struggles with text-heavy or high-detail images
- MLP projection is brittle vs newer cross-attention vision connectors
Tags
transformerssafetensorsllavaimage-text-to-textvisionconversationalendataset:liuhaotian/LLaVA-Instruct-150Klicense:llama2endpoints_compatibleregion:us