AI Tools.

Search

image text to text by llava-hf

llava-1.5-7b-hf

LLaVA 1.5 7B connects a CLIP ViT-L/14@336 vision encoder to Vicuna 7B via a simple MLP projection. It was a state-of-the-art open multimodal model at release and remains widely used as a baseline for vision-language research.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository llava-hf/llava-1.5-7b-hf at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
llava-hf
Pipeline tag
image-text-to-text
Library
Transformers
Weight formats
safetensors
License tag
llama2 — read the license file in the repo before relying on it
Language tags
English (en)
Datasets declared
liuhaotian/LLaVA-Instruct-150K
Downloads (HF counter at last fetch)
1,927,882
Likes (HF counter at last fetch)
372
Model card
https://huggingface.co/llava-hf/llava-1.5-7b-hf

Use cases

  • Visual question answering on natural images
  • Image captioning and description generation
  • Multimodal chat prototyping and experimentation
  • Baseline for evaluating newer vision-language models

Pros

  • Simple architecture makes it easy to understand and modify
  • Strong performance on standard VQA benchmarks for its size
  • Converted to HuggingFace Transformers format for easy loading
  • Apache-2.0 licensed

Cons

  • Superseded by LLaVA-1.6, Qwen2-VL, and InternVL at the same scale
  • Single image input only — no video or multi-image context
  • 336px crop resolution struggles with text-heavy or high-detail images
  • MLP projection is brittle vs newer cross-attention vision connectors

Tags

transformerssafetensorsllavaimage-text-to-textvisionconversationalendataset:liuhaotian/LLaVA-Instruct-150Klicense:llama2endpoints_compatibleregion:us