Llava

orchestra-research/ai-research-skills/18-multimodal/llava

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

AI 產生的概覽

指導使用 LLaVA 視覺語言模型進行影像對話、視覺問答與自訂訓練。

功能
說明如何安裝並執行開源 LLaVA 視覺語言模型,該模型將 CLIP 視覺編碼器與 Vicuna 或 LLaMA 語言模型結合。內容涵蓋載入預訓練權重、單張影像與多輪影像對話、影像描述、視覺問答、文件理解、命令列與 Gradio 介面、4 位元與 8 位元量化,以及自訂模型的兩階段訓練。也列出模型規模、顯示記憶體與速度數據、基準分數、限制與框架整合方式。
適用情境
適合在建立視覺語言聊天機器人或影像理解功能、執行視覺問答或影像描述,以及在自訂影像指令資料上微調 LLaVA 模型時使用。也適合在顯示記憶體受限時選擇模型規模或量化等級。
執行需求
需要 Python 環境以及 transformers、torch 和 pillow,實際推論需要 GPU,並需要連網從 Hugging Face 下載模型權重。此技能僅為說明文件,不附帶指令碼;訓練指令引用上游 LLaVA 儲存庫中的指令碼。

LLaVA - Large Language and Vision Assistant

Open-source vision-language model for conversational image understanding.

When to use LLaVA

Use when:

  • Building vision-language chatbots
  • Visual question answering (VQA)
  • Image description and captioning
  • Multi-turn image conversations
  • Visual instruction following
  • Document understanding with images

Metrics:

  • 23,000+ GitHub stars
  • GPT-4V level capabilities (targeted)
  • Apache 2.0 License
  • Multiple model sizes (7B-34B params)

Use alternatives instead:

  • GPT-4V: Highest quality, API-based
  • CLIP: Simple zero-shot classification
  • BLIP-2: Better for captioning only
  • Flamingo: Research, not open-source

Quick start

Installation

bash
# Clone repositorygit clone https://github.com/haotian-liu/LLaVAcd LLaVA
# Installpip install -e .

Basic usage

python
from llava.model.builder import load_pretrained_modelfrom llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_tokenfrom llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKENfrom llava.conversation import conv_templatesfrom PIL import Imageimport torch
# Load modelmodel_path = "liuhaotian/llava-v1.5-7b"tokenizer, model, image_processor, context_len = load_pretrained_model(    model_path=model_path,    model_base=None,    model_name=get_model_name_from_path(model_path))
# Load imageimage = Image.open("image.jpg")image_tensor = process_images([image], image_processor, model.config)image_tensor = image_tensor.to(model.device, dtype=torch.float16)
# Create conversationconv = conv_templates["llava_v1"].copy()conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")conv.append_message(conv.roles[1], None)prompt = conv.get_prompt()
# Generate responseinput_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).to(model.device)
with torch.inference_mode():    output_ids = model.generate(        input_ids,        images=image_tensor,        do_sample=True,        temperature=0.2,        max_new_tokens=512    )
response = tokenizer.decode(output_ids[0], skip_special_tokens=True).strip()print(response)

Available models

ModelParametersVRAMQuality
LLaVA-v1.5-7B7B~14 GBGood
LLaVA-v1.5-13B13B~28 GBBetter
LLaVA-v1.6-34B34B~70 GBBest
python
# Load different modelsmodel_7b = "liuhaotian/llava-v1.5-7b"model_13b = "liuhaotian/llava-v1.5-13b"model_34b = "liuhaotian/llava-v1.6-34b"
# 4-bit quantization for lower VRAMload_4bit = True  # Reduces VRAM by ~4×

CLI usage

bash
# Single image querypython -m llava.serve.cli \    --model-path liuhaotian/llava-v1.5-7b \    --image-file image.jpg \    --query "What is in this image?"
# Multi-turn conversationpython -m llava.serve.cli \    --model-path liuhaotian/llava-v1.5-7b \    --image-file image.jpg# Then type questions interactively

Web UI (Gradio)

bash
# Launch Gradio interfacepython -m llava.serve.gradio_web_server \    --model-path liuhaotian/llava-v1.5-7b \    --load-4bit  # Optional: reduce VRAM
# Access at http://localhost:7860

Multi-turn conversations

python
# Initialize conversationconv = conv_templates["llava_v1"].copy()
# Turn 1conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")conv.append_message(conv.roles[1], None)response1 = generate(conv, model, image)  # "A dog playing in a park"
# Turn 2conv.messages[-1][1] = response1  # Add previous responseconv.append_message(conv.roles[0], "What breed is the dog?")conv.append_message(conv.roles[1], None)response2 = generate(conv, model, image)  # "Golden Retriever"
# Turn 3conv.messages[-1][1] = response2conv.append_message(conv.roles[0], "What time of day is it?")conv.append_message(conv.roles[1], None)response3 = generate(conv, model, image)

Common tasks

Image captioning

python
question = "Describe this image in detail."response = ask(model, image, question)

Visual question answering

python
question = "How many people are in the image?"response = ask(model, image, question)

Object detection (textual)

python
question = "List all the objects you can see in this image."response = ask(model, image, question)

Scene understanding

python
question = "What is happening in this scene?"response = ask(model, image, question)

Document understanding

python
question = "What is the main topic of this document?"response = ask(model, document_image, question)

Training custom model

bash
# Stage 1: Feature alignment (558K image-caption pairs)bash scripts/v1_5/pretrain.sh
# Stage 2: Visual instruction tuning (150K instruction data)bash scripts/v1_5/finetune.sh

Quantization (reduce VRAM)

python
# 4-bit quantizationtokenizer, model, image_processor, context_len = load_pretrained_model(    model_path="liuhaotian/llava-v1.5-13b",    model_base=None,    model_name=get_model_name_from_path("liuhaotian/llava-v1.5-13b"),    load_4bit=True  # Reduces VRAM ~4×)
# 8-bit quantizationload_8bit=True  # Reduces VRAM ~2×

Best practices

  1. Start with 7B model - Good quality, manageable VRAM
  2. Use 4-bit quantization - Reduces VRAM significantly
  3. GPU required - CPU inference extremely slow
  4. Clear prompts - Specific questions get better answers
  5. Multi-turn conversations - Maintain conversation context
  6. Temperature 0.2-0.7 - Balance creativity/consistency
  7. max_new_tokens 512-1024 - For detailed responses
  8. Batch processing - Process multiple images sequentially

Performance

ModelVRAM (FP16)VRAM (4-bit)Speed (tokens/s)
7B~14 GB~4 GB~20
13B~28 GB~8 GB~12
34B~70 GB~18 GB~5

On A100 GPU

Benchmarks

LLaVA achieves competitive scores on:

  • VQAv2: 78.5%
  • GQA: 62.0%
  • MM-Vet: 35.4%
  • MMBench: 64.3%

Limitations

  1. Hallucinations - May describe things not in image
  2. Spatial reasoning - Struggles with precise locations
  3. Small text - Difficulty reading fine print
  4. Object counting - Imprecise for many objects
  5. VRAM requirements - Need powerful GPU
  6. Inference speed - Slower than CLIP

Integration with frameworks

LangChain

python
from langchain.llms.base import LLM
class LLaVALLM(LLM):    def _call(self, prompt, stop=None):        # Custom LLaVA inference        return response
llm = LLaVALLM()

Gradio App

python
import gradio as gr
def chat(image, text, history):    response = ask_llava(model, image, text)    return response
demo = gr.ChatInterface(    chat,    additional_inputs=[gr.Image(type="pil")],    title="LLaVA Chat")demo.launch()

Resources

來源與署名

來源:orchestra-research/ai-research-skills位於18-multimodal/llava提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架