Clip

orchestra-research/ai-research-skills/18-multimodal/clip

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

僅含說明AI & Agents
AI 產生的概覽

說明如何使用 OpenAI 的 CLIP 模型進行零樣本影像分類、圖文匹配、語意影像搜尋與內容審核。

功能
說明如何載入 CLIP 檢查點(RN50、ViT-B/32、ViT-B/16、ViT-L/14),並將影像與文字編碼為正規化向量。提供零樣本分類、圖文相似度評分、以索引進行的語意影像搜尋、內容審核類別、批次處理,以及把向量存入 Chroma 或 FAISS 等向量資料庫的程式碼。同時列出最佳實務、效能數據與模型限制。
適用情境
適用於不需微調、以自然語言標籤對影像進行分類或檢索、圖文匹配、建立語意影像搜尋,或篩檢不安全影像內容的情境。它較適合廣泛類別而非細緻辨識;文件建議在影像描述、視覺對話或分割工作上改用 BLIP-2、LLaVA 或 Segment Anything。
執行需求
需要 Python 環境以及 transformers、torch、pillow(安裝範例中還包含 torchvision、ftfy、regex、tqdm);CLIP 權重需從網路下載;建議使用 GPU,但並非必要。僅為說明文件,未附帶指令碼。

CLIP - Contrastive Language-Image Pre-Training

OpenAI's model that understands images from natural language.

When to use CLIP

Use when:

  • Zero-shot image classification (no training data needed)
  • Image-text similarity/matching
  • Semantic image search
  • Content moderation (detect NSFW, violence)
  • Visual question answering
  • Cross-modal retrieval (image→text, text→image)

Metrics:

  • 25,300+ GitHub stars
  • Trained on 400M image-text pairs
  • Matches ResNet-50 on ImageNet (zero-shot)
  • MIT License

Use alternatives instead:

  • BLIP-2: Better captioning
  • LLaVA: Vision-language chat
  • Segment Anything: Image segmentation

Quick start

Installation

bash
pip install git+https://github.com/openai/CLIP.gitpip install torch torchvision ftfy regex tqdm

Zero-shot classification

python
import torchimport clipfrom PIL import Image
# Load modeldevice = "cuda" if torch.cuda.is_available() else "cpu"model, preprocess = clip.load("ViT-B/32", device=device)
# Load imageimage = preprocess(Image.open("photo.jpg")).unsqueeze(0).to(device)
# Define possible labelstext = clip.tokenize(["a dog", "a cat", "a bird", "a car"]).to(device)
# Compute similaritywith torch.no_grad():    image_features = model.encode_image(image)    text_features = model.encode_text(text)
    # Cosine similarity    logits_per_image, logits_per_text = model(image, text)    probs = logits_per_image.softmax(dim=-1).cpu().numpy()
# Print resultslabels = ["a dog", "a cat", "a bird", "a car"]for label, prob in zip(labels, probs[0]):    print(f"{label}: {prob:.2%}")

Available models

python
# Models (sorted by size)models = [    "RN50",           # ResNet-50    "RN101",          # ResNet-101    "ViT-B/32",       # Vision Transformer (recommended)    "ViT-B/16",       # Better quality, slower    "ViT-L/14",       # Best quality, slowest]
model, preprocess = clip.load("ViT-B/32")
ModelParametersSpeedQuality
RN50102MFastGood
ViT-B/32151MMediumBetter
ViT-L/14428MSlowBest

Image-text similarity

python
# Compute embeddingsimage_features = model.encode_image(image)text_features = model.encode_text(text)
# Normalizeimage_features /= image_features.norm(dim=-1, keepdim=True)text_features /= text_features.norm(dim=-1, keepdim=True)
# Cosine similaritysimilarity = (image_features @ text_features.T).item()print(f"Similarity: {similarity:.4f}")

Semantic image search

python
# Index imagesimage_paths = ["img1.jpg", "img2.jpg", "img3.jpg"]image_embeddings = []
for img_path in image_paths:    image = preprocess(Image.open(img_path)).unsqueeze(0).to(device)    with torch.no_grad():        embedding = model.encode_image(image)        embedding /= embedding.norm(dim=-1, keepdim=True)    image_embeddings.append(embedding)
image_embeddings = torch.cat(image_embeddings)
# Search with text queryquery = "a sunset over the ocean"text_input = clip.tokenize([query]).to(device)with torch.no_grad():    text_embedding = model.encode_text(text_input)    text_embedding /= text_embedding.norm(dim=-1, keepdim=True)
# Find most similar imagessimilarities = (text_embedding @ image_embeddings.T).squeeze(0)top_k = similarities.topk(3)
for idx, score in zip(top_k.indices, top_k.values):    print(f"{image_paths[idx]}: {score:.3f}")

Content moderation

python
# Define categoriescategories = [    "safe for work",    "not safe for work",    "violent content",    "graphic content"]
text = clip.tokenize(categories).to(device)
# Check imagewith torch.no_grad():    logits_per_image, _ = model(image, text)    probs = logits_per_image.softmax(dim=-1)
# Get classificationmax_idx = probs.argmax().item()max_prob = probs[0, max_idx].item()
print(f"Category: {categories[max_idx]} ({max_prob:.2%})")

Batch processing

python
# Process multiple imagesimages = [preprocess(Image.open(f"img{i}.jpg")) for i in range(10)]images = torch.stack(images).to(device)
with torch.no_grad():    image_features = model.encode_image(images)    image_features /= image_features.norm(dim=-1, keepdim=True)
# Batch texttexts = ["a dog", "a cat", "a bird"]text_tokens = clip.tokenize(texts).to(device)
with torch.no_grad():    text_features = model.encode_text(text_tokens)    text_features /= text_features.norm(dim=-1, keepdim=True)
# Similarity matrix (10 images × 3 texts)similarities = image_features @ text_features.Tprint(similarities.shape)  # (10, 3)

Integration with vector databases

python
# Store CLIP embeddings in Chroma/FAISSimport chromadb
client = chromadb.Client()collection = client.create_collection("image_embeddings")
# Add image embeddingsfor img_path, embedding in zip(image_paths, image_embeddings):    collection.add(        embeddings=[embedding.cpu().numpy().tolist()],        metadatas=[{"path": img_path}],        ids=[img_path]    )
# Query with textquery = "a sunset"text_embedding = model.encode_text(clip.tokenize([query]))results = collection.query(    query_embeddings=[text_embedding.cpu().numpy().tolist()],    n_results=5)

Best practices

  1. Use ViT-B/32 for most cases - Good balance
  2. Normalize embeddings - Required for cosine similarity
  3. Batch processing - More efficient
  4. Cache embeddings - Expensive to recompute
  5. Use descriptive labels - Better zero-shot performance
  6. GPU recommended - 10-50× faster
  7. Preprocess images - Use provided preprocess function

Performance

OperationCPUGPU (V100)
Image encoding~200ms~20ms
Text encoding~50ms~5ms
Similarity compute<1ms<1ms

Limitations

  1. Not for fine-grained tasks - Best for broad categories
  2. Requires descriptive text - Vague labels perform poorly
  3. Biased on web data - May have dataset biases
  4. No bounding boxes - Whole image only
  5. Limited spatial understanding - Position/counting weak

Resources

來源與署名

來源:orchestra-research/ai-research-skills位於18-multimodal/clip提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架