Chroma

orchestra-research/ai-research-skills/15-rag/chroma

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects.

僅含說明AI & Agents
AI 產生的概覽

說明如何使用開源嵌入資料庫 Chroma 進行向量檢索、中繼資料篩選與 RAG 檢索。

功能
說明如何安裝與使用 Chroma 這套面向 AI 應用的開源嵌入資料庫。內容涵蓋建立集合、加入帶中繼資料與嵌入的檔案、相似度查詢、中繼資料篩選、更新與刪除記錄,以及持久化到磁碟或執行伺服器模式。也介紹嵌入函式選項以及與 LangChain 和 LlamaIndex 的整合。
適用情境
適合在建立檢索增強生成、語意搜尋或檔案檢索功能,並希望使用本機或自架向量儲存時使用。既適用於在 notebook 中做原型,也適用於切換到伺服器模式用於正式環境。不適用於託管雲端向量資料庫,也不適用於沒有中繼資料的純相似度檢索。
執行需求
需要 Python 與 chromadb 套件(預設嵌入函式另需 sentence-transformers),或 JavaScript/TypeScript 與 chromadb 及 @chroma-core/default-embed。使用 OpenAI 或 HuggingFace 嵌入函式時可選用 API 金鑰。僅為說明文件,未附帶指令碼。

Chroma - Open-Source Embedding Database

The AI-native database for building LLM applications with memory.

When to use Chroma

Use Chroma when:

  • Building RAG (retrieval-augmented generation) applications
  • Need local/self-hosted vector database
  • Want open-source solution (Apache 2.0)
  • Prototyping in notebooks
  • Semantic search over documents
  • Storing embeddings with metadata

Metrics:

  • 24,300+ GitHub stars
  • 1,900+ forks
  • v1.3.3 (stable, weekly releases)
  • Apache 2.0 license

Use alternatives instead:

  • Pinecone: Managed cloud, auto-scaling
  • FAISS: Pure similarity search, no metadata
  • Weaviate: Production ML-native database
  • Qdrant: High performance, Rust-based

Quick start

Installation

bash
# Pythonpip install chromadb
# JavaScript/TypeScriptnpm install chromadb @chroma-core/default-embed

Basic usage (Python)

python
import chromadb
# Create clientclient = chromadb.Client()
# Create collectioncollection = client.create_collection(name="my_collection")
# Add documentscollection.add(    documents=["This is document 1", "This is document 2"],    metadatas=[{"source": "doc1"}, {"source": "doc2"}],    ids=["id1", "id2"])
# Queryresults = collection.query(    query_texts=["document about topic"],    n_results=2)
print(results)

Core operations

1. Create collection

python
# Simple collectioncollection = client.create_collection("my_docs")
# With custom embedding functionfrom chromadb.utils import embedding_functions
openai_ef = embedding_functions.OpenAIEmbeddingFunction(    api_key="your-key",    model_name="text-embedding-3-small")
collection = client.create_collection(    name="my_docs",    embedding_function=openai_ef)
# Get existing collectioncollection = client.get_collection("my_docs")
# Delete collectionclient.delete_collection("my_docs")

2. Add documents

python
# Add with auto-generated IDscollection.add(    documents=["Doc 1", "Doc 2", "Doc 3"],    metadatas=[        {"source": "web", "category": "tutorial"},        {"source": "pdf", "page": 5},        {"source": "api", "timestamp": "2025-01-01"}    ],    ids=["id1", "id2", "id3"])
# Add with custom embeddingscollection.add(    embeddings=[[0.1, 0.2, ...], [0.3, 0.4, ...]],    documents=["Doc 1", "Doc 2"],    ids=["id1", "id2"])

3. Query (similarity search)

python
# Basic queryresults = collection.query(    query_texts=["machine learning tutorial"],    n_results=5)
# Query with filtersresults = collection.query(    query_texts=["Python programming"],    n_results=3,    where={"source": "web"})
# Query with metadata filtersresults = collection.query(    query_texts=["advanced topics"],    where={        "$and": [            {"category": "tutorial"},            {"difficulty": {"$gte": 3}}        ]    })
# Access resultsprint(results["documents"])      # List of matching documentsprint(results["metadatas"])      # Metadata for each docprint(results["distances"])      # Similarity scoresprint(results["ids"])            # Document IDs

4. Get documents

python
# Get by IDsdocs = collection.get(    ids=["id1", "id2"])
# Get with filtersdocs = collection.get(    where={"category": "tutorial"},    limit=10)
# Get all documentsdocs = collection.get()

5. Update documents

python
# Update document contentcollection.update(    ids=["id1"],    documents=["Updated content"],    metadatas=[{"source": "updated"}])

6. Delete documents

python
# Delete by IDscollection.delete(ids=["id1", "id2"])
# Delete with filtercollection.delete(    where={"source": "outdated"})

Persistent storage

python
# Persist to diskclient = chromadb.PersistentClient(path="./chroma_db")
collection = client.create_collection("my_docs")collection.add(documents=["Doc 1"], ids=["id1"])
# Data persisted automatically# Reload later with same pathclient = chromadb.PersistentClient(path="./chroma_db")collection = client.get_collection("my_docs")

Embedding functions

Default (Sentence Transformers)

python
# Uses sentence-transformers by defaultcollection = client.create_collection("my_docs")# Default model: all-MiniLM-L6-v2

OpenAI

python
from chromadb.utils import embedding_functions
openai_ef = embedding_functions.OpenAIEmbeddingFunction(    api_key="your-key",    model_name="text-embedding-3-small")
collection = client.create_collection(    name="openai_docs",    embedding_function=openai_ef)

HuggingFace

python
huggingface_ef = embedding_functions.HuggingFaceEmbeddingFunction(    api_key="your-key",    model_name="sentence-transformers/all-mpnet-base-v2")
collection = client.create_collection(    name="hf_docs",    embedding_function=huggingface_ef)

Custom embedding function

python
from chromadb import Documents, EmbeddingFunction, Embeddings
class MyEmbeddingFunction(EmbeddingFunction):    def __call__(self, input: Documents) -> Embeddings:        # Your embedding logic        return embeddings
my_ef = MyEmbeddingFunction()collection = client.create_collection(    name="custom_docs",    embedding_function=my_ef)

Metadata filtering

python
# Exact matchresults = collection.query(    query_texts=["query"],    where={"category": "tutorial"})
# Comparison operatorsresults = collection.query(    query_texts=["query"],    where={"page": {"$gt": 10}}  # $gt, $gte, $lt, $lte, $ne)
# Logical operatorsresults = collection.query(    query_texts=["query"],    where={        "$and": [            {"category": "tutorial"},            {"difficulty": {"$lte": 3}}        ]    }  # Also: $or)
# Containsresults = collection.query(    query_texts=["query"],    where={"tags": {"$in": ["python", "ml"]}})

LangChain integration

python
from langchain_chroma import Chromafrom langchain_openai import OpenAIEmbeddingsfrom langchain.text_splitter import RecursiveCharacterTextSplitter
# Split documentstext_splitter = RecursiveCharacterTextSplitter(chunk_size=1000)docs = text_splitter.split_documents(documents)
# Create Chroma vector storevectorstore = Chroma.from_documents(    documents=docs,    embedding=OpenAIEmbeddings(),    persist_directory="./chroma_db")
# Queryresults = vectorstore.similarity_search("machine learning", k=3)
# As retrieverretriever = vectorstore.as_retriever(search_kwargs={"k": 5})

LlamaIndex integration

python
from llama_index.vector_stores.chroma import ChromaVectorStorefrom llama_index.core import VectorStoreIndex, StorageContextimport chromadb
# Initialize Chromadb = chromadb.PersistentClient(path="./chroma_db")collection = db.get_or_create_collection("my_collection")
# Create vector storevector_store = ChromaVectorStore(chroma_collection=collection)storage_context = StorageContext.from_defaults(vector_store=vector_store)
# Create indexindex = VectorStoreIndex.from_documents(    documents,    storage_context=storage_context)
# Queryquery_engine = index.as_query_engine()response = query_engine.query("What is machine learning?")

Server mode

python
# Run Chroma server# Terminal: chroma run --path ./chroma_db --port 8000
# Connect to serverimport chromadbfrom chromadb.config import Settings
client = chromadb.HttpClient(    host="localhost",    port=8000,    settings=Settings(anonymized_telemetry=False))
# Use as normalcollection = client.get_or_create_collection("my_docs")

Best practices

  1. Use persistent client - Don't lose data on restart
  2. Add metadata - Enables filtering and tracking
  3. Batch operations - Add multiple docs at once
  4. Choose right embedding model - Balance speed/quality
  5. Use filters - Narrow search space
  6. Unique IDs - Avoid collisions
  7. Regular backups - Copy chroma_db directory
  8. Monitor collection size - Scale up if needed
  9. Test embedding functions - Ensure quality
  10. Use server mode for production - Better for multi-user

Performance

OperationLatencyNotes
Add 100 docs~1-3sWith embedding
Query (top 10)~50-200msDepends on collection size
Metadata filter~10-50msFast with proper indexing

Resources

來源與署名

來源:orchestra-research/ai-research-skills位於15-rag/chroma提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架