Faiss

orchestra-research/ai-research-skills/15-rag/faiss

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications.

Instructions onlyAI & Agents
AI-generated overview

Reference guide for using FAISS to build and tune vector similarity search indexes.

What it does
This skill is a reference guide for Facebook AI's FAISS library for similarity search and clustering of dense vectors. It documents installation, index types (Flat, IVF, HNSW, Product Quantization), saving and loading indexes, GPU acceleration, and LangChain/LlamaIndex integration, with code examples and performance comparisons. It produces guidance and example code rather than executable artifacts.
When to use it
Use it when you need fast k-nearest-neighbor search or large-scale vector retrieval, including billion-scale datasets or GPU-accelerated search. It is also relevant when choosing between FAISS index types or comparing FAISS with alternatives such as Chroma, Pinecone, Weaviate, or Annoy.
Requirements
Requires the faiss-cpu or faiss-gpu Python package and numpy; GPU acceleration requires faiss-gpu and suitable hardware. Optional integrations reference LangChain, LlamaIndex, and OpenAI embeddings. The skill ships no scripts, only instructions and a reference document.

FAISS - Efficient Similarity Search

Facebook AI's library for billion-scale vector similarity search.

When to use FAISS

Use FAISS when:

  • Need fast similarity search on large vector datasets (millions/billions)
  • GPU acceleration required
  • Pure vector similarity (no metadata filtering needed)
  • High throughput, low latency critical
  • Offline/batch processing of embeddings

Metrics:

  • 31,700+ GitHub stars
  • Meta/Facebook AI Research
  • Handles billions of vectors
  • C++ with Python bindings

Use alternatives instead:

  • Chroma/Pinecone: Need metadata filtering
  • Weaviate: Need full database features
  • Annoy: Simpler, fewer features

Quick start

Installation

bash
# CPU onlypip install faiss-cpu
# GPU supportpip install faiss-gpu

Basic usage

python
import faissimport numpy as np
# Create sample data (1000 vectors, 128 dimensions)d = 128nb = 1000vectors = np.random.random((nb, d)).astype('float32')
# Create indexindex = faiss.IndexFlatL2(d)  # L2 distanceindex.add(vectors)             # Add vectors
# Searchk = 5  # Find 5 nearest neighborsquery = np.random.random((1, d)).astype('float32')distances, indices = index.search(query, k)
print(f"Nearest neighbors: {indices}")print(f"Distances: {distances}")

Index types

1. Flat (exact search)

python
# L2 (Euclidean) distanceindex = faiss.IndexFlatL2(d)
# Inner product (cosine similarity if normalized)index = faiss.IndexFlatIP(d)
# Slowest, most accurate

2. IVF (inverted file) - Fast approximate

python
# Create quantizerquantizer = faiss.IndexFlatL2(d)
# IVF index with 100 clustersnlist = 100index = faiss.IndexIVFFlat(quantizer, d, nlist)
# Train on dataindex.train(vectors)
# Add vectorsindex.add(vectors)
# Search (nprobe = clusters to search)index.nprobe = 10distances, indices = index.search(query, k)

3. HNSW (Hierarchical NSW) - Best quality/speed

python
# HNSW indexM = 32  # Number of connections per layerindex = faiss.IndexHNSWFlat(d, M)
# No training neededindex.add(vectors)
# Searchdistances, indices = index.search(query, k)

4. Product Quantization - Memory efficient

python
# PQ reduces memory by 16-32×m = 8   # Number of subquantizersnbits = 8index = faiss.IndexPQ(d, m, nbits)
# Train and addindex.train(vectors)index.add(vectors)

Save and load

python
# Save indexfaiss.write_index(index, "large.index")
# Load indexindex = faiss.read_index("large.index")
# Continue usingdistances, indices = index.search(query, k)

GPU acceleration

python
# Single GPUres = faiss.StandardGpuResources()index_cpu = faiss.IndexFlatL2(d)index_gpu = faiss.index_cpu_to_gpu(res, 0, index_cpu)  # GPU 0
# Multi-GPUindex_gpu = faiss.index_cpu_to_all_gpus(index_cpu)
# 10-100× faster than CPU

LangChain integration

python
from langchain_community.vectorstores import FAISSfrom langchain_openai import OpenAIEmbeddings
# Create FAISS vector storevectorstore = FAISS.from_documents(docs, OpenAIEmbeddings())
# Savevectorstore.save_local("faiss_index")
# Loadvectorstore = FAISS.load_local(    "faiss_index",    OpenAIEmbeddings(),    allow_dangerous_deserialization=True)
# Searchresults = vectorstore.similarity_search("query", k=5)

LlamaIndex integration

python
from llama_index.vector_stores.faiss import FaissVectorStoreimport faiss
# Create FAISS indexd = 1536faiss_index = faiss.IndexFlatL2(d)
vector_store = FaissVectorStore(faiss_index=faiss_index)

Best practices

  1. Choose right index type - Flat for <10K, IVF for 10K-1M, HNSW for quality
  2. Normalize for cosine - Use IndexFlatIP with normalized vectors
  3. Use GPU for large datasets - 10-100× faster
  4. Save trained indices - Training is expensive
  5. Tune nprobe/ef_search - Balance speed/accuracy
  6. Monitor memory - PQ for large datasets
  7. Batch queries - Better GPU utilization

Performance

Index TypeBuild TimeSearch TimeMemoryAccuracy
FlatFastSlowHigh100%
IVFMediumFastMedium95-99%
HNSWSlowFastestHigh99%
PQMediumFastLow90-95%

Resources

Source and attribution

Source:orchestra-research/ai-research-skillsin15-rag/faissat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal