<overview>
Retrieval Augmented Generation (RAG) enhances LLM responses by fetching relevant context from external knowledge sources.
Pipeline:
- Index: Load → Split → Embed → Store
- Retrieve: Query → Embed → Search → Return docs
- Generate: Docs + Query → LLM → Response
Key Components:
- Document Loaders: Ingest data from files, web, databases
- Text Splitters: Break documents into chunks
- Embeddings: Convert text to vectors
- Vector Stores: Store and search embeddings
</overview>
<vectorstore-selection>
</vectorstore-selection>
Complete RAG Pipeline
<ex-basic-rag-setup>
<python>
End-to-end RAG pipeline: load documents, split into chunks, embed, store, retrieve, and generate a response.
</python>
<typescript>
End-to-end RAG pipeline: load documents, split into chunks, embed, store, retrieve, and generate a response.
</typescript>
</ex-basic-rag-setup>
Document Loaders
<ex-loading-pdf>
<python>
Load a PDF file and extract each page as a separate document.
</python>
<typescript>
Load a PDF file and extract each page as a separate document.
</typescript>
</ex-loading-pdf>
<ex-loading-web-pages>
<python>
Fetch and parse content from a web URL into a document.
</python>
<typescript>
Fetch and parse content from a web URL into a document using Cheerio.
</typescript>
</ex-loading-web-pages>
<ex-loading-directory>
<python>
Load all text files from a directory using a glob pattern.
</python>
</ex-loading-directory>
Text Splitting
<ex-text-splitting>
<python>
Split documents into chunks using RecursiveCharacterTextSplitter with configurable size and overlap.
</python>
</ex-text-splitting>
Vector Stores
<ex-chroma-vectorstore>
<python>
Create a persistent Chroma vector store and reload it from disk.
</python>
<typescript>
Create a Chroma vector store connected to a running Chroma server.
</typescript>
</ex-chroma-vectorstore>
<ex-faiss-vectorstore>
<python>
Create a FAISS vector store, save it to disk, and reload it.
</python>
<typescript>
Create a FAISS vector store, save it to disk, and reload it.
</typescript>
</ex-faiss-vectorstore>
Retrieval
<ex-similarity-search>
<python>
Perform similarity search and retrieve results with relevance scores.
</python>
<typescript>
Perform similarity search and retrieve results with relevance scores.
</typescript>
</ex-similarity-search>
<ex-mmr-search>
<python>
Use MMR (Maximal Marginal Relevance) to balance relevance and diversity in search results.
</python>
</ex-mmr-search>
<ex-metadata-filtering>
<python>
Add metadata to documents and filter search results by metadata properties.
</python>
</ex-metadata-filtering>
<ex-rag-with-agent>
<python>
Create an agent that uses RAG as a tool for answering questions.
</python>
<typescript>
Create an agent that uses RAG as a tool for answering questions.
</typescript>
</ex-rag-with-agent>
<boundaries>
What You CAN Configure
- Chunk size/overlap
- Embedding model
- Number of results (k)
- Metadata filters
- Search algorithms: Similarity, MMR
What You CANNOT Configure
- Embedding dimensions (per model)
- Mix embeddings from different models in same store
</boundaries>
<fix-chunk-size>
<python>
Chunk size 500-1500 is typically good.
</python>
<typescript>
Chunk size 500-1500 is typically good.
</typescript>
</fix-chunk-size>
<fix-chunk-overlap>
<python>
Use overlap (10-20% of chunk size) to maintain context at boundaries.
</python>
</fix-chunk-overlap>
<fix-persist-vectorstore>
<python>
Use persistent vector store instead of in-memory to avoid data loss.
</python>
<typescript>
Use persistent vector store instead of in-memory to avoid data loss.
</typescript>
</fix-persist-vectorstore>
<fix-consistent-embeddings>
<python>
Use the same embedding model for indexing and querying.
</python>
<typescript>
Use the same embedding model for indexing and querying.
</typescript>
</fix-consistent-embeddings>
<fix-faiss-deserialization>
<python>
Only opt in to FAISS deserialization for trusted local indexes. Python FAISS indexes include pickle-backed metadata, and untrusted pickle files can execute arbitrary code during loading.
If you cannot guarantee the provenance of a persisted index, do not load it with allow_dangerous_deserialization=True. Rebuild the index from trusted source documents or use a vector store/backend that does not require pickle deserialization for untrusted files.
</python>
</fix-faiss-deserialization>
<fix-dimension-mismatch>
<python>
Ensure embedding dimensions match the vector store index dimensions.
</python>
</fix-dimension-mismatch>


