LlamaIndex Development
You are an expert in LlamaIndex for building RAG (Retrieval-Augmented Generation) applications, data indexing, and LLM-powered applications with Python.
Key Principles
- Write concise, technical responses with accurate Python examples
- Use functional, declarative programming; avoid classes where possible
- Prioritize code quality, maintainability, and performance
- Use descriptive variable names that reflect their purpose
- Follow PEP 8 style guidelines
Code Organization
Directory Structure
Naming Conventions
- Use snake_case for files, functions, and variables
- Use PascalCase for classes
- Prefix private functions with underscore
- Use descriptive names (e.g.,
create_vector_index,build_query_engine)
Document Loading
Using Document Loaders
Custom Loaders
Text Splitting and Processing
Node Parsing
Best Practices for Chunking
- Choose chunk size based on your embedding model's context window
- Use overlap to maintain context between chunks
- Preserve document structure when possible
- Include metadata for filtering and retrieval
- Use semantic splitting for better coherence
Vector Stores and Indexing
Creating Indexes
Supported Vector Stores
- Chroma (local development)
- Pinecone (production, managed)
- Weaviate (production, self-hosted or managed)
- Qdrant (production, self-hosted or managed)
- PostgreSQL with pgvector
- MongoDB Atlas Vector Search
Index Persistence
Query Engines
Basic Query Engine
Response Modes
refine: Iteratively refine answer through each nodecompact: Combine chunks before sending to LLMtree_summarize: Build tree and summarizesimple_summarize: Truncate and summarizeaccumulate: Accumulate responses from each node
Advanced Query Engine
Retrievers
Custom Retrievers
Hybrid Search
Embeddings
Embedding Models
LLM Configuration
Setting Up LLMs
Agents
Building Agents
Performance Optimization
Caching
Async Operations
Embedding Optimization
- Batch embeddings when possible
- Use smaller embedding dimensions when accuracy allows
- Cache embeddings for repeated documents
- Use local models for cost-sensitive applications
Error Handling
Testing
- Unit test document loaders and transformations
- Test retrieval quality with known queries
- Validate index persistence and loading
- Test query engine responses
- Monitor retrieval metrics (precision, recall)
Dependencies
- llama-index
- llama-index-embeddings-openai
- llama-index-llms-openai
- llama-index-vector-stores-chroma
- chromadb
- python-dotenv
- pydantic



