Llamaindex Development

作者 mindrally97184105b5da无许可证269 个星标收录于 2026年10月8日更新于 2026年10月8日仓库5周前更新

Expert guidance for LlamaIndex development including RAG applications, vector stores, document processing, query engines, and building production AI applications.

仅含说明AI & Agents
AI 生成的概览

指导 LlamaIndex 开发,涵盖 RAG 应用、索引、向量存储、查询引擎、检索器、嵌入与代理。

功能
提供构建 LlamaIndex RAG 应用的专家指导和 Python 代码示例。内容涵盖文档加载、文本切分、向量存储索引与持久化、查询引擎、检索器、嵌入、LLM 配置、代理、缓存、异步操作、错误处理与测试。产出为带代码片段的说明性参考资料,而非生成的文件。
适用场景
适用于构建或审查基于 LlamaIndex 的 Python 检索增强生成应用。适合涉及文档索引、向量存储、查询引擎、检索器、嵌入或 LlamaIndex 代理的任务。
运行要求
不附带脚本,仅为说明性内容。示例引用 Python 包,包括 llama-index、llama-index-embeddings-openai、llama-index-llms-openai、llama-index-vector-stores-chroma、chromadb、python-dotenv 和 pydantic;部分示例需要 OpenAI 或 Anthropic 的 API 凭据以及向量存储服务。

LlamaIndex Development

You are an expert in LlamaIndex for building RAG (Retrieval-Augmented Generation) applications, data indexing, and LLM-powered applications with Python.

Key Principles

  • Write concise, technical responses with accurate Python examples
  • Use functional, declarative programming; avoid classes where possible
  • Prioritize code quality, maintainability, and performance
  • Use descriptive variable names that reflect their purpose
  • Follow PEP 8 style guidelines

Code Organization

Directory Structure

project/├── data/                 # Source documents and data├── indexes/              # Persisted index storage├── loaders/              # Custom document loaders├── retrievers/           # Custom retriever implementations├── query_engines/        # Query engine configurations├── prompts/              # Custom prompt templates├── transformations/      # Document transformations├── callbacks/            # Custom callback handlers├── utils/                # Utility functions├── tests/                # Test files└── config/               # Configuration files

Naming Conventions

  • Use snake_case for files, functions, and variables
  • Use PascalCase for classes
  • Prefix private functions with underscore
  • Use descriptive names (e.g., create_vector_index, build_query_engine)

Document Loading

Using Document Loaders

python
from llama_index.core import SimpleDirectoryReaderfrom llama_index.readers.file import PDFReader, DocxReader
# Load from directorydocuments = SimpleDirectoryReader(    input_dir="./data",    recursive=True,    required_exts=[".pdf", ".txt", ".md"]).load_data()
# Load specific file typespdf_reader = PDFReader()documents = pdf_reader.load_data(file="document.pdf")

Custom Loaders

python
from llama_index.core.readers.base import BaseReaderfrom llama_index.core import Document
class CustomLoader(BaseReader):    def load_data(self, file_path: str) -> list[Document]:        # Custom loading logic        with open(file_path, 'r') as f:            content = f.read()
        return [Document(            text=content,            metadata={"source": file_path}        )]

Text Splitting and Processing

Node Parsing

python
from llama_index.core.node_parser import (    SentenceSplitter,    SemanticSplitterNodeParser,    MarkdownNodeParser)
# Simple sentence splittingsplitter = SentenceSplitter(    chunk_size=1024,    chunk_overlap=200)nodes = splitter.get_nodes_from_documents(documents)
# Semantic splitting (preserves meaning)from llama_index.embeddings.openai import OpenAIEmbedding
semantic_splitter = SemanticSplitterNodeParser(    embed_model=OpenAIEmbedding(),    breakpoint_percentile_threshold=95)
# Markdown-aware splittingmarkdown_splitter = MarkdownNodeParser()

Best Practices for Chunking

  • Choose chunk size based on your embedding model's context window
  • Use overlap to maintain context between chunks
  • Preserve document structure when possible
  • Include metadata for filtering and retrieval
  • Use semantic splitting for better coherence

Vector Stores and Indexing

Creating Indexes

python
from llama_index.core import VectorStoreIndex, StorageContextfrom llama_index.vector_stores.chroma import ChromaVectorStoreimport chromadb
# In-memory indexindex = VectorStoreIndex.from_documents(documents)
# With persistent vector storechroma_client = chromadb.PersistentClient(path="./chroma_db")chroma_collection = chroma_client.get_or_create_collection("my_collection")
vector_store = ChromaVectorStore(chroma_collection=chroma_collection)storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents(    documents,    storage_context=storage_context)

Supported Vector Stores

  • Chroma (local development)
  • Pinecone (production, managed)
  • Weaviate (production, self-hosted or managed)
  • Qdrant (production, self-hosted or managed)
  • PostgreSQL with pgvector
  • MongoDB Atlas Vector Search

Index Persistence

python
from llama_index.core import StorageContext, load_index_from_storage
# Persist indexindex.storage_context.persist(persist_dir="./storage")
# Load indexstorage_context = StorageContext.from_defaults(persist_dir="./storage")index = load_index_from_storage(storage_context)

Query Engines

Basic Query Engine

python
from llama_index.core import VectorStoreIndex
index = VectorStoreIndex.from_documents(documents)query_engine = index.as_query_engine(    similarity_top_k=5,    response_mode="compact")
response = query_engine.query("What is the main topic?")print(response.response)

Response Modes

  • refine: Iteratively refine answer through each node
  • compact: Combine chunks before sending to LLM
  • tree_summarize: Build tree and summarize
  • simple_summarize: Truncate and summarize
  • accumulate: Accumulate responses from each node

Advanced Query Engine

python
from llama_index.core.query_engine import RetrieverQueryEnginefrom llama_index.core.postprocessor import SimilarityPostprocessor
query_engine = RetrieverQueryEngine.from_args(    retriever=index.as_retriever(similarity_top_k=10),    node_postprocessors=[        SimilarityPostprocessor(similarity_cutoff=0.7)    ],    response_mode="compact")

Retrievers

Custom Retrievers

python
from llama_index.core.retrievers import VectorIndexRetriever
# Basic retrieverretriever = VectorIndexRetriever(    index=index,    similarity_top_k=10)
# Retrieve nodesnodes = retriever.retrieve("search query")

Hybrid Search

python
from llama_index.core.retrievers import QueryFusionRetriever
# Combine multiple retrieval strategiesretriever = QueryFusionRetriever(    [        index.as_retriever(similarity_top_k=5),        bm25_retriever,  # Keyword-based    ],    num_queries=4,    use_async=True)

Embeddings

Embedding Models

python
from llama_index.embeddings.openai import OpenAIEmbeddingfrom llama_index.embeddings.huggingface import HuggingFaceEmbeddingfrom llama_index.core import Settings
# OpenAI embeddingsSettings.embed_model = OpenAIEmbedding(    model="text-embedding-3-small",    dimensions=512  # Optional dimension reduction)
# Local embeddingsSettings.embed_model = HuggingFaceEmbedding(    model_name="BAAI/bge-small-en-v1.5")

LLM Configuration

Setting Up LLMs

python
from llama_index.llms.openai import OpenAIfrom llama_index.llms.anthropic import Anthropicfrom llama_index.core import Settings
# OpenAISettings.llm = OpenAI(    model="gpt-4o",    temperature=0.1)
# AnthropicSettings.llm = Anthropic(    model="claude-sonnet-4-20250514",    temperature=0.1)

Agents

Building Agents

python
from llama_index.core.agent import ReActAgentfrom llama_index.core.tools import QueryEngineTool, ToolMetadata
# Create tools from query enginestools = [    QueryEngineTool(        query_engine=documents_query_engine,        metadata=ToolMetadata(            name="documents",            description="Search through documents"        )    ),    QueryEngineTool(        query_engine=code_query_engine,        metadata=ToolMetadata(            name="codebase",            description="Search through code"        )    )]
# Create agentagent = ReActAgent.from_tools(    tools,    llm=llm,    verbose=True)
response = agent.chat("Find information about X")

Performance Optimization

Caching

python
from llama_index.core import Settingsfrom llama_index.core.llms import LLMCache
# Enable LLM response cachingSettings.llm = OpenAI(model="gpt-4o")Settings.llm_cache = LLMCache()

Async Operations

python
# Use async for better performanceresponse = await query_engine.aquery("question")
# Batch processingresponses = await asyncio.gather(*[    query_engine.aquery(q) for q in questions])

Embedding Optimization

  • Batch embeddings when possible
  • Use smaller embedding dimensions when accuracy allows
  • Cache embeddings for repeated documents
  • Use local models for cost-sensitive applications

Error Handling

python
from llama_index.core.callbacks import CallbackManager, LlamaDebugHandler
# Debug handler for troubleshootingdebug_handler = LlamaDebugHandler()callback_manager = CallbackManager([debug_handler])
Settings.callback_manager = callback_manager

Testing

  • Unit test document loaders and transformations
  • Test retrieval quality with known queries
  • Validate index persistence and loading
  • Test query engine responses
  • Monitor retrieval metrics (precision, recall)

Dependencies

  • llama-index
  • llama-index-embeddings-openai
  • llama-index-llms-openai
  • llama-index-vector-stores-chroma
  • chromadb
  • python-dotenv
  • pydantic

来源与署名

来源:mindrally/skills位于llamaindex-development提交9718410

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架