Llamaindex

orchestra-research/ai-research-skills/14-agents/llamaindex

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications.

僅含說明AI & Agents
AI 產生的概覽

指導使用 LlamaIndex 框架建構 RAG 與文件問答應用,涵蓋資料接入、索引、查詢與代理。

功能
此技能是一份使用 LlamaIndex 的參考指南,LlamaIndex 是一個將大型語言模型與私有資料連接的 Python 資料框架。內容涵蓋安裝該套件、透過連接器載入文件、建立向量索引、清單索引與樹狀索引,並使用查詢引擎、檢索器與聊天引擎進行查詢。此外也介紹搭配工具的代理、中繼資料篩選、結構化輸出、向量儲存整合、多模態檢索、評估與最佳實務,並附有程式碼範例。
適用情境
適用於建構檢索增強生成(RAG)流程、以私有資料為基礎的文件問答、知識庫,或以企業資料支撐的聊天機器人。也適合在 LlamaIndex 與 LangChain 等替代方案之間做選擇,用於以資料為核心的 LLM 應用。
執行需求
需要 Python 環境以及 llama-index 套件與其整合元件,還需要 OpenAI 或 Anthropic 等 LLM 與嵌入服務供應商,並設定對應的 API 憑證。部分範例涉及 Pinecone、Chroma、FAISS、HuggingFace 與資料庫等外部服務,網頁或 API 讀取器需要網路存取。此技能不附帶指令碼,僅包含說明文件與參考文件。

LlamaIndex - Data Framework for LLM Applications

The leading framework for connecting LLMs with your data.

When to use LlamaIndex

Use LlamaIndex when:

  • Building RAG (retrieval-augmented generation) applications
  • Need document question-answering over private data
  • Ingesting data from multiple sources (300+ connectors)
  • Creating knowledge bases for LLMs
  • Building chatbots with enterprise data
  • Need structured data extraction from documents

Metrics:

  • 45,100+ GitHub stars
  • 23,000+ repositories use LlamaIndex
  • 300+ data connectors (LlamaHub)
  • 1,715+ contributors
  • v0.14.7 (stable)

Use alternatives instead:

  • LangChain: More general-purpose, better for agents
  • Haystack: Production search pipelines
  • txtai: Lightweight semantic search
  • Chroma: Just need vector storage

Quick start

Installation

bash
# Starter package (recommended)pip install llama-index
# Or minimal core + specific integrationspip install llama-index-corepip install llama-index-llms-openaipip install llama-index-embeddings-openai

5-line RAG example

python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Load documentsdocuments = SimpleDirectoryReader("data").load_data()
# Create indexindex = VectorStoreIndex.from_documents(documents)
# Queryquery_engine = index.as_query_engine()response = query_engine.query("What did the author do growing up?")print(response)

Core concepts

1. Data connectors - Load documents

python
from llama_index.core import SimpleDirectoryReader, Documentfrom llama_index.readers.web import SimpleWebPageReaderfrom llama_index.readers.github import GithubRepositoryReader
# Directory of filesdocuments = SimpleDirectoryReader("./data").load_data()
# Web pagesreader = SimpleWebPageReader()documents = reader.load_data(["https://example.com"])
# GitHub repositoryreader = GithubRepositoryReader(owner="user", repo="repo")documents = reader.load_data(branch="main")
# Manual document creationdoc = Document(    text="This is the document content",    metadata={"source": "manual", "date": "2025-01-01"})

2. Indices - Structure data

python
from llama_index.core import VectorStoreIndex, ListIndex, TreeIndex
# Vector index (most common - semantic search)vector_index = VectorStoreIndex.from_documents(documents)
# List index (sequential scan)list_index = ListIndex.from_documents(documents)
# Tree index (hierarchical summary)tree_index = TreeIndex.from_documents(documents)
# Save indexindex.storage_context.persist(persist_dir="./storage")
# Load indexfrom llama_index.core import load_index_from_storage, StorageContextstorage_context = StorageContext.from_defaults(persist_dir="./storage")index = load_index_from_storage(storage_context)

3. Query engines - Ask questions

python
# Basic queryquery_engine = index.as_query_engine()response = query_engine.query("What is the main topic?")print(response)
# Streaming responsequery_engine = index.as_query_engine(streaming=True)response = query_engine.query("Explain quantum computing")for text in response.response_gen:    print(text, end="", flush=True)
# Custom configurationquery_engine = index.as_query_engine(    similarity_top_k=3,          # Return top 3 chunks    response_mode="compact",     # Or "tree_summarize", "simple_summarize"    verbose=True)

4. Retrievers - Find relevant chunks

python
# Vector retrieverretriever = index.as_retriever(similarity_top_k=5)nodes = retriever.retrieve("machine learning")
# With filteringretriever = index.as_retriever(    similarity_top_k=3,    filters={"metadata.category": "tutorial"})
# Custom retrieverfrom llama_index.core.retrievers import BaseRetriever
class CustomRetriever(BaseRetriever):    def _retrieve(self, query_bundle):        # Your custom retrieval logic        return nodes

Agents with tools

Basic agent

python
from llama_index.core.agent import FunctionAgentfrom llama_index.llms.openai import OpenAI
# Define toolsdef multiply(a: int, b: int) -> int:    """Multiply two numbers."""    return a * b
def add(a: int, b: int) -> int:    """Add two numbers."""    return a + b
# Create agentllm = OpenAI(model="gpt-4o")agent = FunctionAgent.from_tools(    tools=[multiply, add],    llm=llm,    verbose=True)
# Use agentresponse = agent.chat("What is 25 * 17 + 142?")print(response)

RAG agent (document search + tools)

python
from llama_index.core.tools import QueryEngineTool
# Create index as beforeindex = VectorStoreIndex.from_documents(documents)
# Wrap query engine as toolquery_tool = QueryEngineTool.from_defaults(    query_engine=index.as_query_engine(),    name="python_docs",    description="Useful for answering questions about Python programming")
# Agent with document search + calculatoragent = FunctionAgent.from_tools(    tools=[query_tool, multiply, add],    llm=llm)
# Agent decides when to search docs vs calculateresponse = agent.chat("According to the docs, what is Python used for?")

Advanced RAG patterns

Chat engine (conversational)

python
from llama_index.core.chat_engine import CondensePlusContextChatEngine
# Chat with memorychat_engine = index.as_chat_engine(    chat_mode="condense_plus_context",  # Or "context", "react"    verbose=True)
# Multi-turn conversationresponse1 = chat_engine.chat("What is Python?")response2 = chat_engine.chat("Can you give examples?")  # Remembers contextresponse3 = chat_engine.chat("What about web frameworks?")

Metadata filtering

python
from llama_index.core.vector_stores import MetadataFilters, ExactMatchFilter
# Filter by metadatafilters = MetadataFilters(    filters=[        ExactMatchFilter(key="category", value="tutorial"),        ExactMatchFilter(key="difficulty", value="beginner")    ])
retriever = index.as_retriever(    similarity_top_k=3,    filters=filters)
query_engine = index.as_query_engine(filters=filters)

Structured output

python
from pydantic import BaseModelfrom llama_index.core.output_parsers import PydanticOutputParser
class Summary(BaseModel):    title: str    main_points: list[str]    conclusion: str
# Get structured responseoutput_parser = PydanticOutputParser(output_cls=Summary)query_engine = index.as_query_engine(output_parser=output_parser)
response = query_engine.query("Summarize the document")summary = response  # Pydantic modelprint(summary.title, summary.main_points)

Data ingestion patterns

Multiple file types

python
# Load all supported formatsdocuments = SimpleDirectoryReader(    "./data",    recursive=True,    required_exts=[".pdf", ".docx", ".txt", ".md"]).load_data()

Web scraping

python
from llama_index.readers.web import BeautifulSoupWebReader
reader = BeautifulSoupWebReader()documents = reader.load_data(urls=[    "https://docs.python.org/3/tutorial/",    "https://docs.python.org/3/library/"])

Database

python
from llama_index.readers.database import DatabaseReader
reader = DatabaseReader(    sql_database_uri="postgresql://user:pass@localhost/db")documents = reader.load_data(query="SELECT * FROM articles")

API endpoints

python
from llama_index.readers.json import JSONReader
reader = JSONReader()documents = reader.load_data("https://api.example.com/data.json")

Vector store integrations

Chroma (local)

python
from llama_index.vector_stores.chroma import ChromaVectorStoreimport chromadb
# Initialize Chromadb = chromadb.PersistentClient(path="./chroma_db")collection = db.get_or_create_collection("my_collection")
# Create vector storevector_store = ChromaVectorStore(chroma_collection=collection)
# Use in indexfrom llama_index.core import StorageContextstorage_context = StorageContext.from_defaults(vector_store=vector_store)index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)

Pinecone (cloud)

python
from llama_index.vector_stores.pinecone import PineconeVectorStoreimport pinecone
# Initialize Pineconepinecone.init(api_key="your-key", environment="us-west1-gcp")pinecone_index = pinecone.Index("my-index")
# Create vector storevector_store = PineconeVectorStore(pinecone_index=pinecone_index)storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)

FAISS (fast)

python
from llama_index.vector_stores.faiss import FaissVectorStoreimport faiss
# Create FAISS indexd = 1536  # Dimension of embeddingsfaiss_index = faiss.IndexFlatL2(d)
vector_store = FaissVectorStore(faiss_index=faiss_index)storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)

Customization

Custom LLM

python
from llama_index.llms.anthropic import Anthropicfrom llama_index.core import Settings
# Set global LLMSettings.llm = Anthropic(model="claude-sonnet-4-5-20250929")
# Now all queries use Anthropicquery_engine = index.as_query_engine()

Custom embeddings

python
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
# Use HuggingFace embeddingsSettings.embed_model = HuggingFaceEmbedding(    model_name="sentence-transformers/all-mpnet-base-v2")
index = VectorStoreIndex.from_documents(documents)

Custom prompt templates

python
from llama_index.core import PromptTemplate
qa_prompt = PromptTemplate(    "Context: {context_str}\n"    "Question: {query_str}\n"    "Answer the question based only on the context. "    "If the answer is not in the context, say 'I don't know'.\n"    "Answer: ")
query_engine = index.as_query_engine(text_qa_template=qa_prompt)

Multi-modal RAG

Image + text

python
from llama_index.core import SimpleDirectoryReaderfrom llama_index.multi_modal_llms.openai import OpenAIMultiModal
# Load images and documentsdocuments = SimpleDirectoryReader(    "./data",    required_exts=[".jpg", ".png", ".pdf"]).load_data()
# Multi-modal indexindex = VectorStoreIndex.from_documents(documents)
# Query with multi-modal LLMmulti_modal_llm = OpenAIMultiModal(model="gpt-4o")query_engine = index.as_query_engine(llm=multi_modal_llm)
response = query_engine.query("What is in the diagram on page 3?")

Evaluation

Response quality

python
from llama_index.core.evaluation import RelevancyEvaluator, FaithfulnessEvaluator
# Evaluate relevancerelevancy = RelevancyEvaluator()result = relevancy.evaluate_response(    query="What is Python?",    response=response)print(f"Relevancy: {result.passing}")
# Evaluate faithfulness (no hallucination)faithfulness = FaithfulnessEvaluator()result = faithfulness.evaluate_response(    query="What is Python?",    response=response)print(f"Faithfulness: {result.passing}")

Best practices

  1. Use vector indices for most cases - Best performance
  2. Save indices to disk - Avoid re-indexing
  3. Chunk documents properly - 512-1024 tokens optimal
  4. Add metadata - Enables filtering and tracking
  5. Use streaming - Better UX for long responses
  6. Enable verbose during dev - See retrieval process
  7. Evaluate responses - Check relevance and faithfulness
  8. Use chat engine for conversations - Built-in memory
  9. Persist storage - Don't lose your index
  10. Monitor costs - Track embedding and LLM usage

Common patterns

Document Q&A system

python
# Complete RAG pipelinedocuments = SimpleDirectoryReader("docs").load_data()index = VectorStoreIndex.from_documents(documents)index.storage_context.persist(persist_dir="./storage")
# Queryquery_engine = index.as_query_engine(    similarity_top_k=3,    response_mode="compact",    verbose=True)response = query_engine.query("What is the main topic?")print(response)print(f"Sources: {[node.metadata['file_name'] for node in response.source_nodes]}")

Chatbot with memory

python
# Conversational interfacechat_engine = index.as_chat_engine(    chat_mode="condense_plus_context",    verbose=True)
# Multi-turn chatwhile True:    user_input = input("You: ")    if user_input.lower() == "quit":        break    response = chat_engine.chat(user_input)    print(f"Bot: {response}")

Performance benchmarks

OperationLatencyNotes
Index 100 docs~10-30sOne-time, can persist
Query (vector)~0.5-2sRetrieval + LLM
Streaming query~0.5s first tokenBetter UX
Agent with tools~3-8sMultiple tool calls

LlamaIndex vs LangChain

FeatureLlamaIndexLangChain
Best forRAG, document Q&AAgents, general LLM apps
Data connectors300+ (LlamaHub)100+
RAG focusCore featureOne of many
Learning curveEasier for RAGSteeper
CustomizationHighVery high
DocumentationExcellentGood

Use LlamaIndex when:

  • Your primary use case is RAG
  • Need many data connectors
  • Want simpler API for document Q&A
  • Building knowledge retrieval system

Use LangChain when:

  • Building complex agents
  • Need more general-purpose tools
  • Want more flexibility
  • Complex multi-step workflows

References

  • Query Engines Guide [blocked] - Query modes, customization, streaming
  • Agents Guide [blocked] - Tool creation, RAG agents, multi-step reasoning
  • Data Connectors Guide [blocked] - 300+ connectors, custom loaders

Resources

來源與署名

來源:orchestra-research/ai-research-skills位於14-agents/llamaindex提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架