Llama Cpp

orchestra-research/ai-research-skills/12-inference-serving/llama-cpp

作者 orchestra-research773a52944ba4MIT13K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3个月前更新

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

AI 生成的概览

指导在 CPU、Apple Silicon 和非 NVIDIA GPU 上使用 llama.cpp 进行本地 LLM 推理。

功能
该技能提供安装和使用 llama.cpp(一个 C/C++ LLM 推理引擎)的说明,包括 Metal、CUDA 和 ROCm 的构建选项。内容涵盖下载或转换 GGUF 模型、运行聊天和批量推理、启动兼容 OpenAI 的服务器,以及选择量化格式。它还包含硬件加速说明、性能基准测试和支持的模型系列。
适用场景
适用于在仅有 CPU 的机器、Apple Silicon Mac、AMD 或 Intel GPU,或没有 NVIDIA CUDA 硬件的边缘设备上运行 LLM 推理。当你需要通过 GGUF 量化减少内存占用,或希望无需 Docker 或 Python 进行简单部署时,也适合使用。
运行要求
需要安装 llama.cpp(通过 Homebrew 或从源码构建),并准备 GGUF 模型文件。可选的加速需要 Metal、CUDA 或 ROCm 构建支持。该技能将 llama-cpp-python 列为依赖,并引用 Hugging Face 下载模型。它不附带脚本,仅为说明文档。

llama.cpp

Pure C/C++ LLM inference with minimal dependencies, optimized for CPUs and non-NVIDIA hardware.

When to use llama.cpp

Use llama.cpp when:

  • Running on CPU-only machines
  • Deploying on Apple Silicon (M1/M2/M3/M4)
  • Using AMD or Intel GPUs (no CUDA)
  • Edge deployment (Raspberry Pi, embedded systems)
  • Need simple deployment without Docker/Python

Use TensorRT-LLM instead when:

  • Have NVIDIA GPUs (A100/H100)
  • Need maximum throughput (100K+ tok/s)
  • Running in datacenter with CUDA

Use vLLM instead when:

  • Have NVIDIA GPUs
  • Need Python-first API
  • Want PagedAttention

Quick start

Installation

bash
# macOS/Linuxbrew install llama.cpp
# Or build from sourcegit clone https://github.com/ggerganov/llama.cppcd llama.cppmake
# With Metal (Apple Silicon)make LLAMA_METAL=1
# With CUDA (NVIDIA)make LLAMA_CUDA=1
# With ROCm (AMD)make LLAMA_HIP=1

Download model

bash
# Download from HuggingFace (GGUF format)huggingface-cli download \    TheBloke/Llama-2-7B-Chat-GGUF \    llama-2-7b-chat.Q4_K_M.gguf \    --local-dir models/
# Or convert from HuggingFacepython convert_hf_to_gguf.py models/llama-2-7b-chat/

Run inference

bash
# Simple chat./llama-cli \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    -p "Explain quantum computing" \    -n 256  # Max tokens
# Interactive chat./llama-cli \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    --interactive

Server mode

bash
# Start OpenAI-compatible server./llama-server \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    --host 0.0.0.0 \    --port 8080 \    -ngl 32  # Offload 32 layers to GPU
# Client requestcurl http://localhost:8080/v1/chat/completions \  -H "Content-Type: application/json" \  -d '{    "model": "llama-2-7b-chat",    "messages": [{"role": "user", "content": "Hello!"}],    "temperature": 0.7,    "max_tokens": 100  }'

Quantization formats

GGUF format overview

FormatBitsSize (7B)SpeedQualityUse Case
Q4_K_M4.54.1 GBFastGoodRecommended default
Q4_K_S4.33.9 GBFasterLowerSpeed critical
Q5_K_M5.54.8 GBMediumBetterQuality critical
Q6_K6.55.5 GBSlowerBestMaximum quality
Q8_08.07.0 GBSlowExcellentMinimal degradation
Q2_K2.52.7 GBFastestPoorTesting only

Choosing quantization

bash
# General use (balanced)Q4_K_M  # 4-bit, medium quality
# Maximum speed (more degradation)Q2_K or Q3_K_M
# Maximum quality (slower)Q6_K or Q8_0
# Very large models (70B, 405B)Q3_K_M or Q4_K_S  # Lower bits to fit in memory

Hardware acceleration

Apple Silicon (Metal)

bash
# Build with Metalmake LLAMA_METAL=1
# Run with GPU acceleration (automatic)./llama-cli -m model.gguf -ngl 999  # Offload all layers
# Performance: M3 Max 40-60 tokens/sec (Llama 2-7B Q4_K_M)

NVIDIA GPUs (CUDA)

bash
# Build with CUDAmake LLAMA_CUDA=1
# Offload layers to GPU./llama-cli -m model.gguf -ngl 35  # Offload 35/40 layers
# Hybrid CPU+GPU for large models./llama-cli -m llama-70b.Q4_K_M.gguf -ngl 20  # GPU: 20 layers, CPU: rest

AMD GPUs (ROCm)

bash
# Build with ROCmmake LLAMA_HIP=1
# Run with AMD GPU./llama-cli -m model.gguf -ngl 999

Common patterns

Batch processing

bash
# Process multiple prompts from filecat prompts.txt | ./llama-cli \    -m model.gguf \    --batch-size 512 \    -n 100

Constrained generation

bash
# JSON output with grammar./llama-cli \    -m model.gguf \    -p "Generate a person: " \    --grammar-file grammars/json.gbnf
# Outputs valid JSON only

Context size

bash
# Increase context (default 512)./llama-cli \    -m model.gguf \    -c 4096  # 4K context window
# Very long context (if model supports)./llama-cli -m model.gguf -c 32768  # 32K context

Performance benchmarks

CPU performance (Llama 2-7B Q4_K_M)

CPUThreadsSpeedCost
Apple M3 Max1650 tok/s$0 (local)
AMD Ryzen 9 7950X3235 tok/s$0.50/hour
Intel i9-13900K3230 tok/s$0.40/hour
AWS c7i.16xlarge6440 tok/s$2.88/hour

GPU acceleration (Llama 2-7B Q4_K_M)

GPUSpeedvs CPUCost
NVIDIA RTX 4090120 tok/s3-4×$0 (local)
NVIDIA A1080 tok/s2-3×$1.00/hour
AMD MI25070 tok/s2×$2.00/hour
Apple M3 Max (Metal)50 tok/s~Same$0 (local)

Supported models

LLaMA family:

  • Llama 2 (7B, 13B, 70B)
  • Llama 3 (8B, 70B, 405B)
  • Code Llama

Mistral family:

  • Mistral 7B
  • Mixtral 8x7B, 8x22B

Other:

  • Falcon, BLOOM, GPT-J
  • Phi-3, Gemma, Qwen
  • LLaVA (vision), Whisper (audio)

Find models: https://huggingface.co/models?library=gguf

References

  • Quantization Guide [blocked] - GGUF formats, conversion, quality comparison
  • Server Deployment [blocked] - API endpoints, Docker, monitoring
  • Optimization [blocked] - Performance tuning, hybrid CPU+GPU

Resources

来源与署名

来源:orchestra-research/ai-research-skills位于12-inference-serving/llama-cpp提交773a529

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架