Llama Cpp

orchestra-research/ai-research-skills/12-inference-serving/llama-cpp

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

Instructions onlyAI & AgentsDevOps & Cloud
AI-generated overview

Guides running local LLM inference with llama.cpp on CPUs, Apple Silicon, and non-NVIDIA GPUs.

What it does
This skill provides instructions for installing and using llama.cpp, a C/C++ LLM inference engine, including build options for Metal, CUDA, and ROCm. It covers downloading or converting GGUF models, running chat and batch inference, launching an OpenAI-compatible server, and selecting quantization formats. It also includes hardware acceleration notes, performance benchmarks, and supported model families.
When to use it
Use it when running LLM inference on CPU-only machines, Apple Silicon Macs, AMD or Intel GPUs, or edge devices without NVIDIA CUDA hardware. It is also relevant when you need GGUF quantization to reduce memory use or want a simple deployment without Docker or Python.
Requirements
Requires llama.cpp installed via Homebrew or built from source, and a GGUF model file. Optional acceleration requires Metal, CUDA, or ROCm build support. The skill lists llama-cpp-python as a dependency and references Hugging Face for model downloads. It ships no scripts; it is instructions only.

llama.cpp

Pure C/C++ LLM inference with minimal dependencies, optimized for CPUs and non-NVIDIA hardware.

When to use llama.cpp

Use llama.cpp when:

  • Running on CPU-only machines
  • Deploying on Apple Silicon (M1/M2/M3/M4)
  • Using AMD or Intel GPUs (no CUDA)
  • Edge deployment (Raspberry Pi, embedded systems)
  • Need simple deployment without Docker/Python

Use TensorRT-LLM instead when:

  • Have NVIDIA GPUs (A100/H100)
  • Need maximum throughput (100K+ tok/s)
  • Running in datacenter with CUDA

Use vLLM instead when:

  • Have NVIDIA GPUs
  • Need Python-first API
  • Want PagedAttention

Quick start

Installation

bash
# macOS/Linuxbrew install llama.cpp
# Or build from sourcegit clone https://github.com/ggerganov/llama.cppcd llama.cppmake
# With Metal (Apple Silicon)make LLAMA_METAL=1
# With CUDA (NVIDIA)make LLAMA_CUDA=1
# With ROCm (AMD)make LLAMA_HIP=1

Download model

bash
# Download from HuggingFace (GGUF format)huggingface-cli download \    TheBloke/Llama-2-7B-Chat-GGUF \    llama-2-7b-chat.Q4_K_M.gguf \    --local-dir models/
# Or convert from HuggingFacepython convert_hf_to_gguf.py models/llama-2-7b-chat/

Run inference

bash
# Simple chat./llama-cli \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    -p "Explain quantum computing" \    -n 256  # Max tokens
# Interactive chat./llama-cli \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    --interactive

Server mode

bash
# Start OpenAI-compatible server./llama-server \    -m models/llama-2-7b-chat.Q4_K_M.gguf \    --host 0.0.0.0 \    --port 8080 \    -ngl 32  # Offload 32 layers to GPU
# Client requestcurl http://localhost:8080/v1/chat/completions \  -H "Content-Type: application/json" \  -d '{    "model": "llama-2-7b-chat",    "messages": [{"role": "user", "content": "Hello!"}],    "temperature": 0.7,    "max_tokens": 100  }'

Quantization formats

GGUF format overview

FormatBitsSize (7B)SpeedQualityUse Case
Q4_K_M4.54.1 GBFastGoodRecommended default
Q4_K_S4.33.9 GBFasterLowerSpeed critical
Q5_K_M5.54.8 GBMediumBetterQuality critical
Q6_K6.55.5 GBSlowerBestMaximum quality
Q8_08.07.0 GBSlowExcellentMinimal degradation
Q2_K2.52.7 GBFastestPoorTesting only

Choosing quantization

bash
# General use (balanced)Q4_K_M  # 4-bit, medium quality
# Maximum speed (more degradation)Q2_K or Q3_K_M
# Maximum quality (slower)Q6_K or Q8_0
# Very large models (70B, 405B)Q3_K_M or Q4_K_S  # Lower bits to fit in memory

Hardware acceleration

Apple Silicon (Metal)

bash
# Build with Metalmake LLAMA_METAL=1
# Run with GPU acceleration (automatic)./llama-cli -m model.gguf -ngl 999  # Offload all layers
# Performance: M3 Max 40-60 tokens/sec (Llama 2-7B Q4_K_M)

NVIDIA GPUs (CUDA)

bash
# Build with CUDAmake LLAMA_CUDA=1
# Offload layers to GPU./llama-cli -m model.gguf -ngl 35  # Offload 35/40 layers
# Hybrid CPU+GPU for large models./llama-cli -m llama-70b.Q4_K_M.gguf -ngl 20  # GPU: 20 layers, CPU: rest

AMD GPUs (ROCm)

bash
# Build with ROCmmake LLAMA_HIP=1
# Run with AMD GPU./llama-cli -m model.gguf -ngl 999

Common patterns

Batch processing

bash
# Process multiple prompts from filecat prompts.txt | ./llama-cli \    -m model.gguf \    --batch-size 512 \    -n 100

Constrained generation

bash
# JSON output with grammar./llama-cli \    -m model.gguf \    -p "Generate a person: " \    --grammar-file grammars/json.gbnf
# Outputs valid JSON only

Context size

bash
# Increase context (default 512)./llama-cli \    -m model.gguf \    -c 4096  # 4K context window
# Very long context (if model supports)./llama-cli -m model.gguf -c 32768  # 32K context

Performance benchmarks

CPU performance (Llama 2-7B Q4_K_M)

CPUThreadsSpeedCost
Apple M3 Max1650 tok/s$0 (local)
AMD Ryzen 9 7950X3235 tok/s$0.50/hour
Intel i9-13900K3230 tok/s$0.40/hour
AWS c7i.16xlarge6440 tok/s$2.88/hour

GPU acceleration (Llama 2-7B Q4_K_M)

GPUSpeedvs CPUCost
NVIDIA RTX 4090120 tok/s3-4×$0 (local)
NVIDIA A1080 tok/s2-3×$1.00/hour
AMD MI25070 tok/s2×$2.00/hour
Apple M3 Max (Metal)50 tok/s~Same$0 (local)

Supported models

LLaMA family:

  • Llama 2 (7B, 13B, 70B)
  • Llama 3 (8B, 70B, 405B)
  • Code Llama

Mistral family:

  • Mistral 7B
  • Mixtral 8x7B, 8x22B

Other:

  • Falcon, BLOOM, GPT-J
  • Phi-3, Gemma, Qwen
  • LLaVA (vision), Whisper (audio)

Find models: https://huggingface.co/models?library=gguf

References

  • Quantization Guide [blocked] - GGUF formats, conversion, quality comparison
  • Server Deployment [blocked] - API endpoints, Docker, monitoring
  • Optimization [blocked] - Performance tuning, hybrid CPU+GPU

Resources

Source and attribution

Source:orchestra-research/ai-research-skillsin12-inference-serving/llama-cppat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal