Huggingface Local Models

作者 huggingfaceabc20ae526d8无许可证11K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

仅含说明AI & Agents
AI 生成的概览

指导在本地用 llama.cpp 选择并运行 GGUF 模型,涵盖 Hub 搜索、量化选择与本地服务。

功能
该技能介绍如何在 Hugging Face Hub 上查找兼容 llama.cpp 的 GGUF 仓库、选择合适的量化版本,并用 llama-cli 或 llama-server 启动模型。内容涵盖精确查找 GGUF 文件名、在仓库使用自定义命名时的回退参数、在没有 GGUF 时从 Transformers 权重转换,以及对本地 OpenAI 兼容接口做冒烟测试。它还指向关于 Hub 发现、量化和硬件加速的参考文件。
适用场景
当你希望在 CPU、Mac Metal、CUDA 或 ROCm 上用 llama.cpp 和 GGUF 在本地运行语言模型时使用。它适合为内存预算挑选量化版本、在仓库中定位确切的 GGUF 文件,或启动本地 OpenAI 兼容服务等任务。
运行要求
需要 llama.cpp(通过 brew、winget 安装或从源码构建)以及访问 Hugging Face Hub 的网络;受限仓库需要 Hugging Face 身份验证。转换需要 Python、convert_hf_to_gguf.py 和 llama-quantize。不附带脚本,仅为说明文档,另有三份参考文档。

Hugging Face Local Models

Search the Hugging Face Hub for llama.cpp-compatible GGUF repos, choose the right quant, and launch the model with llama-cli or llama-server.

Default Workflow

  1. Search the Hub with apps=llama.cpp.
  2. Open https://huggingface.co/<repo>?local-app=llama.cpp.
  3. Prefer the exact HF local-app snippet and quant recommendation when it is visible.
  4. Confirm exact .gguf filenames with https://huggingface.co/api/models/<repo>/tree/main?recursive=true.
  5. Launch with llama-cli -hf <repo>:<QUANT> or llama-server -hf <repo>:<QUANT>.
  6. Fall back to --hf-repo plus --hf-file when the repo uses custom file naming.
  7. Convert from Transformers weights only if the repo does not already expose GGUF files.

Quick Start

Install llama.cpp

bash
brew install llama.cppwinget install llama.cpp
bash
git clone https://github.com/ggml-org/llama.cppcd llama.cppmake

Authenticate for gated repos

bash
hf auth login

Search the Hub

text
https://huggingface.co/models?apps=llama.cpp&sort=trendinghttps://huggingface.co/models?search=Qwen3.6&apps=llama.cpp&sort=trendinghttps://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending

Run directly from the Hub

bash
llama-cli -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_Mllama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M

Run an exact GGUF file

bash
llama-server \    --hf-repo unsloth/Qwen3.6-35B-A3B-GGUF \    --hf-file Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \    -c 4096

Convert only when no GGUF is available

bash
hf download <repo-without-gguf> --local-dir ./model-srcpython convert_hf_to_gguf.py ./model-src \    --outfile model-f16.gguf \    --outtype f16llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M

Smoke test a local server

bash
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
bash
curl http://localhost:8080/v1/chat/completions \  -H "Content-Type: application/json" \  -H "Authorization: Bearer no-key" \  -d '{    "messages": [      {"role": "user", "content": "Write a limerick about exception handling"}    ]  }'

Quant Choice

  • Prefer the exact quant that HF marks as compatible on the ?local-app=llama.cpp page.
  • Keep repo-native labels such as UD-Q4_K_M instead of normalizing them.
  • Default to Q4_K_M unless the repo page or hardware profile suggests otherwise.
  • Prefer Q5_K_M or Q6_K for code or technical workloads when memory allows.
  • Consider Q3_K_M, Q4_K_S, or repo-specific IQ / UD-* variants for tighter RAM or VRAM budgets.
  • Treat mmproj-*.gguf files as projector weights, not the main checkpoint.

Load References

  • Read hub-discovery.md [blocked] for URL-first workflows, model search, tree API extraction, and command reconstruction.
  • Read quantization.md [blocked] for format tables, model scaling, quality tradeoffs, and imatrix.
  • Read hardware.md [blocked] for Metal, CUDA, ROCm, or CPU build and acceleration details.

Resources

  • llama.cpp: https://github.com/ggml-org/llama.cpp
  • Hugging Face GGUF + llama.cpp docs: https://huggingface.co/docs/hub/gguf-llamacpp
  • Hugging Face Local Apps docs: https://huggingface.co/docs/hub/main/local-apps
  • Hugging Face Local Agents docs: https://huggingface.co/docs/hub/agents-local
  • GGUF converter Space: https://huggingface.co/spaces/ggml-org/gguf-my-repo

来源与署名

来源:huggingface/skills位于skills/huggingface-local-models提交abc20ae

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架