Huggingface Local Models

作者 huggingfaceabc20ae526d8無授權條款11K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

僅含說明AI & Agents
AI 產生的概覽

指導在本機以 llama.cpp 挑選並執行 GGUF 模型,涵蓋 Hub 搜尋、量化選擇與本機服務。

功能
此技能說明如何在 Hugging Face Hub 上尋找相容於 llama.cpp 的 GGUF 儲存庫、挑選合適的量化版本,並以 llama-cli 或 llama-server 啟動模型。內容涵蓋精確查詢 GGUF 檔名、儲存庫使用自訂命名時的回退參數、在沒有 GGUF 時從 Transformers 權重轉換,以及對本機 OpenAI 相容端點進行冒煙測試。它也指向關於 Hub 探索、量化與硬體加速的參考檔案。
適用情境
當你想在 CPU、Mac Metal、CUDA 或 ROCm 上以 llama.cpp 和 GGUF 在本機執行語言模型時使用。它適合為記憶體預算挑選量化版本、在儲存庫中定位確切的 GGUF 檔案,或啟動本機 OpenAI 相容服務等任務。
執行需求
需要 llama.cpp(透過 brew、winget 安裝或從原始碼建置)以及存取 Hugging Face Hub 的網路;受限制的儲存庫需要 Hugging Face 身分驗證。轉換需要 Python、convert_hf_to_gguf.py 和 llama-quantize。不附帶指令碼,僅為說明文件,另有三份參考文件。

Hugging Face Local Models

Search the Hugging Face Hub for llama.cpp-compatible GGUF repos, choose the right quant, and launch the model with llama-cli or llama-server.

Default Workflow

  1. Search the Hub with apps=llama.cpp.
  2. Open https://huggingface.co/<repo>?local-app=llama.cpp.
  3. Prefer the exact HF local-app snippet and quant recommendation when it is visible.
  4. Confirm exact .gguf filenames with https://huggingface.co/api/models/<repo>/tree/main?recursive=true.
  5. Launch with llama-cli -hf <repo>:<QUANT> or llama-server -hf <repo>:<QUANT>.
  6. Fall back to --hf-repo plus --hf-file when the repo uses custom file naming.
  7. Convert from Transformers weights only if the repo does not already expose GGUF files.

Quick Start

Install llama.cpp

bash
brew install llama.cppwinget install llama.cpp
bash
git clone https://github.com/ggml-org/llama.cppcd llama.cppmake

Authenticate for gated repos

bash
hf auth login

Search the Hub

text
https://huggingface.co/models?apps=llama.cpp&sort=trendinghttps://huggingface.co/models?search=Qwen3.6&apps=llama.cpp&sort=trendinghttps://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending

Run directly from the Hub

bash
llama-cli -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_Mllama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M

Run an exact GGUF file

bash
llama-server \    --hf-repo unsloth/Qwen3.6-35B-A3B-GGUF \    --hf-file Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \    -c 4096

Convert only when no GGUF is available

bash
hf download <repo-without-gguf> --local-dir ./model-srcpython convert_hf_to_gguf.py ./model-src \    --outfile model-f16.gguf \    --outtype f16llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M

Smoke test a local server

bash
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
bash
curl http://localhost:8080/v1/chat/completions \  -H "Content-Type: application/json" \  -H "Authorization: Bearer no-key" \  -d '{    "messages": [      {"role": "user", "content": "Write a limerick about exception handling"}    ]  }'

Quant Choice

  • Prefer the exact quant that HF marks as compatible on the ?local-app=llama.cpp page.
  • Keep repo-native labels such as UD-Q4_K_M instead of normalizing them.
  • Default to Q4_K_M unless the repo page or hardware profile suggests otherwise.
  • Prefer Q5_K_M or Q6_K for code or technical workloads when memory allows.
  • Consider Q3_K_M, Q4_K_S, or repo-specific IQ / UD-* variants for tighter RAM or VRAM budgets.
  • Treat mmproj-*.gguf files as projector weights, not the main checkpoint.

Load References

  • Read hub-discovery.md [blocked] for URL-first workflows, model search, tree API extraction, and command reconstruction.
  • Read quantization.md [blocked] for format tables, model scaling, quality tradeoffs, and imatrix.
  • Read hardware.md [blocked] for Metal, CUDA, ROCm, or CPU build and acceleration details.

Resources

  • llama.cpp: https://github.com/ggml-org/llama.cpp
  • Hugging Face GGUF + llama.cpp docs: https://huggingface.co/docs/hub/gguf-llamacpp
  • Hugging Face Local Apps docs: https://huggingface.co/docs/hub/main/local-apps
  • Hugging Face Local Agents docs: https://huggingface.co/docs/hub/agents-local
  • GGUF converter Space: https://huggingface.co/spaces/ggml-org/gguf-my-repo

來源與署名

來源:huggingface/skills位於skills/huggingface-local-models提交abc20ae

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架