Gptq

orchestra-research/ai-research-skills/10-optimization/gptq

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning.

僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。

路徑大小類型
references/calibration.md8 KBtext/markdown
references/integration.md2.7 KBtext/markdown
references/troubleshooting.md1.9 KBtext/markdown
SKILL.md11.3 KBtext/markdown

來源與署名

來源:orchestra-research/ai-research-skills位於10-optimization/gptq提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架