Gptq

orchestra-research/ai-research-skills/10-optimization/gptq

作者 orchestra-research773a52944ba4MIT13K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3个月前更新

Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning.

  1. 773a52944ba4当前提交 773a529发布于 2026年10月8日

来源与署名

来源:orchestra-research/ai-research-skills位于10-optimization/gptq提交773a529

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架