Moe Training

orchestra-research/ai-research-skills/19-emerging-techniques/moe-training

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.

僅含說明AI & Agents

僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。

路徑大小類型
references/architectures.md13.7 KBtext/markdown
references/inference.md8 KBtext/markdown
references/training.md9.5 KBtext/markdown
SKILL.md14.6 KBtext/markdown

來源與署名

來源:orchestra-research/ai-research-skills位於19-emerging-techniques/moe-training提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架