Preference Optimization

作者 wshobson46891e7e60da無授權條款40K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 天前更新

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

僅含說明AI & Agents

僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。

路徑大小類型
references/method-configs.md6.4 KBtext/markdown
SKILL.md7.7 KBtext/markdown

來源與署名

來源:wshobson/agents位於plugins/llm-finetuning/skills/preference-optimization提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架