Preference Optimization

作者 wshobson46891e7e60da無授權條款40K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 天前更新

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

僅含說明AI & Agents
  1. 46891e7e60da目前提交 46891e7發布於 2026年10月8日

來源與署名

來源:wshobson/agents位於plugins/llm-finetuning/skills/preference-optimization提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架