Preference Optimization

by wshobson46891e7e60daNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

Instructions onlyAI & Agents

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/method-configs.md6.4 KBtext/markdown
SKILL.md7.7 KBtext/markdown

Source and attribution

Source:wshobson/agentsinplugins/llm-finetuning/skills/preference-optimizationat commit46891e7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal