Constitutional Ai

orchestra-research/ai-research-skills/07-safety-alignment/constitutional-ai

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.

僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。

路徑大小類型
SKILL.md8 KBtext/markdown

來源與署名

來源:orchestra-research/ai-research-skills位於07-safety-alignment/constitutional-ai提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架