Constitutional Ai

orchestra-research/ai-research-skills/07-safety-alignment/constitutional-ai

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.

Instructions onlyAI & AgentsSecurity

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
SKILL.md8 KBtext/markdown

Source and attribution

Source:orchestra-research/ai-research-skillsin07-safety-alignment/constitutional-aiat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal