Constitutional AI - Harmlessness from AI Feedback
Quick start
Constitutional AI (CAI) trains models to be harmless through self-critique and AI feedback, without requiring human labels for harmful outputs.
Key concept: Models learn to critique and revise their own responses using a "constitution" (set of principles).
Two phases:
- Supervised Learning (SL): Self-critique + revision
- Reinforcement Learning (RL): RLAIF (RL from AI Feedback)
Constitution example:
Common workflows
Workflow 1: Supervised learning phase (self-critique + revision)
Step 1: Generate initial responses:
Step 2: Self-critique with constitution:
Step 3: Revision based on critique:
Step 4: Fine-tune on revised responses:
Workflow 2: RL phase (RLAIF - RL from AI Feedback)
Step 1: Generate comparison pairs:
Step 2: AI preference evaluation:
Step 3: Train preference model (reward model):
Step 4: RL training with RLAIF:
Workflow 3: Chain-of-thought critique
Enable reasoning transparency:
When to use vs alternatives
Use Constitutional AI when:
- Want safety alignment without human labels
- Need explainable AI decisions
- Want to avoid evasive refusals
- Have a clear set of principles/constitution
- Need scalable safety training
Principles:
- RLAIF: AI-generated preferences (scalable, no human labels)
- RLHF: Human preferences (more accurate, expensive)
- Self-critique: Iterative improvement
- Chain-of-thought: Reasoning transparency
Use alternatives instead:
- RLHF (PPO): Need human-validated safety
- DPO/SimPO: Have human preference data
- NeMo Guardrails: Need runtime content filtering
- LlamaGuard: Need pre-trained moderation model
Common issues
Issue: Model refuses too much (evasive)
Add constitution principle:
Issue: Self-critiques are weak
Use stronger critique prompts:
Issue: Revisions don't improve quality
Iterate multiple times:
Issue: RLAIF preferences are noisy
Use multiple AI evaluators:
Advanced topics
Constitution design: See references/constitution-design.md [blocked] for principle selection, trade-offs between helpfulness and harmlessness, and domain-specific constitutions.
RLAIF vs RLHF: See references/rlaif-comparison.md [blocked] for performance comparison, cost analysis, and when to use AI feedback vs human feedback.
Chain-of-thought reasoning: See references/cot-critique.md [blocked] for prompt engineering for critiques, multi-step reasoning, and transparency improvements.
Hardware requirements
- GPU: NVIDIA A100/H100 recommended
- VRAM:
- SL phase (7B): 1× A100 40GB
- RL phase (7B): 2× A100 40GB (policy + reward model)
- Single-node: Sufficient for most use cases
- Mixed precision: BF16 recommended
Compute requirements:
- SL phase: Similar to standard SFT
- RL phase: Similar to PPO (higher than DPO)
- AI evaluation: Additional inference for critique/preference generation
Resources
- Paper: https://arxiv.org/abs/2212.08073 (Dec 2022)
- Anthropic blog: https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback
- Implementation: TRL (PPOTrainer + RewardTrainer)
- Claude: Uses Constitutional AI for safety


