01
Alignment2022
Constitutional AI: Harmlessness from AI Feedback
Bai, Kadavath, et al. — Anthropic
Alignment without an army of human labelers. Worth reading for the governance pattern alone.
Original paper ↗
Summary · Glossary
The white papers that actually moved the field, each linked to the original and paired with my plain-language summary, a glossary of the terms, and what it changes in practice.
01
Bai, Kadavath, et al. — Anthropic
Alignment without an army of human labelers. Worth reading for the governance pattern alone.
Summary · Glossary
02
Rafailov, Sharma, et al. — Stanford
RLHF without the RL. The math is dense; the summary isn’t. If you fine-tune in production, this is the one to understand.
Summary
All original papers remain the work of their authors and link to the official source (arXiv or publisher). Summaries and glossaries on this site are my own commentary.