Standard 1 stop to get here · leads to 2
AI Alignment
Ensuring AI systems behave in accordance with human values and intentions, a central challenge in AI safety.
Your route here
1 stop · basics first
- AI Safety ✓ understood
Research and practices aimed at ensuring AI systems are safe, reliable, and beneficial, especially as capabilities increase.
- AI Alignment · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs RLHF Reinforcement Learning from Human Feedback - training models using human preferences to align behavior with human values. Language & LLMs Constitutional AI Training AI systems using principles and rules rather than only human feedback, developed by Anthropic for Claude. Language & LLMs Alignment Tax Performance degradation that may occur when making models safer and more aligned with human values. Shipping AI Interpretability Understanding the internal workings of AI models, including which features influence predictions and why. Shipping AI Fairness Ensuring AI systems treat all individuals and groups equitably, without discrimination based on protected attributes.
In the research
All papers →A paper that builds on AI Alignment .