Anthropic describes Constitutional AI as an approach to shaping model behavior around a set of principles. I understand the broad idea, but I’m not sure how it differs in practice from relying mainly on human feedback during training.
For people following Anthropic’s work, what changes does this approach make to Claude’s responses, especially when a request is ambiguous or touches on a sensitive subject? I’m looking for a clear explanation of the trade-offs, not just the company’s description of the method.