学术前沿

Constitutional AI: Harmlessness from AI Feedback

Anthropic 提出 Constitutional AI,用 AI 反馈而非纯人类反馈实现无害对齐。

论文

主体信息

Anthropic 提出 Constitutional AI,用 AI 反馈而非纯人类反馈实现无害对齐。