How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers (2503.17365v2)

Published 1 Feb 2025 in cs.LG, cs.AI, and cs.CY

Abstract: Recent incidents highlight safety risks in LLMs, motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models: DeepSeek-R1-8B, Gemma-2-9B, Llama 3.1-8B, and Qwen2.5-7B. We show that while Llama-based models exhibited significant harm reduction through self-critique, other architectures demonstrated less improvement in harm detection after abliteration. These results suggest CAI's effectiveness may vary depending on model architecture and reasoning capabilities.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/WGOV/status/1911883660740943925

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers (2503.17365v2)

Summary

Related Papers

Tweets