---
title: 'Vicuna-moderator-7B: Safety & Moderation Insights'
url: https://www.emergentmind.com/topics/vicuna-moderator-7b
type: topic
---

# Vicuna-moderator-7B: Safety & Moderation Insights

Vicuna-moderator-7B is an informally designated reference to the built-in moderation and safety behavior of the Vicuna-7B v1.5 large language model (LLM), as probed in the context of forbidden task robustness in "In-Context Learning Can Re-learn Forbidden Tasks" [2402.05723]. Derived from Llama2-7B using QLoRA and LoRA techniques on approximately 70,000 ShareGPT human–ChatGPT conversation pairs, Vicuna-7B's moderation capability depends primarily on a system prompt prepended at inference, rather than on hard-coded or separately fine-tuned refusal weights.

## 1. Model Origin and Safety Training Mechanism

Vicuna-7B v1.5, a 7B-parameter chat model

Source: https://www.emergentmind.com/topics/vicuna-moderator-7b