AI Rookies

Model censorship — Model Censorship

Fact

Rules that limit what an AI model is allowed to say.

In Plain Words

Model censorship is like a parent with the TV remote. Things get too spicy, and click, the show is over.

It helps with safety, legal rules, and brand image. It can also block normal questions by mistake.

Related Concepts

Alignment
Model censorship often shows the output limits set by Alignment.

RLHF
RLHF can train the model on what to refuse.

Constitutional AI
Constitutional AI uses rule lists to say what the model should refuse.

Jailbreak
Jailbreak prompts often test and bypass model censorship.