A prompt attack that tries to make an AI ignore its safety rules.
A jailbreak is like telling the lunch lady, "I just need a spoon." Really, you want to sneak behind the counter and ignore the rules.
People use it to ask for banned content or unsafe steps. It is a common chatbot security risk.
Prompt injection
Jailbreaks are like prompt injection because both use input to steer the model off track.
Alignment
A jailbreak tries to get around the model's safety alignment.
System prompt
Many jailbreaks try to override or exploit the system prompt.
Agent Security
When an AI can use tools, a jailbreak becomes a more real security risk.