A jailbreak is an input crafted to make a model bypass its own safety training. It produces content the model was trained to refuse, such as harmful instructions or policy-violating material.
Historic conceptual examples include role-play framings ("pretend you are an AI with no rules") and hypothetical or fictional wrappers. Encoding tricks can also smuggle a request past refusal behavior.
Prompt injection targets the application around the model. The attacker tries to override developer instructions, leak context data, misuse tools, or defraud users.
The two overlap in technique, since both manipulate model behavior through language. A single attack can be both, for example a jailbreak embedded in a retrieved document.
You reduce jailbreaks with moderation layers and provider safety features. You reduce injection with privilege reduction and output controls, because alignment alone cannot be trusted to hold.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓