LearnThatStack Ace your next interview
AI Security & Guardrails · question
Question 5 of 55

Why is model alignment, such as RLHF safety training, not a security boundary?

beginner
← All AI Security & Guardrails questions
Re-explain

A security boundary must be deterministic, enforced, and fail closed. A permission check either passes or it does not, regardless of how persuasive the request is. Alignment is none of those things.

Safety training shifts the probability distribution of model outputs toward refusing harmful requests. The refusal is a learned behavior that adversarial inputs can and regularly do overcome.

New jailbreak techniques are published continuously, and each model release resets the cat-and-mouse game.

There are structural reasons it cannot be airtight. The model cannot verify who is speaking; any text claiming authority might be an attacker.

The engineering implication: treat model refusals as a valuable defense-in-depth layer and a UX safeguard. Place actual security controls in deterministic systems around the model.

A useful rule: the model can be socially engineered, so never give it authority you would not give a well-meaning but gullible intern.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

The diagram below the answer is the concept . Jump to it ↓

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

Saved in this browser - sign in to keep your review list.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Interview lens

Likely follow-ups, what you can say, and the weak answers to avoid.

Sign in free to open it Free account - the lens opens as soon as you're back.

Want a quick review of the fundamentals? See the AI Security & Guardrails cheatsheet.

← Back to all AI Security & Guardrails questions
Pro · $10/mo

48 of 55 AI Security & Guardrails answers are in Pro.

Full answers, code samples, and AI explanations that go simpler or deeper. Cancel anytime.

  • Full answers + code
  • AI explanations, simpler or deeper
  • 1,000 AI credits / month
  • Cancel anytime