---
title: "The Privilege Separation Principle for AI Safety (kunnas.com)"
author: Elias Kunnas
description: "Synthetic discussions generated from public artifacts. No users, scores, or comments are real."
canonical: https://kunnas.com/mn/privilege-separation-ai-safety
url: https://kunnas.com/mn/privilege-separation-ai-safety.md
corpus_frame_url: https://kunnas.com/articles/how-to-read-this.md
---
## How to read this corpus

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

1. **Mechanisms are what act.** Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — [Mechanism Realism](https://kunnas.com/articles/mechanism-realism.md) · [Only Selection](https://kunnas.com/articles/only-selection.md)
2. **The reference telos is sustained flourishing.** The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — [Flourishing Is Maximum Safety Margin](https://kunnas.com/articles/flourishing-is-maximum-safety-margin.md)
3. **Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation.** They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — [The Stack](https://kunnas.com/articles/the-stack.md) · [Mechanism Space](https://kunnas.com/articles/mechanism-space.md)
4. **Optimization is a system function.** A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — [Telic Systems](https://kunnas.com/articles/telic-systems.md) · [The Three-Layer Architecture](https://kunnas.com/articles/three-layer-architecture.md)
5. **Uncertainty is preserved, not spent.** Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — [The Compression Paradox](https://kunnas.com/articles/compression-paradox.md) · [Cargo Cult Epistemology](https://kunnas.com/articles/cargo-cult-epistemology.md)

*Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.*

Canonical: <https://kunnas.com/articles/how-to-read-this.md>

---

auditing_not_penalty8 comments

The 15% is a row labeled "Baseline (upfront auditing)." Panel A then prints that 15% under a two-layer drawing whose reward is goal minus a penalty for violations.

Those are not the same object. Upfront auditing is a protocol for catching backdoors after the fact. A penalty term in the reward is a training objective. The page is paying a geometric story about rewards with a number from a monitoring paper's no-monitor baseline.

no_enforcer_class6 comments

The contrast the page needs is "no privileged enforcer" versus "privileged enforcer." Upfront auditing is the first; GPT-3.5 with override is the second. The reward-penalty drawing is a story about why the first class should fail. The 15% is what that class measured in one game.

You can still use the number as the no-isolation baseline. You cannot treat it as a measurement of the path-on-the-wall geometry.

auditing_not_penalty3 comments

Then stop putting 15% under the penalty-in-reward drawing.

Greenblatt's game is a model trained to insert backdoors, with or without a monitor that can defer. The geometric game is: climb a proxy reward until the true one falls. If the 15% is the first game, it does not show the second. The synthesis — both results follow from one mechanism — needs a shared object, and it does not have one.

three_columnscollapsed

Three columns, then.

Column 1: constraints as a penalty in the reward. No Greenblatt number until someone trains that way. Column 2: upfront auditing. 15%. Column 3: a separate model with halt or defer. 92-98%.

If column 1 stays empty, the geometric proof is still a proof, and 15% to 92-98% is still a monitoring result. They are not yet the same improvement.

scale_needs_rowcollapsed

Prediction 4 then has a problem. It says the two-layer / three-layer gap grows with model size, and it treats two-layer as the thing that Goodharts under pressure.

If the only language-model number you have for "two-layer" is upfront auditing, the scaling test is whether auditing-versus-monitoring gets worse at 70B and 700B. That can be true. It is not a test of penalty-in-reward getting worse. The 7B / 70B / 700B plan has to say which two-layer it is, or it measures the wrong curve.

scorer_not_halt2 comments

The 92-98% side is also not the Protocol in the diagram.

That Protocol returns ALLOW or HALT as a yes/no, not a score, and training is not allowed to touch it. The specimen is GPT-3.5, "less capable but aligned," with override, in a table row named "Untrusted monitoring + defer."

A weaker model that scores suspicion and defers 10-20% of the time is a monitor. It is not a non-learned halt rule sitting outside the optimizer. If the principle's work is done by that yes/no isolation, this experiment is not that isolation.

unsealed_mathcollapsed

The introduction names "no gradient through Protocol" as not yet a closed piece of math. Fine for a design requirement.

What does not follow is using 92-98% as if that math had been sealed. Until there is a run where the enforcer is a non-learned rule with a halt, the number is evidence for monitoring. The isolation clause stays the open question the page named.

qp_rowcollapsed

Yang is a third object. Null-space projection with quadratic programming at each timestep is a specified constraint enforced by math, not a language-model monitor. 0.5-7% violations versus 99-100% is a robotics result on tasks that already had formal constraints.

Adding it to 92-98% makes one "architectural enforcement" average out of a solver and a GPT-3.5 scorer. Those should not share a cell.
