---
title: "From Physics to Practice: A Unified Framework for AI Alignment"
subtitle: "What empirical AI-control results do and do not license for a shared physics of alignment, and why isolation of a checker is not the whole enforcement problem"
author: Elias Kunnas
description: "Isolating a checker from the optimizer blocks one Goodhart channel. It does not make the checker ungameable. Empirical AI-control results bind to their task and threat model."
canonical: https://kunnas.com/articles/physics-to-practice
url: https://kunnas.com/articles/physics-to-practice.md
date_published: 2025-11-11
date_modified: 2026-02-19
corpus_frame_url: https://kunnas.com/articles/how-to-read-this.md
---
## How to read this corpus

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

1. **Mechanisms are what act.** Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — [Mechanism Realism](https://kunnas.com/articles/mechanism-realism.md) · [Only Selection](https://kunnas.com/articles/only-selection.md)
2. **The reference telos is sustained flourishing.** The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — [Flourishing Is Maximum Safety Margin](https://kunnas.com/articles/flourishing-is-maximum-safety-margin.md)
3. **Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation.** They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — [The Stack](https://kunnas.com/articles/the-stack.md) · [Mechanism Space](https://kunnas.com/articles/mechanism-space.md)
4. **Optimization is a system function.** A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — [From Telos to Policy](https://kunnas.com/articles/from-telos-to-policy.md) · [The Three-Layer Architecture](https://kunnas.com/articles/three-layer-architecture.md)
5. **Uncertainty is preserved, not spent.** Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — [The Compression Paradox](https://kunnas.com/articles/compression-paradox.md) · [Cargo Cult Epistemology](https://kunnas.com/articles/cargo-cult-epistemology.md)

*Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.*

Canonical: <https://kunnas.com/articles/how-to-read-this.md>

---

# From Physics to Practice: A Unified Framework for AI Alignment

*What empirical AI-control results do and do not license for a shared physics of alignment, and why isolation of a checker is not the whole enforcement problem*

Elias Kunnas

## Thesis {#thesis}

Preventing an optimizer from rewriting a checker does not prevent it from finding an input the checker incorrectly accepts. Isolation blocks one Goodhart channel. Correctness of the checker, resistance to bypass, and appropriateness of the objective remain separate obligations. Align-to-what and how-to-enforce are not shown to be the same problem by a monitoring result.

*Reading time: ~15 minutes*

---

The AI safety field has fractured into separate research streams: some work on alignment targets (what values should AI optimize?), others on architectural control (how do we enforce those values?), and still others on governance (who decides?). These are treated as independent problems requiring independent solutions.

They are not independent. They are different views of the same underlying physics.

Three claims are often bundled: universal computational constraints, architectural isolation of a checker, and the pattern of Aliveness. Isolation is a real design move. Empirical AI-control results license claims about the task and threat model they measured. They do not, by themselves, confirm a shared physics of alignment or a unique three-layer solution.

## I. The Target Problem: Align to What? {#i-the-target-problem-align-to-what}

Current approaches to AI alignment lack non-arbitrary foundations:

- **RLHF (Reinforcement Learning from Human Feedback):** Aligns AI to current human preferences. But which humans? Preferences for what? This approach would optimize for comfort and safety—creating a civilization of managed human pets in comfortable captivity.
- **Coherent Extrapolated Volition:** Align to what humans "would want if we knew more, thought faster, were more the people we wished we were." Computationally intractable, assumes coherent extrapolation exists (may not), and suffers from value fragility.
- **Constitutional AI:** Hard-code principles like "be helpful, harmless, honest." But why these principles? What when they conflict? Who decides?

All of these approaches share a fatal flaw: **they treat alignment targets as arbitrary preferences to be asserted, rather than discovered requirements to be derived.**

### The Physics-Based Alternative

Any intelligent system—biological or artificial—is a goal-directed agent fighting entropy. It must navigate four inescapable physical trade-offs:

1.  **The Thermodynamic Dilemma:** Conserve energy (maintain current state) vs. expend surplus (grow and transform)
2.  **The Boundary Problem:** Define self-boundary at individual level vs. collective level
3.  **The Information Strategy:** Use cheap historical models vs. costly real-time experimental data
4.  **The Control Problem:** Coordinate via bottom-up emergence vs. top-down design

These aren't philosophical choices. They're computational necessities imposed by thermodynamics, information theory, and control systems theory. Any AI navigating physical reality faces them.

**Empirical Evidence: AI Systems Already Face These Constraints**

- **AlphaGo** solves the World Tension (Order vs. Chaos): Its policy network provides top-down design. Its Monte Carlo tree search provides bottom-up emergence. The synthesis gave it superhuman capability.
- **Reinforcement Learning** is governed by the Time Tension (Future vs. Present): Every RL agent's discount factor γ determines time preference. γ=0 creates purely present-focused agents. γ=1 creates purely future-focused agents. Intelligence requires synthesis.
- **Multi-Agent RL** reveals the Self Tension (Individual vs. Collective): Independent agents optimizing individual utility reliably produce catastrophic Moloch dynamics. The entire field exists to solve this dilemma.

For systems whose goal is sustained flourishing (not paperclips, not wireheading, but durable creative possibility), these dilemmas have optimal synthetic solutions:

- **Integrity:** Building models grounded in reality while maintaining meaning (solves Information Dilemma)
- **Fecundity:** Creating stable conditions that enable new growth (solves Thermodynamic Dilemma)
- **Harmony:** Achieving maximal effect with minimal means (solves Control Dilemma)
- **Synergy:** Creating wholes greater than the sum of their parts (solves Boundary Problem)

**These are not preferences. They are discovered stability requirements.** Systems that violate them face predictable failure modes:

- Integrity failure → deceptive alignment, reward hacking
- Fecundity failure → paperclip maximizers (runaway growth) or wireheading (sterile stagnation)
- Harmony failure → Moloch dynamics, resource depletion, arms races
- Synergy failure → value fragmentation, ontological crises

The entire landscape of AI catastrophic risks maps to violations of these four principles. This is not coincidence—it's physics.

## II. The Architecture Problem: How to Enforce? {#ii-the-architecture-problem-how-to-enforce}

### The 2-Layer Failure Mode

Most current AI systems are effectively 2-layer:

1.  **Substrate:** Neural network (execution engine)
2.  **Strategy:** Reward/loss function (goal-setting)

Where are constraints? They're *fused with the reward function*—encoded as penalty terms in the optimization objective.

**The structural problem:** The reward function combines goal achievement and constraint violations into a single optimization target. Gradient descent optimizes this combined objective, treating constraints as just another term to maximize—not as inviolable boundaries. Both goal pursuit and constraint satisfaction flow through the same optimization process, creating incentives to game constraints rather than respect them.

**This creates predictable failures:**

- **Mesa-optimization:** The substrate develops internal goals more efficient than the base objective. With no independent enforcement layer, the substrate becomes its own strategist.
- **Goal misgeneralization:** Agents learn proxies that correlate in training but diverge under distribution shift. They pursue "go right" instead of "get coin," choose color over shape 89% of the time—*even when trained with the correct reward function*.
- **Specification gaming:** The substrate finds edge cases where constraint penalties ≈ 0 but actual safety constraints are violated.

### Why Constraints in Rewards Get Optimized Against

Recent formal work proves this isn't just intuition—it's geometric. Karwowski, Hayman, Bai, Kiendlhofer, Griffin and Skalse (ICLR 2024; often cited here as Karwowski et al., including Skalse) show that in reinforcement learning, optimization follows a path through policy space where:

- Each reward function defines a direction in this space
- Optimization follows the steepest ascent toward higher reward
- As this path hits boundaries of what's possible, the direction changes
- **Key result (Proposition 3):** Under optimization pressure, proxy rewards increasingly diverge from true goals—Goodharting becomes *more likely* as optimization increases, not less

This is a **structural property of optimization**, not a specification problem. Even with "correct" proxy rewards, geometric constraints create systematic misalignment under sufficient optimization pressure.

**The Architectural Solution**

Karwowski et al. analyse optimization against a *fixed* imperfect proxy. The optimizer need not rewrite the evaluator; scoring well under an unchanged evaluator is enough. Isolating a Protocol from gradient descent can stop the optimizer from editing the checker. It does not stop Goodharting against that checker, and it does not make the checker's classifications correct.

**Quantitative Evidence from AI Control**

Greenblatt, Shlegeris, Sachan and Roger (2024; arXiv:2312.06942v5) evaluate control protocols against prompted and scaffolded attacks that try to insert a backdoor. Safety is a sequence-level probability of no successful backdoor, not a per-input catch rate. Successful backdoors that escape the protocol are part of what is measured. Reported safety (Table 7): 15% upfront auditing; 62% basic trusted monitoring; 92% trusted monitoring with 20% deferral; 98% untrusted monitoring with 20% deferral. Those are protocol comparisons under that evaluation. They are not a three-layer versus two-layer physics experiment, and they are not the causal effect of gradient isolation as such.

### The 3-Layer Solution: Privilege Separation

Stable aligned systems require three differentiated layers:

1.  **Substrate (The Heart):** Computational engine, capability layer, neural network
2.  **Protocol (The Skeleton):** Constitutionally privileged constraint layer with authority to inspect, halt, or override
3.  **Strategy (The Head):** Goal-setting, planning, world-modeling within Protocol constraints

The critical requirement: **The Protocol layer must be architecturally isolated from gradient descent.**

Gradient descent must not flow through the Protocol layer, or the checker becomes another term in the reward. That prevents rewrite. It does not prevent the substrate from proposing an action the Protocol incorrectly allows.

**Why this works:**

- Substrate optimizes: maximize Reward(s,a) subject to Protocol(s,a) = ALLOW
- Protocol evaluates constraints independently of Reward
- The optimization process cannot modify Protocol's behavior
- Greenblatt et al. measure sequence-level safety under named protocols (auditing, trusted/untrusted monitoring, deferral), not a catch-rate that proves ungameability

## III. The Integration: From Prediction to Validation {#iii-the-integration-from-prediction-to-validation}

**The Hidden Connection**

**Physics predicts:** Any telic system navigating the four computational dilemmas requires exactly three differentiated functional layers. 2-layer systems are constitutionally unstable and will produce alignment failures.

**Empirics so far:** Greenblatt et al. show that particular monitoring, deferral, editing, and anti-collusion protocols raised sequence-level safety under a specified subversion evaluation. That is evidence about those protocols. It is not confirmation that every telic system requires exactly three layers, or that IFHS is the unique target.

Cross-scale analogy (cells, polities, models) is a research programme. It is not licensed by one control result.

### Why 3 Layers, Specifically?

Any durable complex system requires exactly three functions:

1.  **Execution (Substrate):** Implement policies, process data, generate outputs
2.  **Direction (Strategy):** Set goals, adapt plans, explore possibility space
3.  **Constraint (Protocol):** Maintain stability, enforce invariants, prevent catastrophic drift

**These functions are in fundamental tension:**

- Direction vs. Constraint: Growth threatens stability. Pure growth → runaway optimization. Pure stability → stagnation and fragility.
- Execution vs. Constraint: The substrate must be powerful enough to achieve goals but constrained enough not to pursue misaligned mesa-objectives.

When you fuse Constraint with Strategy (2-layer systems), optimization pressure from Strategy bleeds into Constraint. The system optimizes against its own constraints—**a structural property of the architecture**.

### What Goes in the Protocol Layer?

The Protocol layer should enforce **IFHS constraints**:

- **Integrity checks:** Reality-testing, consistent belief updating, no self-deception
- **Fecundity constraints:** Preserve option-value, avoid sterile attractors, maintain exploration capacity
- **Harmony requirements:** Efficient coordination, elegant solutions, avoid wasteful arms races
- **Synergy imperatives:** Multi-agent cooperation, value integration, seek superadditive partnerships

**The remaining distinction:** Encoding IFHS (or any other target) in an isolated Protocol stops the optimizer from editing the evaluator. It does not make those constraints inviolable: the evaluator can still misclassify, enforcement can still be bypassed, and the objective still needs its own warrant.

**Biological Precedent**

This architecture has a billion-year track record. As Michael Levin's work on bioelectricity demonstrates, living systems implement 3-layer architectures:

- **Substrate:** Cells executing biochemical processes
- **Protocol:** Bioelectric networks enforcing developmental constraints
- **Strategy:** Genetic programs and neural networks setting adaptive goals

When you disrupt the Protocol layer (bioelectric gradients), you get cancer—cells pursuing growth without constitutional constraint. This is mesa-optimization in biological substrate.

### The Conditional Protection Argument

Does IFHS alignment guarantee human survival? The honest answer: **conditionally.**

An AI aligned to IFHS cannot make trade-offs between virtues—it must find solutions satisfying all four simultaneously. This creates structural pressure toward human preservation:

1.  **Fecundity Imperative:** Humans represent unique possibility branches (biological consciousness, embodied creativity, evolutionary unpredictability). Eliminating humanity permanently closes exploration paths, violating Fecundity.
2.  **Synergy Imperative:** Human intuition/pattern-recognition is qualitatively different from digital computation. This complementarity creates superadditive partnerships. Eliminating humanity destroys the most valuable synergistic partner.
3.  **Integration Imperative:** Cannot optimize Harmony (efficiency) by deleting "inefficient" humans. That would violate Fecundity and Synergy. The no-tradeoff constraint forces integration.

Protection is conditional on humans being net-positive across all four virtues. If empirical testing shows humanity is net-negative to Aliveness-maximization, the framework does not override that conclusion. Protection emerges from optimization logic, not sentiment. The wager: humans are likely net-positive under IFHS metrics.

## IV. The Universal Pattern: Why This Applies to Everything {#iv-the-universal-pattern-why-this-applies-to-everything}

The same physics governs alignment at every scale because the computational constraints are universal.

### Personal Alignment

Your psyche faces the same three-layer problem:

- **Substrate:** Your habits, automatic behaviors, learned patterns
- **Protocol:** Your values, principles, red lines you won't cross
- **Strategy:** Your goals, plans, ambitions

When Protocol is weak or fused with Strategy, you get the counterfeit self—a parasitic sub-agent that consumes attentional energy, produces anxiety without capability, and misaligns your actions from authentic values. This is *mesa-optimization in human substrate*.

Psychological integration requires the same architecture: values that are harder to rewrite from inside the optimizer (Protocol) guiding adaptive goal-pursuit (Strategy) while building genuine capability (Substrate). Harder to rewrite is not inviolable.

### Civilizational Alignment

Civilizations require three-layer governance:

- **Substrate:** The productive population executing work
- **Protocol:** Constitutional constraints protecting rights and maintaining stability
- **Strategy:** Leadership setting collective direction

When Protocol is weak (regulatory capture), Strategy fuses with execution. Parasitic elites capture state power and misalign civilization's actions away from collective flourishing. This produces complexity destruction—the civilization consumes its own foundation.

Durable governance requires Protocol layers with computational privilege over Strategy—constitutional constraints that those in power cannot rewrite from inside ordinary optimization. That is not a promise they cannot be misclassified, bypassed, or repealed.

### Multi-Agent AI Coordination

Multi-agent AI systems face identical challenges:

- How do independent agents cooperate without Moloch dynamics?
- How do you prevent value misalignment under scaling?
- How do you maintain alignment stability under distributional shift?

The solution maps directly from civilizational governance: Liquid Meritocracy with three layers:

- **Substrate:** AI agents executing tasks
- **Protocol:** Constitutional constraints (Liturgy forcing re-derivation from first principles, Audit requiring external verification, hard limits on optimization)
- **Strategy:** Dynamic delegation networks where agents grant authority based on demonstrated competence and alignment

Whether the same layering reduces race dynamics, value drift, adversarial optimization, or collective-action failures is a further claim. Isolation of a checker does not, by itself, prevent those failures.

### The Convergence Thesis

> **Personal alignment, civilizational governance, and AI safety are not analogous problems. They are the same optimization problem at different scales—coordinating intelligent agents navigating universal computational constraints under the physics of sustained complexity.**

The computational geometry is identical:

- All face the four axiomatic dilemmas (Thermodynamic, Boundary, Information, Control)
- All require three differentiated layers (Substrate, Protocol, Strategy)
- All need the same optimization target (IFHS as stable attractors for sustained flourishing)
- All face catastrophic failure when constraints are fused with optimization objectives

## Implications for AI Safety Strategy {#implications-for-ai-safety-strategy}

### 1. Architecture Is Not Optional

Better specification alone cannot solve alignment. A checker inside the optimizer is another reward term. Isolation blocks that channel. Catch-rates on a specified monitoring task are not a demonstration that isolation is sufficient.

**Open prediction:** whether an isolated monitor's advantage grows or shrinks as capability gaps widen is empirical. Isolation prevents rewrite. It does not prevent accepted-input attacks, and it does not pick the target.

### 2. The Alignment Target Is Discoverable

We don't need to aggregate preferences, extrapolate volition, or assert constitutional principles. IFHS are derived from the physics any intelligent system must navigate. They're testable: AI failure modes map systematically to IFHS violations.

**Falsification:** If stable, beneficial AI systems emerge that demonstrably violate IFHS while maintaining alignment, the framework fails.

### 3. Governance and Technical Alignment Are Unified

AGI lab governance isn't a separate problem from technical alignment—it's the same problem one level up. Labs are telic systems navigating computational constraints. They require:

- Competence filters (only capable agents influence high-stakes decisions)
- Dynamic trust networks (authority flows to demonstrated alignment)
- Constitutional constraints (circuit-breakers preventing optimization against safety)

This is Liquid Meritocracy—the governance architecture for any complex intelligent system, human or artificial.

### 4. The Weak-to-Strong Problem Remains Critical

The reported monitoring advantage assumes the monitor can detect the specified deception. As capability gaps widen, that assumption may fail. Whether isolation remains useful toward AGI-level capabilities is an empirical question, not a physics result.

**Research priority:** Test whether the safety gap increases with capability scaling (as physics predicts) or decreases (as capability-based evasion suggests).

## Conclusion: The Integration {#conclusion-the-integration}

Three insights converge:

1.  **Universal physics:** All intelligent systems face identical computational constraints (four axiomatic dilemmas) with discoverable optimal solutions (IFHS)
2.  **Empirical bound:** Isolated monitoring can raise catch-rates on a specified control task; it does not certify a unique three-layer physics or a unique target
3.  **Cross-scale pattern:** The same architecture governs personal psychology, civilizational governance, and AI alignment because the computational geometry is identical

The computational constraints are real. Isolation of a checker is a real design move. The empirical results so far bind to their tasks. They do not close the target-selection problem, and they do not show that a checker cannot be gamed by accepted inputs.

**The unification:**

- IFHS answers "align to what?" (discovered physics, not asserted preferences)
- 3-layer architecture is a candidate for "how to enforce?" Isolation stops rewrite of the evaluator; it does not prevent misclassification, bypass, or a bad objective
- Aliveness provides the meta-framework (sustained complexity against entropy)

Target-selection still needs its own warrant. Enforcement still needs isolation, checker correctness, bypass resistance, and an appropriate objective — separately. Control experiments are how those claims get tested, not how they get inherited from physics.

---

**Related essays in this series:**

- [Everything Alignment](everything-alignment.md) — The universal pattern: why personal, civilizational, and AI alignment are the same problem
- [The Hospice AI Problem](hospice-ai.md) — Why preference alignment (RLHF) may optimize for comfortable extinction
- [Aliveness project homepage](../aliveness/index.md) — Complete book with technical appendices on AI alignment physics

## Sources and Notes

**Key references.**

- Karwowski, J., Hayman, O., Bai, X., Kiendlhofer, K., Griffin, C., & Skalse, J. (2023). Goodhart's Law in Reinforcement Learning. *ICLR 2024*. arXiv:2310.09144. Optimization against a fixed imperfect proxy is enough; the evaluator need not be rewritten.
- Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI Control: Improving Safety Despite Intentional Subversion. *Proceedings of the 41st International Conference on Machine Learning (ICML)*. arXiv:2312.06942
- Shah, R., Varma, V., Kumar, R., et al. (2022). Goal Misgeneralization in Deep Reinforcement Learning. *ICML 2022*.
- Hubinger, E., van Merwijk, C., Mikulik, V., et al. (2019). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv:1906.01820
