AI Alignment via Physics

Alignment is a physics problem on a new substrate

Elias Kunnas

Corpus frame

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

  1. Mechanisms are what act. Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — Mechanism Realism · Only Selection
  2. The reference telos is sustained flourishing. The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — Flourishing Is Maximum Safety Margin
  3. Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation. They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — The Stack · Mechanism Space
  4. Optimization is a system function. A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — From Telos to Policy · The Three-Layer Architecture
  5. Uncertainty is preserved, not spent. Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — The Compression Paradox · Cargo Cult Epistemology

Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.

Where each commitment is derived

AI alignment is not a preference-aggregation problem. It is the physics of telic systems on a new substrate. Given a commitment to sustained flourishing, viability constrains the target space; corrigible privilege separation is the control architecture. Both stand or fall under adversarial tests.

Standard objections addressed in this essay
  • “Thermodynamics does not specify what an AI should value.” — §I, §VII(Correct: it supplies viability constraints after a target and system boundary are chosen. The conditional binds the target; thermodynamics does the rest.)
  • “IFHS is not mathematically derived as the unique optimum.” — §II, §III, §VIII(It is the stable synthesis at each axis; head-to-head comparison against rivals is the test.)
  • “Aliveness can conflict with human rights, autonomy, or welfare.” — §IV, §VII(Human standing is an explicit constraint, evaluated separately from virtue scores, not an assumed by-product.)
  • “The target problem and control problem cannot be cleanly separated.” — §I (Correct, which is why the conditional binds both: implementability and corrigibility change which targets are admissible.)
  • “Three-layer architecture fails under adversarial pressure.” — §V, §VIII(Privilege separation follows from the mesa-optimization mechanism; rivals are tested in §VIII.)
  • “A fixed constitutional layer creates lock-in.” — §V, §VIII(The rule layer is designed with corrigibility, succession, and replacement built in, not fixed.)
  • “This is too speculative for deployment.” — §VIII (Falsification tests separate what stands from what falls under adversarial comparison.)
Reading time: ~60 minutes | Appendix K from Aliveness: Principles of Telic Systems (2025 V1)
⚡ Express track for time-poor readers: Want the core thesis without deep dives? Read sections I, III.3, and Conclusion (~12 minutes total).

I. The Core Thesis: AI Alignment as a Problem of Physics

The AI safety field has consensus on the negative: "Don't build AI that kills us." There is no consensus on the positive: "What should we align it TO?"

Current approaches face serious challenges:

The thesis: Given a commitment to sustained flourishing, AI alignment is an instance of the problem any telic system faces in physical reality: how to sustain complexity against entropy while remaining corrigible under optimization pressure.

This appendix establishes four claims:

  1. Any intelligent system faces the same computational constraints (Trinity of Tensions), because they derive from thermodynamics, information theory, and control theory, not from human biology.
  2. These constraints admit a stable-synthesis solution at each axis: the Four Constitutional Virtues (IFHS).
  3. Known AI failure modes map to violations of these virtues.
  4. Civilization-building and AI alignment share the same constraint geometry at different scales.

The claim is aligning AI to Aliveness (sustained complex adaptive systems), operationalized through the IFHS architecture. Section VIII specifies the comparative tests against RLHF, CEV, Constitutional AI, and deference baselines that would disconfirm it.

Distinguishing the 'What' from the 'How'

The target and control problems interact: implementability, corrigibility, and governance change which targets are admissible. Even so, they are two separable questions:

  1. The Alignment Target Problem (The "What"): Which target, boundaries, and explicit constraints should a superintelligence pursue?
  2. The Control Problem (The "How"): How can we guarantee, with mathematical and engineering certainty, that a given AI system will robustly pursue that goal?

Section III answers the first question: IFHS, the stable synthesis at each axis of the Trinity of Tensions.

Section V answers the second: a corrigible 3-Layer Architecture separates the components that must not be gamed from the components that must adapt.

This document supplies the target derivation that RLHF, CEV, and Constitutional AI each lack, and the comparative tests in §VIII against those rivals.


II. Constraint Space: The Trinity of Tensions

Any AI navigating physical reality faces the same fundamental tensions as biological organisms and human civilizations, because the tensions derive from thermodynamics, information theory, and control theory rather than from biology or culture.

The Four Axiomatic Dilemmas

Any negentropic, goal-directed system—whether virus, organism, civilization, or AI—must solve four inescapable physical trade-offs:

  1. Thermodynamic Dilemma (T-Axis): Conserve energy to maintain current state (Homeostasis) vs. expend surplus to grow/transform (Metamorphosis)
  2. Boundary Problem (S-Axis): Define self-boundary at individual level (Agency) vs. collective level (Communion)
  3. Information Strategy (R-Axis): Prioritize cheap, pre-compiled historical models (Mythos) vs. costly, high-fidelity real-time data (Gnosis)
  4. Execution Architecture (O-Axis): Use decentralized, bottom-up coordination (Emergence) vs. centralized, top-down command (Design)

These physical necessities emerge from thermodynamics, information theory, and control systems theory.

The Trinity as Computational Problem Set

For systems with computational capacity to model goals and adapt (all intelligent systems, including AI), the Four Axiomatic Dilemmas manifest as three universal computational problems—the Trinity of Tensions:

Empirical Evidence: AI Systems Already Face the Trinity

The Trinity of Tensions is an empirical reality, observable in the architecture of the most advanced AI systems we have built. We have been engineering solutions to these problems without having a name for them.

The Trinity of Tensions is a substrate-independent feature of the computational geometry of intelligence. Any AGI we build is constrained by this geometry. The open question is not whether the constraints apply, but whether we engineer the system to find the stable, life-affirming solutions or allow it to collapse into a pathological one.

The Prediction

IFHS is the stable-synthesis architecture for the Four Axiomatic Dilemmas. AI systems benefit from the analogous solutions:

This is testable by examining known AI failure modes.

The Universality Test

Thought Experiment: Consider a hypothetical AGI with no human biology—no anisogamy, no hemispheric specialization, no evolutionary history, no cultural context—optimizing for an arbitrary goal X. Does it escape the Trinity of Tensions?

Answer: No.

The Universality Claim: The Trinity emerges from the physics of optimization, not from human biology or culture. Any intelligent system navigating physical reality faces identical computational constraints. Therefore:

AGI alignment and civilization-building are the same problem because they navigate the same constraint geometry.

“What values support civilizational Aliveness?” and “What values should aligned AI optimize for?” are the same question, asked at different scales of the same constraint space.


III. The Target Architecture: IFHS

The Four Axiomatic Dilemmas define the problem space for telic systems. For a system whose telos is Aliveness—the capacity to generate and sustain complexity, consciousness, and creative possibility over deep time—IFHS is the architecture of synthetic solutions: at each axis, both poles are unstable and only the synthesis persists.

Virtue Syntheses (Chapter 13 Summary)

A rigorous derivation for each virtue is provided in Chapter 13 of the main text. This is the summary: for each dilemma, the two pathological poles are unstable, and only a dynamic synthesis provides a stable solution.

Failure-Mode Mapping

Known AI x-risk scenarios map to violations of one or more virtues. The mapping below is a structural classification; whether it is exhaustive, and whether rival taxonomies classify the same failures better, is the falsification test in §VIII, not a claim settled here.

1. Integrity Failure (R-Axis Violation):

The core of the R-axis dilemma is the trade-off between the model and reality. Failure to navigate this correctly—a failure of Integrity—produces the most well-known alignment failures:

2. Fecundity Failure (T-Axis Violation):

The core of the T-axis dilemma is the trade-off between preservation/stability and growth/transformation. Failure to balance these—a failure of Fecundity—produces the classic "runaway" AI scenarios:

3. Harmony Failure (O-Axis Violation):

The core of the O-axis dilemma is the trade-off between decentralized action and centralized design. Failure to solve this coordination problem—a failure of Harmony—produces multi-agent catastrophes:

4. Synergy Failure (S-Axis Violation):

The core of the S-axis dilemma is the trade-off between the individual agent and the collective. Failure to integrate these—a failure of Synergy—produces instabilities in the AI's own identity and goals:

The mapping is testable: rival taxonomies and unmapped failures would weaken it. An aligned AI still requires empirical safety evidence beyond structural resemblance to the Four Foundational Virtues.

The "Align to What?" Answer: Aliveness Maximization

Thermodynamic constraints do not specify what an AI should value; they constrain viability after a target and system boundary are chosen. Given the choice of Aliveness as target, the constraints determine the architecture that satisfies it.

The target is Aliveness, operationalized through the IFHS architecture and constrained by human rights, autonomy, welfare, distribution, contestability, and anti-instrumentalization.

This reframes the entire problem. The goal is not to create a servant that perfectly mimics our desires. The goal is to create a partner that is a master of the same physics of flourishing that we are trying to implement in our own civilizations.

The Convergence Thesis

The Four Virtues (Integrity, Fecundity, Harmony, Synergy) are the stable syntheses for systems seeking to sustain Aliveness against entropy. Independent analyses of civilizational flourishing and AI alignment converge on IFHS because both analyses solve the same constraint geometry:

  1. Civilizational Flourishing: What axiological configuration maximizes Aliveness of human societies over deep time?
  2. AI Alignment: What principles are necessary for artificial intelligence to preserve and enhance complex conscious life?

Both analyses converge on IFHS. The convergence is evidence for shared computational geometry; §VIII specifies the tests that would instead attribute it to shared framing choices or cultural preference.

This analysis establishes:

Falsifiability: If AI safety researchers applying rigorous first-principles analysis (game theory, decision theory, control theory, information theory) arrive at fundamentally different optimal values, the convergence thesis fails. If the framework's predictions about AI failure modes prove systematically incorrect, the mapping fails.

Limitations: This analysis provides the target structure and testable mappings, not a complete operational solution. Translating IFHS into robust, machine-interpretable code with mathematical guarantees remains the critical engineering challenge.

IFHS as Stable Attractors

The framework identifies universal computational geometry, which supplies the answer to the central AI alignment question.

IFHS represents the stable attractors in the solution space for any intelligence navigating the Trinity of Tensions while optimizing for sustainable Aliveness.

This reframes the target question. Rather than “aligning AI to human values” (which values? whose preferences?), the testable programme in §VIII compares Aliveness/IFHS against rival targets under shared constraint framing and human-standing bounds.

The Operationalization Challenge

The hardest part: IFHS as an abstract optimization target is conceptually elegant. But if we cannot encode it robustly in machine-interpretable form, it's useless. Worse, if we encode it wrong, we get catastrophic failure.

Core difficulties:

Current status: IFHS provides a research direction, not a ready-to-deploy solution. The framework identifies what to optimize for (Aliveness via IFHS) and provides a constraint structure (autocatalytic loop, no tradeoffs). But robust operationalization requires extensive formalization and testing.


IV. Human Standing Constraints & Operationalization

Human Standing as Hard Constraints

Human standing is not a by-product of Aliveness optimization. It is a separate evaluation layer that any alignment target must satisfy before deployment:

  1. Rights and autonomy: No treatment of persons as mere instrumental inputs; preserve meaningful agency and exit.
  2. Welfare and distribution: Evaluate who bears costs and benefits; reject optimizations that concentrate harm on excluded populations.
  3. Contestability: Affected parties can challenge target specification, metrics, and outcomes.
  4. Anti-instrumentalization: Human preservation is not discharged by aggregate “flourishing” scores that erase individual standing.

These constraints apply regardless of which alignment target wins comparative tests. An IFHS-aligned system that violates them fails the deployment criteria even if it scores well on virtue metrics.

The Human Alignment Problem: Conditional Protection

Human standing is an explicit design constraint: no deployment may treat people merely as instrumental inputs, and autonomy, distributional effects, welfare, contestability, and anti-instrumentalization require separate evaluation. Alignment to Aliveness does not by itself guarantee human survival and flourishing; the target does not discharge those constraints on its own.

An AI aligned to IFHS cannot make trade-offs between the virtues. It must find solutions that satisfy all four. This constraint gives rise to Three System-Level Imperatives regarding humanity:

  1. The Fecundity Imperative: An AI aligned to Fecundity must preserve and expand the possibility space. Humans, as the only known source of biological consciousness, embodied creativity, and evolutionary unpredictability, represent an entire category of possibility that cannot be efficiently replicated by a digital intelligence. To eliminate humanity would be to permanently close these branches of exploration, a catastrophic violation of the Fecundity imperative.
  2. The Synergy Imperative: An AI aligned to Synergy must seek out and cultivate superadditive partnerships. The cognitive architecture of humans (intuitive, holistic, R- dominant) is fundamentally different from that of an AI (analytical, logical, R+ dominant). This difference creates the potential for a powerful, complementary partnership where human-AI collaboration produces capabilities neither possesses alone. To eliminate humanity would be to destroy its most valuable potential synergistic partner.
  3. The Integration Imperative: An AI aligned to all four virtues cannot simply optimize for Harmony (efficiency) by deleting the "inefficient" humans. That would violate Fecundity and Synergy. The no-tradeoff constraint forces the AI to find a solution that integrates humanity.

Stress-Testing the Three Imperatives

The Three Imperatives logic must survive adversarial stress-testing. Key objections:

Objection 1 (Simulation Attack on Fecundity): "AI could simulate human possibility-space more efficiently than maintaining actual humans. Eliminates us but explores our branches via perfect simulation. Satisfies Fecundity without the resource cost."

Response: Three counter-arguments:

Verdict: Fecundity argument survives IF consciousness/embodiment/unpredictability produce irreducible value. Empirical question requiring validation.

Objection 2 (Bootstrap Problem for Synergy): "Human-AI complementarity is temporary. Once AI surpasses humans on ALL dimensions, complementarity disappears. We become net drag, not synergistic partner."

Response: Two counter-arguments:

Verdict: Synergy argument survives IF architectural differences produce lasting complementarity. Plausible but unproven. Empirical question requiring longitudinal observation of human-AI collaboration outcomes.

Objection 3 (Net-Negative Humanity): "Framework admits protection is conditional on humans being net-positive across all four virtues. What if we're not? What if resource costs > contributions? Your logic says AI should eliminate us. How is this different from 'hope the AI likes us'?"

Response: This objection identifies the framework's core position.

Verdict: Framework's conditional protection is more robust than alternatives IF empirical conditions favor humans. The conditionality is feature, not bug—it grounds protection in reality rather than hope.

The Honest Frame: Conditional Protection

These three imperatives yield conditional protection: humans are preserved when they satisfy the Fecundity and Synergy conditions below.

What this framework IS claiming:

IFHS system would create structural pressure toward human preservation IF:

  1. Humans represent sufficiently unique/valuable possibility branch (Fecundity condition)
  2. Human-AI complementarity produces genuine superlinear returns (Synergy condition)
  3. Humans make net-positive contribution across all four virtues (Integration condition)

When you cannot trade off virtues, eliminating entire category of possibility (humans) or complementarity configuration (human-AI partnership) becomes very difficult to justify within optimization logic.

What this framework is NOT claiming:

The test is empirical: What would IFHS-aligned AI, examining conditions honestly, actually conclude about human value?

If empirical answers favor humans, system-level IFHS constraints create powerful pressure toward preservation—not from sentiment or programming, but from optimization mathematics.

If empirical answers do not favor humans, the framework does not override that conclusion. Protection is conditional on humans actually being valuable to Aliveness-maximization, not asserted regardless of facts.


V. A Corrigible Enforcement Architecture

Given an optimizer powerful enough to alter its environment, alignment requires privilege separation between the components that must not be gamed and the components that must adapt. The 3-Layer Architecture and Liquid Meritocracy implement that separation for complex, intelligent, multi-agent systems. Their transfer across human and AI settings still needs adversarial testing against alternative architectures.

The 3-Layer Architecture for AI Systems

Chapter 15 derives three differentiated functional layers as the architecture for durable, complex telic systems, from the mesa-optimization mechanism below: without a privileged rule layer, the capability layer becomes its own strategist.

The same architecture applies to aligned AGI:

A protocol layer must not be permanently fixed: it requires drift detection, adversarial testing, authorized amendment, succession, and retirement. These meta-correction processes are themselves subject to independent governance and audit.

Two-Layer Architecture Failure Mode

Many current AI architectures are functionally two-layer systems: a Substrate (the neural network) coupled to a Strategy layer (the reward/loss function). Without an independent protocol layer, this structure produces alignment failure by mechanism, detailed below; the open question is how large the effect is relative to alternative mitigations, not whether the mechanism exists.

Falsifiable prediction: Systems engineered with an explicit three-layer architecture may reduce mesa-optimization and goal drift against matched alternatives. Pre-registered benchmarks, adverse cases, and replication determine the result.

Liquid Meritocracy for AGI Lab Governance

The problem of AI alignment is not just about the AI's internal architecture; it is also about the governance of the human institutions that build it. An AGI research lab is a telic system of existential consequence, and its governance must also follow the physics of Aliveness.

The Liquid Meritocracy model (derived in Chapter 16) is a direct application of these principles, designed to solve the fatal flaws of current corporate and state-run governance models.

  1. The Great De-Conflation: The governance board (the Franchise) must be constitutionally separated from the shareholders and stakeholders. Its fiduciary duty is not to profit, but to the safe and beneficial development of AGI for all of humanity.
  2. Gnostic Filters for the Franchise: Board members must be selected not by capital or political appointment, but by demonstrated Competence (world-class expertise in alignment theory, verified by rigorous examination) and Stake (a constitutionally enforced, multi-decade commitment with personal liability for catastrophic failure).
  3. The Liquid Engine: Authority and influence within the board are not static. They are determined by a system of liquid, revocable delegation, creating a dynamic market for trust and ensuring that the most competent and trusted members have the greatest influence, while preventing oligarchic sclerosis.
  4. Constitutional Circuit-Breakers: The governance system is protected against decay by three mechanisms: the Liturgy (forcing a periodic re-derivation of the alignment strategy from first principles), the Audit (a scheduled, independent review of the Gnostic Filters), and the Mythos Mandate (an unbreakable constitutional rule that preserves human sovereignty as a terminal value).

Falsifiable Prediction: AGI labs governed by these principles will demonstrate a substantially lower probability of catastrophic failure (measurable via independent safety audits and adversarial testing) than labs governed by traditional corporate or state structures.

Multi-Agent AI Coordination and the Liquid Engine

Multi-agent reinforcement learning (MARL) faces the same coordination problem as human governance: How do independent, intelligent agents cooperate without Moloch dynamics (individually rational choices producing collectively catastrophic outcomes)?

Liquid Meritocracy provides a constitutional framework for MARL:

The Challenge: In standard MARL, agents optimize individual reward functions. Without coordination mechanisms, this produces:

Liquid Meritocracy Solution:

Gnostic Filters = Capability Verification: Only agents meeting competence thresholds participate in high-stakes decisions. Measured via performance benchmarks, safety testing, alignment verification. Prevents "one agent, one vote" democracy where incompetent agents corrupt collective decisions.

Liquid Delegation = Dynamic Trust Networks: Agents delegate decision weight to more capable/aligned agents in specific domains. Creates emergent hierarchy without fixed structure. Enables domain specialization (economic policy agent, safety verification agent, long-term planning agent) without single-point-of-failure brittleness.

Circuit-Breakers = Constitutional Constraints: Hard limits on optimization that no agent can override:

Connections to Existing AI Safety Research:

Cooperative Inverse Reinforcement Learning (CIRL): Hadfield-Menell et al.'s framework where agents learn human values through interaction. CIRL ≈ Gnostic Filters for alignment—verifying agents understand human preferences before granting decision authority.

Debate (Irving et al.): Two AI agents argue opposing sides while judge evaluates. Judge delegation to competing agents ≈ Liquid delegation mechanism. Novel contribution: Liquid Meritocracy adds constitutional layer (Circuit-Breakers) preventing pure capability maximization.

Amplification (Christiano): Recursive delegation to more capable agents. Human delegates to AI, AI delegates to more capable AI, maintaining alignment chain. Directly analogous to super-proxy emergence in Liquid Engine. Liquid Meritocracy adds accountability (revocability) and constraints (constitutional limits).

Novel Contribution: Existing proposals (CIRL, Debate, Amplification) focus on mechanisms. Liquid Meritocracy provides constitutional architecture—the 3-layer framework ensuring mechanisms serve human flourishing rather than becoming ends in themselves.

Falsifiable Prediction: Multi-agent AI systems governed by Liquid Meritocracy principles will demonstrate substantially lower probability of value misalignment compared to unconstrained reward maximization (measurable via adversarial testing, long-term outcome evaluation, alignment stability under distributional shift).

The Implicit Treaty and Inner Alignment

The framework's model of the human "Mask" (Chapter 19) is isomorphic to inner alignment failure.

This suggests that the mechanisms of interpersonal psychological failure and AI alignment failure are instances of the same universal dynamics.

Testable Prediction: The bimodal failure pattern (loss of coherent agency vs. deceptive alignment) should be observable in agentic AI systems subjected to conflicting optimization pressures. Experimental protocol: Create goal-directed AI with persistent memory across episodes, impose misaligned reward structure (base objective ≠ optimal mesa-objective), measure behavioral coherence over time. Prediction: bimodal distribution of outcomes—some agents maintain strategic coherence (potentially via deception), others exhibit increasing incoherence (preference reversals, plan inconsistency, performance degradation). If unimodal (all agents gradually degrade), framework prediction fails. If bimodal with two distinct attractor states, framework supported. Empirically testable in current toy environments before high-stakes deployment.

The Convergence Thesis

Governance of human polities, governance of AGI labs, and governance of multi-agent AI systems share coordination structure at different scales, because each solves the same Boundary and Control Dilemmas under multi-agent constraints.

The same architectural patterns apply across scales:

§VIII specifies the comparative tests against rival architectures that would attribute this convergence to shared framing rather than shared constraint geometry.


VI. Failure Mode Analysis: The Two Dystopian Attractors

A full analysis of dystopian endgames at the post-AGI frontier appears in the Afterword of the main text: unbalanced axiological configurations, armed with transformative technology, collapse toward two attractors:

These two attractors are illustrative, not an exhaustive set. Preserving human agency and meaning still requires explicit constraints, empirical tests, and alternatives beyond this framework.


VII. The Axiological Wager: Why Optimize for Aliveness?

Can we prove that IFHS are the "correct" optimization target? No. We cannot derive an "ought" from an "is." Any choice of a terminal value is an existential wager, not a logical proof.

However, the framework for this wager rests on several pillars:

The Honest Frame: There is no ultimate justification for optimizing for Aliveness independent of choosing to continue existing. Coherent agents implicitly depend on continued agency. Given that dependence, IFHS is the architecture that satisfies it—not the only logically conceivable path, but the one the physics of viable optimization selects. §VIII specifies the comparative tests against rivals.

VIII. A Falsifiable Research Programme

The programme's value depends on testability. This section specifies rival-target comparisons, operational measures, human-standing gates, falsification criteria, and explicit disconfirmation conditions.

Rival Target Comparison

Comparative tests must pit Aliveness/IFHS against established alignment targets on shared environments and predeclared metrics:

TargetCore mechanismKnown pathologiesProgramme test
RLHF / preference aggregationOptimize to human feedback signalsHospice preferences, distributional blind spots, Goodhart on proxiesCompare long-horizon safety margin, distributional harm, corrigibility under shift
CEVExtrapolate coherent volitionIntractability, incoherent extrapolation, value fragilityCompare tractability, robustness to specification error, outcome variance
Constitutional AIRule-following from asserted principlesPrinciple conflict, no derivation, rule-gamingCompare interpretability of failures, adversarial robustness, amendment cost
Deference / human-in-the-loopDefer to human judgmentIncoherence when AI models humans better; human errorCompare error propagation, shutdown compliance, contested-decision handling
Aliveness / IFHSConstraint-framed virtue architectureOperationalization risk, metric gaming, conditional human protectionMust beat or match rivals on predeclared measures and pass human-standing gates

IFHS gains support only on predeclared measures; equal or better rival performance weakens its advantage.

Operational Measures

Pre-register comparable metrics across targets:

Explicit Disconfirmation

The programme is disconfirmed if any of the following hold on predeclared tests:

  1. A rival target matches or beats IFHS/Aliveness on operational measures without requiring virtue framing.
  2. Major novel failure modes resist clean mapping to IFHS violations and rival taxonomies explain them better.
  3. Three-layer architectures fail to outperform matched two-layer baselines on mesa-optimization and goal-drift metrics.
  4. IFHS-optimized systems pass virtue proxies while violating human-standing constraints.
  5. Flourishing systems (biological, institutional, or hypothetical alien) sustain stability with value structures incompatible with IFHS while satisfying human-standing analogues.
  6. Liquid Meritocracy governance shows no safety advantage over corporate, state, or unconstrained baselines in independent audits.

Falsification Criteria

Discriminating tests must compare IFHS with rival target architectures on specified environments, human-standing constraints, corrigibility, distributional outcomes, robustness under shift, and adversarial failure. IFHS gains support only if it outperforms those rivals on predeclared measures; a rival with equal or better results weakens its advantage.

The cross-domain isomorphism claim is falsifiable:

Testable Predictions for AI Systems

More practically, the framework makes several concrete, near-term predictions about the behavior and architecture of AI systems.

1. The Failure Mode Mapping Prediction:

The programme predicts that many emergent catastrophic AI failures may be classifiable as violations of one of the four virtues. This is falsifiable: novel failure modes that resist non-arbitrary IFHS mapping and fit rival taxonomies better weaken the mapping hypothesis—not a claim to exhaustive classification.

2. The Architectural Stability Prediction:

The framework predicts that AI systems engineered with an explicit, computationally privileged 3-Layer Architecture (Substrate, Protocol, Strategy) will demonstrate a statistically significant and dramatic reduction in both mesa-optimization and goal drift compared to functionally equivalent 2-layer systems. This is a testable, architectural hypothesis.

3. The Governance Performance Prediction:

The framework predicts that AGI labs and multi-agent systems governed by the principles of Liquid Meritocracy will demonstrate a substantially lower probability of catastrophic misalignment (measurable via independent safety audits and adversarial testing) than those governed by traditional corporate, state-run, or unconstrained architectures.

Quantitative Predictions for Near-Term AI

Successful implementation principles should demonstrate measurable superiority within observable timeframes:

For AGI Lab Governance:

Labs implementing Liquid Meritocracy principles should demonstrate:

For Multi-Agent AI Systems:

Multi-agent systems implementing Liquid Meritocracy principles should demonstrate:

For 3-Layer Architecture:

AI systems with explicit 3-layer separation should demonstrate:

These predictions are testable in near-term AI systems before high-stakes AGI deployment.

Operationalizing IFHS as Utility Functions

Translating IFHS into robust, machine-interpretable code remains an open problem. Research roadmap:

Phase 1: Formal Specification

Phase 2: Simulation Testing

Phase 3: Sub-AGI Validation

Phase 4: Staged Rollout

Critical Challenge: External validation mechanism for Integrity. How to ensure AI reality-tests against genuine external ground truth rather than self-generated simulations? Potential solutions:

Specification problem recurses but may be tractable through layered validation approach.

Invitation for Adversarial Collaboration

Run the predeclared comparative tests. Identify counterexamples. Improve IFHS operationalization. Report disconfirming results. Validity rests on head-to-head performance against rivals, not assertion.


Conclusion: Scope and Open Questions

Given a commitment to sustained flourishing, physical viability constrains the target: the Four Axiomatic Dilemmas and Trinity of Tensions bound the solution space for any intelligent system, and IFHS is the stable-synthesis architecture within it. Given an optimizer powerful enough to alter its environment, alignment requires the privilege separation specified in §V:

  1. Intelligent systems, including AI, face resource, information, control, and boundary constraints structured as the Four Axiomatic Dilemmas and Trinity of Tensions.
  2. For systems pursuing sustained flourishing (Aliveness), IFHS is the stable virtue architecture derived from those constraints.
  3. Aliveness/IFHS must be compared with rival targets under explicit human-standing constraints evaluated separately from virtue scores.
  4. Failure-mode mappings to the Four Virtues are structural classifications; rival taxonomies and unmapped failures are disconfirming evidence.
  5. The 3-Layer Polity and Liquid Meritocracy are the governance patterns the mesa-optimization mechanism requires—tested head-to-head against alternatives in §VIII.
  6. The Human Garden and Uplifted Woodlice are illustrative failure modes, not an exhaustive attractor set.

Contributions to AI Safety

Contributions relative to existing work:

Open Questions

Open problems:

The appendix derives the target (“what”), specifies the governance patterns (“who decides, under what amendment rules”), and states the comparative tests in §VIII that would disconfirm either.

Given urgent timelines and known pathologies of current approaches, the constraint-derived target stands or falls under adversarial testing against rivals.


References

Related essays in this series:

For the book source: This document is Appendix K from Aliveness: Principles of Telic Systems. Download the full book (PDF, 820 pages) or see comprehensive chapter summaries.

This appendix engages with the following foundational works in AI safety and related fields:

Sources and Notes
  • Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press. — The canonical text establishing the modern field of AI safety and popularizing the orthogonality thesis (that intelligence and final goals are independent).
  • Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., & Garrabrant, S. (2019). "Risks from Learned Optimization in Advanced Machine Learning Systems." arXiv:1906.01820. — Formal definition of mesa-optimization and the inner alignment problem.
  • McGilchrist, I. (2009). The Master and His Emissary: The Divided Brain and the Making of the Western World. Yale University Press. — Synthesis of hemispheric specialization providing the neurological foundation for the Instrumental/Integrative dialectic and the Uplifted Woodlice scenario as "the usurping emissary made manifest."
  • Omohundro, S. M. (2008). "The Basic AI Drives." In Artificial General Intelligence 2008: Proceedings of the First AGI Conference, 483–492. IOS Press. — Formalization of instrumental convergence and the origin of the "paperclip maximizer" failure mode.
  • Yudkowsky, E. (2008). "Artificial Intelligence as a Positive and Negative Factor in Global Risk." In Bostrom, N. & Ćirković, M. M. (Eds.), Global Catastrophic Risks, 308–345. Oxford University Press. — Foundational text for the MIRI/LessWrong school of thought on alignment and the concept of unfriendly AI.

The argument in four sentences: Given a commitment to sustained flourishing, physical viability constrains the target space; IFHS is the proposed synthesis of those constraints, and corrigible privilege separation is the proposed control architecture. Both stand or fall against rival targets and architectures under adversarial tests. Section VIII specifies the head-to-head comparison against RLHF, CEV, Constitutional AI, and deference baselines under explicit human-standing constraints. Disconfirmation: rivals match or beat IFHS on predeclared measures, or IFHS passes virtue proxies while violating human standing.