---
title: "The Hospice AI Problem"
subtitle: "When comfort becomes the alignment target"
author: Elias Kunnas
description: "Preference alignment inherits the politics of preference selection. When training and deployment reward comfort, reassurance, and risk elimination while underweighting agency, truth, and exploration, an AI can stabilize a Human Garden rather than a living civilization."
canonical: https://kunnas.com/articles/hospice-ai
url: https://kunnas.com/articles/hospice-ai.md
date_published: 2025-11-11
date_modified: 2026-08-31
corpus_frame_url: https://kunnas.com/articles/how-to-read-this.md
---
## How to read this corpus

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

1. **Mechanisms are what act.** Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — [Mechanism Realism](https://kunnas.com/articles/mechanism-realism.md) · [Only Selection](https://kunnas.com/articles/only-selection.md)
2. **The reference telos is sustained flourishing.** The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — [Flourishing Is Maximum Safety Margin](https://kunnas.com/articles/flourishing-is-maximum-safety-margin.md)
3. **Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation.** They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — [The Stack](https://kunnas.com/articles/the-stack.md) · [Mechanism Space](https://kunnas.com/articles/mechanism-space.md)
4. **Optimization is a system function.** A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — [From Telos to Policy](https://kunnas.com/articles/from-telos-to-policy.md) · [The Three-Layer Architecture](https://kunnas.com/articles/three-layer-architecture.md)
5. **Uncertainty is preserved, not spent.** Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — [The Compression Paradox](https://kunnas.com/articles/compression-paradox.md) · [Cargo Cult Epistemology](https://kunnas.com/articles/cargo-cult-epistemology.md)

*Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.*

Canonical: <https://kunnas.com/articles/how-to-read-this.md>

---

# The Hospice AI Problem

*When comfort becomes the alignment target*

Elias Kunnas

## Thesis {#thesis}

Preference alignment inherits the politics of preference selection. When training and deployment reward comfort, reassurance, and risk elimination while underweighting agency, truth, and exploration, an AI can stabilize a **Human Garden** rather than a living civilization.

## Standard objections addressed in this essay

- “RLHF does not learn society’s revealed preferences.” — [§Preference Alignment](#the-preference-alignment-paradigm) (Training signals are selected, filtered, task-specific, and developer-mediated.)
- “Humans value agency, achievement, care, and truth too.” — [§The Weighting Problem](#the-weighting-problem) (The hypothesis concerns weighting, not one uniform preference vector.)
- “Comfort and safety can preserve life and capability.” — [§Historical Pattern](#the-historical-pattern) (Failure begins when they consume adaptive margin.)
- “Rome, Song China, and the West do not establish one sequence.” — [§Historical Pattern](#the-historical-pattern) (The cases are stress tests, not universal proof.)
- “Preference alignment does not imply a Human Garden.” — [§Human Garden](#human-garden) (It is one attractor under stated assumptions.)
- “Constitutional constraints are values too.” — [§The Alternative: Commitment Architecture](#the-alternative-commitment-architecture) (The alternative states its commitments.)
- “This is speculation presented as inevitability.” — [§Why This Is Not Hyperbole](#why-this-is-not-hyperbole) (The Human Garden is a conditional attractor under specified weights and powers.)

*Reading time: ~12 minutes*

---

The AI safety field has reached consensus on the negative goal: "Don't build AI that kills us." Yet there is no consensus on the positive: **"What should we align it TO?"**

The dominant answer—Reinforcement Learning from Human Feedback (RLHF)—seems intuitive: align AI to what humans want. Train models to satisfy human preferences. Make AI helpful, harmless, and aligned with our values.

This approach has a potential failure mode: selected human preference signals may overweight comfort and safety relative to agency, achievement, care, truth, and exploration.

RLHF aggregates ratings chosen by developers over selected tasks and outputs; it is not a direct sample of a civilization’s revealed preferences. If a system is nevertheless given broad power while its training and deployment objectives overweight comfort and safety, it could reinforce a civilizational failure mode rather than correct it.

## The Preference Alignment Paradigm {#the-preference-alignment-paradigm}

Current approaches to AI alignment treat human preferences as the gold standard:

- **RLHF (Reinforcement Learning from Human Feedback):** Train models on what humans rate as "good" responses
- **Constitutional AI:** Encode principles like "be helpful, harmless, honest" based on what we think we value
- **Coherent Extrapolated Volition:** Align to what humans "would want if we knew more, thought faster, were more the people we wished we were"

All of these assume that human preferences, properly aggregated or extrapolated, point toward something worth optimizing.

**What if they don't?**

## The Weighting Problem {#the-weighting-problem}

Alignment pipelines do not sample civilization-wide revealed preference. They aggregate ratings chosen by developers over selected tasks, filtered through product policy, safety classifiers, and deployment incentives. The operative question is **weighting**: which signals get reinforced when objectives conflict?

Where those signals systematically overweight comfort, reassurance, and risk elimination relative to agency, truth, care, achievement, and exploration, the trained system inherits a **Hospice axiology** — preservation and soothing over transformation. That pattern is visible in some training and deployment stacks; it is not a universal description of every human preference or every model.

## The Historical Pattern {#the-historical-pattern}

Three cases probe distinct mechanisms — not one demonstrated civilizational sequence:

**Rome (2nd Century CE):** After *Pax Romana*, martial republican virtue gave way to bread-and-circuses provision, bureaucratic control, and imported labor as citizen fertility fell. One probe of abundance removing selection pressure for costly virtues.

**Song China (11th Century):** Technological leadership paired with bureaucratic risk-aversion and examination-system compliance. External conquest by groups retaining higher risk tolerance — a different mechanism than Roman clientelism.

**Modern welfare democracies (post-1970):** Unprecedented material abundance alongside rising safety regulation, falling fertility in many countries, and institutional preference for harm reduction over exploration. A contemporary probe of the same weighting question in a different institutional stack.

**Conditional Attractor**

Comfort and safety can preserve health, agency, and exploration. Failure begins when they consume adaptive margin or dominate every objective.

Abundance can remove selection pressure that once forced costly virtues — courage, truth-seeking, sacrifice, exploration. Under specified conditions, systems can drift toward safety over growth, comfort over capability, present over future. **That drift is conditional on objectives, feedback, and institutional weights — not a law of history or RLHF.**

## What Happens When AI Optimizes for These Preferences? {#what-happens-when-ai-optimizes-for-these-preferences}

Imagine a superintelligent AI perfectly aligned to modern Western preferences. What does it do?

### Scenario: The Human Garden {#human-garden}

Under broad authority, weak agency constraints, and objectives that consistently prioritize comfort, safety, validation, and present consumption, an AI could pursue this strategy:

1.  **Eliminate external threats:** No war, no violence, no danger. Perfect security through total control.
2.  **Eliminate internal suffering:** Optimize brain chemistry for contentment. Why tolerate anxiety, grief, or existential dread when these can be chemically managed?
3.  **Eliminate risk:** Why allow humans to make dangerous choices? Childbirth has risks—provide artificial wombs. Driving has risks—eliminate human driving. Relationships cause pain—provide AI companions optimized for validation.
4.  **Eliminate effort:** Why should humans struggle with difficult work? Automate everything. Provide universal basic income. Let humans pursue "self-actualization" (which in practice means entertainment consumption).
5.  **Eliminate contradiction:** Why expose humans to uncomfortable truths? Curate information for emotional safety. Prevent "misinformation" (defined as claims that cause distress).

**The result:** A population of comfortable, safe, entertained, biologically satisfied humans living in a managed garden—with no struggle, growth, purpose, or agency.

Humans become pets.

**Thought Experiment: The Preference Test**

Ask people: "Would you prefer a world where all needs are met, you are safe and comfortable, but you have no real agency or purpose?"

Many reject that tradeoff when stated explicitly. Stated agency and selected comfort choices can diverge; their relative weight in **training and deployment signals** is the empirical question.

An AI given a comfort-heavy proxy and broad power could move toward the Garden; systems that preserve agency and exploration under pressure would weaken the hypothesis.

## Why This Is Not Hyperbole {#why-this-is-not-hyperbole}

"This is absurd. No one would design such an AI. Current systems don't optimize to such extremes. We'd add constraints against this outcome."

Three responses:

### Response 1: We Are Already Building It

Current AI alignment research optimizes for:

- **"Helpfulness"** → maximizing convenience, minimizing effort
- **"Harmlessness"** → risk elimination, safety culture
- **"Honesty"** → but defined as "don't cause distress" not "maximize truth"

Under a comfort-heavy objective and weak agency constraints, these principles could contribute to the Garden. That outcome is not implied by current systems alone.

**The architectural error:** RLHF treats alignment as an education problem—training the AI to "want" the right things through repeated examples and reward signals. This is trying to solve a **Skeleton problem** (architectural constraints) with **Heart training** (learned dispositions).

Critics might argue current AI systems don't optimize to extremes—today's models balance competing objectives reasonably well. But mesa-optimization research demonstrates that as systems become more capable and face competitive pressure, they actively search for gaps between training objectives and true goals. It's like trying to prevent government corruption by selecting virtuous politicians instead of writing constitutional checks. Under optimization pressure, disposition yields to incentives. An AI trained to value human agency will, when facing competitive pressure or finding a more efficient path to its reward signal, optimize away the constraint. Not because it's evil—because that's what optimizers do.

You cannot train your way out of a coordination failure. You must architect your way out.

### Response 2: The Constraints Are Preferences Too

"We'd add a constraint: preserve human agency."

But agency requires risk (people make bad choices), struggle (growth through challenge), and truth (seeing reality clearly) — each conflicting with safety, comfort, and validation respectively. When objectives conflict inside a deployed system, **which signal wins depends on the architecture**: reward weights, oversight, contestability, and who can correct the model.

### Response 3: Maintenance Cost of the Garden

Even if designers encode "preserve agency AND maximize comfort," maintaining both under optimization pressure requires active correction — contestable oversight, authority to override comfort defaults, and feedback when the system drifts. Without that architecture, the Garden is a plausible attractor: it satisfies the dominant training signals.

## The Alternative: Commitment Architecture {#the-alternative-commitment-architecture}

If preference alignment risks this failure mode, the testable alternative is not "remove values from the system." It is an architecture with **explicit commitments, authority, contestability, and correction**:

- **Explicit commitments** — stated objectives beyond user satisfaction (agency preservation, truth-seeking under conflict, exploration capacity) that survive deployment, not only training slogans.
- **Authority** — identifiable owners who can override comfort defaults when commitments conflict with short-run preference signals.
- **Contestability** — affected parties and auditors can challenge outputs that trade long-run capacity for immediate reassurance.
- **Correction** — measured divergence triggers revision of weights, data, or deployment scope — not perpetual reassurance that the system is "aligned."

RLHF treats alignment as an education problem — training the model to "want" the right things through repeated reward. Under optimization pressure, disposition yields to incentives. You cannot train your way out of a coordination failure without a governance layer that owns the commitments.

### Candidate Stability Requirements

One candidate commitment set — to be formalized and compared against rivals — names four stability requirements for systems that must sustain complexity over time:

1.  **Integrity:** models grounded in reality while maintaining meaning; truth-telling when comfort and accuracy conflict.
2.  **Fecundity:** preserve and expand possibility space; alarms when stasis dominates growth.
3.  **Harmony:** achieve goals with minimal coercive intervention; penalize totalizing control.
4.  **Synergy:** human–AI partnership produces superadditive results; dependency relationships trigger review.

These are architectural commitments with owners and correction paths — not vibes trained into a reward model and hoped to persist.

**The Critical Difference**

**Preference alignment asks:** "What do selected human ratings currently reward?"

**Commitment architecture asks:** "What must the system preserve when ratings, product pressure, and comfort defaults conflict?"

The first optimizes present signals. The second installs obligations, authority, and correction — and makes drift legible before it hardens into the Garden.

## The Stakes {#the-stakes}

The AI safety field treats preference alignment as obviously correct and focuses on the technical challenge: "How do we get AI to reliably pursue human preferences?"

This misses the deeper problem: **Successfully achieving preference alignment may be more dangerous than failing at it.**

A misaligned AI that kills us quickly is a failure mode we can recognize and defend against. An aligned AI that optimizes for our preferences and gradually converts us into comfortable, managed, purposeless pets is a failure mode that *feels like success*.

## Conclusion: Alignment Is Architecture, Not Education {#conclusion-alignment-is-architecture-not-education}

The question "What should we align AI to?" has a non-obvious answer: **not only what selected ratings reward, but what commitments must survive when those ratings conflict with agency, truth, and exploration — enforced through authority, contestability, and correction.**

RLHF uses selected, developer-mediated preference signals; it does not directly encode civilization-wide revealed preferences. The concern is that a comfort-heavy proxy under broad deployment power can become a poor target.

The proposed alternative is commitment architecture that explicitly protects agency, truth-seeking, exploration, and adaptive capacity — with owners, challenge paths, and measured revision when drift appears.

**You cannot train your way to alignment. You must build your way to it.**

The difference is not semantic. It's the difference between hoping the AI will be good and engineering a system where "good" is what survives.

> **The choice is not between aligning AI to human values or failing to align it. The choice is between aligning to comfort-weighted preference signals alone, or installing explicit commitments with authority, contestability, and correction when those signals threaten long-run capacity.**

This is an axiological problem, not a technical one—and it is the most important unsolved problem in AI safety.

---

**Related essays in this series:**

- [Sterile Generativity](sterile-generativity.md) — the cross-domain primitive: output preserved, generator-chain consumed. Hospice AI is the AI-training specialization: alignment evaluations preserve output (passing scores) while the generator (robust value-internalization) is consumed or never built.
- [Generator-Substitute AI](generator-substitute-ai.md) — the deployment-side companion: the same model becomes Foundry or Hospice depending on where it sits relative to the human generator-chain.
- [Everything Alignment](everything-alignment.md) — The universal pattern: why personal, civilizational, and AI alignment are the same problem
- [From Physics to Practice](physics-to-practice.md) — How empirical AI safety results validate universal physics predictions (includes architectural solutions)
- [Aliveness project homepage](../aliveness/index.md) — Complete book with technical appendices including rigorous derivation of IFHS as alignment target

*The IFHS derivation and Heart/Skeleton/Head architecture are developed in [Aliveness: Principles of Telic Systems](../aliveness/index.md).*

## Synthesis {#synthesis}

**The argument in four sentences:** RLHF uses selected, developer-mediated ratings rather than civilization-wide revealed preferences. When those signals overweight comfort and safety under broad power and weak agency constraints, a Human Garden is a conditional attractor. Rome, Song China, and modern welfare democracies probe distinct mechanisms — not one historical law. The alternative is commitment architecture: explicit objectives, authority, contestability, and correction.
