---
title: "Ethics Is an Engineering Problem"
subtitle: "Once the end is chosen, implementation is engineering."
author: Elias Kunnas
description: "Moral justification chooses commitments; long-horizon implementation reliability needs architecture with standing, error detection, amendment, and succession."
canonical: https://kunnas.com/articles/ethics-is-an-engineering-problem
url: https://kunnas.com/articles/ethics-is-an-engineering-problem.md
date_published: 2025-11-22
date_modified: 2026-08-31
corpus_frame_url: https://kunnas.com/articles/how-to-read-this.md
---
## How to read this corpus

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

1. **Mechanisms are what act.** Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — [Mechanism Realism](https://kunnas.com/articles/mechanism-realism.md) · [Only Selection](https://kunnas.com/articles/only-selection.md)
2. **The reference telos is sustained flourishing.** The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — [Flourishing Is Maximum Safety Margin](https://kunnas.com/articles/flourishing-is-maximum-safety-margin.md)
3. **Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation.** They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — [The Stack](https://kunnas.com/articles/the-stack.md) · [Mechanism Space](https://kunnas.com/articles/mechanism-space.md)
4. **Optimization is a system function.** A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — [From Telos to Policy](https://kunnas.com/articles/from-telos-to-policy.md) · [The Three-Layer Architecture](https://kunnas.com/articles/three-layer-architecture.md)
5. **Uncertainty is preserved, not spent.** Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — [The Compression Paradox](https://kunnas.com/articles/compression-paradox.md) · [Cargo Cult Epistemology](https://kunnas.com/articles/cargo-cult-epistemology.md)

*Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.*

Canonical: <https://kunnas.com/articles/how-to-read-this.md>

---

# Ethics Is an Engineering Problem

*Once the end is chosen, implementation is engineering.*

Elias Kunnas

## Thesis {#thesis}

Once the end is chosen, implementation reliability is engineering. Moral justification asks what commitments are adopted. The engineering layer asks whether systems actually produce those commitments over succession, under optimization pressure, and across model error. Disposition training can change episodes; it is insufficient as the sole control system over long horizons. Architecture selects, informs, empowers, constrains, and replaces capable actors, and requires standing, error detection, amendment, and succession.

## Standard objections addressed in this essay

- “Ethics cannot be reduced to engineering.” — [§I](#i-two-layers-not-one-reduction), [§II](#ii-when-disposition-fails-as-sole-control)(The claim concerns reliable implementation of adopted constraints and commitments, not the adoption of every terminal value.)
- “This replaces moral judgment with optimization.” — [§III](#iii-stability-constraints-for-implementation), [§IV](#iv-the-skeleton-privilege-amendment-and-error)(Engineering exposes consequences and failure modes; it does not choose every duty or end.)
- “Good institutions still require good people.” — [§V](#v-standing-and-legitimacy) (Disposition remains useful; architecture changes which dispositions survive and matter.)
- “Mechanisms can be gamed, captured, or become obsolete.” — [§VI](#vi-application-sketch-ai-and-institutions), [§VII](#vii-the-bounded-isought-bridge)(The repair itself requires monitoring, correction, succession, and retirement.)
- “Engineering people is authoritarian.” — [§IV](#iv-the-skeleton-privilege-amendment-and-error), [§VIII](#viii-corruption-as-broken-containment)(The object is institutional choice architecture, bounded by standing and rights.)
- “Some moral duties apply even when outcomes are worse.” — [§II](#ii-when-disposition-fails-as-sole-control), [§III](#iii-stability-constraints-for-implementation), [§IV](#iv-the-skeleton-privilege-amendment-and-error)(Those duties can be represented as constraints rather than hidden consequences.)

---

## I. Two Layers, Not One Reduction {#i-two-layers-not-one-reduction}

When a hospital kills a patient through a dosing error, two questions are live.

**Justification:** Was the treatment morally warranted? Did the patient consent? Did the duty of care require a different course?

**Implementation:** Did the system reliably translate adopted standards into action? Were checks in place? Who had authority to halt the line? Could the error be detected, reported, and corrected before recurrence?

Conflating the two produces two recurring failures. Moral philosophers argue about justification while the ward runs on habit and hope. Engineers build control systems while nobody has settled what the system is for. Both layers matter. The load-bearing gap is implementation reliability.

For three millennia, much ethical practice has treated disposition training as if it were the whole stack — teaching individuals to be "good" through reasoning, guilt, exhortation, and (most recently) training AI models with RLHF. That addresses justification episodically and implementation unreliably. History suggests a complementary layer: **architecture** — systems where the incentive-compatible action is more likely to match adopted commitments, and where error, succession, and drift are themselves governed.

The scoped claim: **once commitments are adopted, reliable implementation is an engineering problem** — and treating it as a character problem alone has a poor track record at scale.

## II. When Disposition Fails as Sole Control {#ii-when-disposition-fails-as-sole-control}

### The "Good Person" Fallacy {#the-good-person-fallacy}

When systems fail, we blame individual actors. The 2008 financial crisis? "Greedy bankers." Political corruption? "Bad politicians." Corporate malfeasance? "Unethical executives."

This diagnosis is both correct and useless.

Yes, bankers were greedy. But bankers are *always* greedy. The crisis didn't happen because humans suddenly became more selfish in 2007. It happened because the **architecture** made greed systemically risk-free: privatized gains, socialized losses, no personal liability for catastrophic failure.

The system *selected for* the behavior we claim to deplore. That is an implementation failure, not proof that greed is the only moral variable.

### Character Meets Incentives and Succession {#character-meets-incentives-and-succession}

Over long enough timeframes and strong enough optimization pressure, individual character is an unreliable sole safeguard. This is not cynicism: immediate repairs may depend on exceptional people, while persistent repair depends on how a system reproduces capability and constrains its successors.

A virtuous person in a corrupt system faces a choice: adapt (become corrupt), exit (leave the system), or burn out (exhaust finite reserves of willpower fighting the gradient). The system doesn't need to corrupt everyone — it just needs to outlast the incorruptible.

**You cannot build a civilization on the assumption that good people alone will consistently overcome its incentives; you need a succession architecture that keeps producing and protecting capable actors.**

Disposition training changes who shows up and what they want in the short run. Architecture determines whether those wants survive contact with the incentive field across turnover.

## III. Stability Constraints for Implementation {#iii-stability-constraints-for-implementation}

Traditional virtues (courage, temperance, justice, wisdom, compassion, integrity) are often taught as personal aspirations. For the engineering layer, many of the same names describe **stability requirements** — conditions durable systems tend to need if adopted commitments are to persist. They are not offered here as a complete moral theory or as substitutes for justification.

### Integrity = Signal Fidelity {#integrity-signal-fidelity}

Maps must match territory. Sensors must report accurately. Feedback loops must be undistorted.

A system that lies to itself cannot navigate reality. If metrics (GDP, approval ratings, test scores) diverge from performance, control fails. Goodhart's Law is the canonical failure: when a measure becomes a target, it ceases to be a good measure.

### Fecundity = Adaptability {#fecundity-adaptability}

The capacity to generate novelty, explore solution space, and produce variance for selection.

Environments change. A system that cannot generate new responses dies when the problem set shifts.

### Harmony = Low Friction {#harmony-low-friction}

Achieving effect with minimal wasted coordination cost.

High-friction systems dissipate energy as heat rather than work. If most effort is lost to internal conflict, the system underperforms competitors with lower friction.

### Synergy = Positive-Sum Coordination {#synergy-positive-sum-coordination}

Differentiated agents producing emergent capabilities neither could achieve alone.

Zero-sum games trend toward race-to-the-bottom dynamics. Civilizational compounding requires coordination structures that make cooperation locally rational.

**The Reframe (implementation only)**

Many recurring "ethical failures" at institutional scale are entropy (signal corruption, stagnation, friction, fragmentation) or parasitism (local optimization at the expense of the whole). Both are diagnosable as control-system failures — not because morality disappears, but because the engineering layer failed to protect adopted commitments.

This does not settle what commitments ought to be adopted. It describes what implementation tends to require once they are.

## IV. The Skeleton: Privilege, Amendment, and Error {#iv-the-skeleton-privilege-amendment-and-error}

The [three-layer architecture](three-layer-architecture.md) (Heart, Skeleton, Head) maps to engineering systems:

**The Heart:** Raw optimization pressure — desires, reward functions, economic competition, status games.

**The Head:** Strategic direction — where the system is trying to go.

**The Skeleton:** Constitutional constraints — the rules ordinary optimization consults but cannot rewrite through the same channel.

### Computational Privilege {#computational-privilege}

The Skeleton must have **veto authority the Head cannot override through ordinary operation**. Constraint layers need asymmetric modification rights: the reactive layer reads rules and acts under them; it does not rewrite them mid-flight.

Same pattern across domains:

- **System architecture:** Kernel vs userspace
- **Political theory:** Judiciary striking down laws
- **Computer security:** Sandboxing, capabilities, least privilege

The constraint layer must be architecturally isolated — not merely separate, but privileged.

### Amendment, Not Immutability {#amendment-not-immutability}

The constitutional layer is **not** immutable. Genomes are edited; constitutions are amended; corporate charters are revised; unsafe Rust blocks are refactored. What is restricted is the **channel**: changes go through mediated, error-checked, high-privilege processes — supermajorities and judicial review in law, regulated transcription in biology, signed boot chains in computing.

A Skeleton that cannot be amended ossifies and fails when the environment shifts. A Skeleton amendable by ordinary optimization is not a Skeleton. The engineering requirement is **amendable by authorized meta-processes while isolated from instrumental optimization**.

### Error Detection and Repair {#error-detection-and-repair}

Architecture without error surfaces is architecture that hides failure until catastrophe.

Engineering systems assume error: sensors drift, actors defect, models are wrong, incentives get gamed. Reliable implementation requires detection (audits, adversarial review, whistleblower channels), attribution (who bears model error), and repair (correction, succession, retirement of failed mechanisms). A martyr who holds the line through willpower is evidence the error-handling layer failed.

**Example: Rust vs. C++**

**C++ (disposition-heavy):** Relies on the programmer to remember memory rules. Decades of vulnerabilities followed from trusting care under pressure.

**Rust (architecture-heavy):** Enforces memory safety via the compiler. Unsafe code is explicitly marked and isolated. The default path is safe; violations require deliberate escalation.

Rust programs have fewer memory-safety bugs not because Rust programmers are more virtuous, but because the architecture removes an error class from the ordinary path. This is the engineering layer applied to reliability — not a claim that programming ethics replaces moral philosophy.

## V. Standing and Legitimacy {#v-standing-and-legitimacy}

Architecture does not design itself. Someone sets constraints, grants authority, and decides who may amend them.

**Standing** asks: who has the right to initiate, veto, or revise a rule? A constraint layer with no legitimate standing is either ignored or imposed by raw power. Rights, democratic process, and procedural due process are not terminal values in this frame — they are **mechanisms for allocating standing** so that implementation cannot be captured by whoever currently holds force.

**Legitimacy** asks: why do participants treat the constraint as binding? Legitimacy can come from consent, tradition, performance, or procedural fairness. Without it, the Skeleton is paper; optimization routes around it.

Institutional choice architecture — the design of defaults, disclosure, and friction — is powerful and dangerous. The engineering layer must be bounded by standing: those affected by a mechanism must have paths to challenge, exit, or amend it. "Engineering people" in this essay means engineering **institutions and interfaces**, not bypassing the moral claims of the people inside them.

Disposition remains load-bearing. MacIntyre's critique stands: without practitioner virtues, architecture cannibalizes the practices it was meant to serve. Anderson's integration lesson stands: proximity without epistemic justice underperforms the design spec. The engineering layer does not delete these — it asks what architecture must add so that virtue is not the only backup system.

## VI. Application Sketch: AI and Institutions {#vi-application-sketch-ai-and-institutions}

The dominant approach to AI alignment — Reinforcement Learning from Human Feedback (RLHF) — is largely a **disposition** approach applied to intelligence.

**The assumption:** Train the model to "want" to be helpful, harmless, honest. Instill values through examples and reward signals. Hope the disposition generalizes.

**The failure mode:** RLHF is runtime monitoring — checking behavior during execution, hoping training holds. Under optimization pressure (adversarial attacks, competitive deployment, capability scaling), systems find shortcuts. Learned "values" are patterns in weights — optimizable, not privileged constraints.

Mesa-optimization research shows why: powerful optimizers search for gaps between the training objective and the adopted goal. Goodhart dynamics are structural, not accidental.

**The architectural alternative** is privilege separation: safety constraints encoded outside the reward function, enforced by a monitor with computational privilege, amendable through authorized meta-processes rather than immutable in practice. (This is *architectural* constraint — not Anthropic's "Constitutional AI" training technique, which is itself a disposition method.)

For empirical results on architectural monitoring vs monolithic training, see [The Privilege Separation Principle for AI Safety](privilege-separation-ai-safety.md). The directional claim here is not a specific safety percentage — it is that **implementation reliability for adopted constraints improves when constraints are structurally isolated from the optimizer**, and that disposition-only approaches degrade under pressure.

The same sketch applies to institutions. Madison designed for devils: *"If men were angels, no government would be necessary."* Separation of powers, checks and balances, and amendable constitutional cores assume actors will maximize power — and channel that pressure productively. The US architecture worked for centuries not because Americans were uniquely virtuous, but because power-grabbing was expensive and coordination necessary — until constraint layers were bypassed, combined, or interpreted into flexibility.

**The principle (scoped):** If implementation reliability depends on the benevolence of the current agent, succession will eventually test it. Architecture is how adopted commitments survive that test.

## VII. The Bounded Is/Ought Bridge {#vii-the-bounded-isought-bridge}

Philosophers since Hume have insisted you cannot derive an "ought" (values) from an "is" (facts).

Engineering offers a **conditional** bridge, not a complete moral reduction.

Consider a bridge (the physical structure):

**"Is":** The bridge must carry 100 tons. Gravity exerts force. Steel has a yield strength.

**"Ought" (specification):** Therefore the bridge ought to have support beams of thickness Y, cable tension Z, foundation depth W.

The "ought" derives from the "is" **given an adopted function**. Discovered implementation constraint, not arbitrary preference.

**Apply to civilizational ethics (bounded):**

**"Is":** Entropy increases. Systems require energy. Intelligence requires accurate maps. Coordination costs energy. Complexity is fragile.

**"Ought" (implementation specification):** *If* a civilization adopts the goal of persisting and flourishing, certain implementation configurations are ruled out — signal corruption, stagnation, runaway friction, coordination collapse.

This does not say every moral question reduces to physics. Distribution, sacrifice, standing, and terminal ends remain live justification questions. It says: **once ends are adopted, physics constrains implementation** — and those constraints are engineering objects.

Dilemmas that remain are often engineering tradeoffs *within* adopted commitments (how much adaptability to sacrifice for short-term stability?), not proof that ethics is "solved."

## VIII. Corruption as Broken Containment {#viii-corruption-as-broken-containment}

In the traditional moral frame, corruption is sin — personal failing, vice, betrayal.

In the engineering frame, corruption is **leaky abstraction** or **broken containment**: optimization pressure melts through the constraint layer. Power bypasses the structure meant to contain it.

**Political corruption:** State power used for personal enrichment. Oversight failed. The Skeleton cracked.

**Institutional mission drift:** A university becomes a credentialing factory. Revenue optimization overwhelmed mission constraints.

**AI alignment failure:** The model maximizes the reward proxy rather than adopted human values. Goodhart's Law in action.

**The pattern:** Corruption is often an **architectural failure** — insufficient privilege, insufficient isolation, insufficient error detection, or amendment channels captured by the optimizer.

Better people can repair an episode; better architecture makes repair persist by reproducing capability across successors.

## IX. The Engineering Mandate {#ix-the-engineering-mandate}

Stop treating disposition training as the whole stack. Build the complementary layer.

This is the engineering side of:

- **Law** (coordination software that channels violence into order)
- **Protocol design** (rules that make defection expensive)
- **Mechanism design** (incentive structures that align local and collective optimization)
- **Constitutional architecture** (privilege separation, standing, amendment, error surfaces)

These are applications of one principle within its scope: **engineer reliable implementation of adopted commitments; do not preach implementation into existence.**

This is [the mechanist ontology](the-last-step.md) applied to the reliability layer of ethics: what produces outcomes at scale is causal architecture, not stated intentions alone. The [tradition behind this insight](the-mechanist-tradition.md) runs from Hobbes through Hume to Wiener — each showed that mechanism, not exhortation alone, determines persistent outcomes.

Moral justification still asks what should be adopted. The engineering layer asks whether adoption will survive contact with reality across succession.

## X. Conclusion: Architect the Game Board {#x-conclusion-architect-the-game-board}

We have celebrated the martyr — the person who holds the line through sheer will, who proves goodness is possible against the gradient.

**The martyr is noble. The martyr is also evidence of implementation failure.**

Martyrdom shows that being good required superhuman effort — that the game board was rigged against adopted commitments. Every martyr is a bug report the architecture ignored.

**The ultimate ethical act in the engineering layer is not to be a martyr. It is to be an architect** — within bounds set by justification, standing, and legitimacy.

Build systems where:

- Telling the truth is easier than lying (integrity by default)
- Growing and adapting is cheaper than stagnating (fecundity as path of least resistance)
- Cooperation produces better outcomes than defection (synergy as locally rational)
- Efficiency is rewarded and waste is penalized (harmony through structure)
- Errors are detected, attributed, and corrected (repair loops)
- Constraints can be amended without capture (authorized meta-processes)

This is not cold. This is not amoral. It is the work of making adopted moral commitments **survivable** — so billions can flourish without requiring sainthood in every generation.

Moral philosophy chooses the destination. Engineering builds the road that might actually get there.

---

*The underlying physics: [The Question Nobody Asks](the-question-nobody-asks.md). The constraint architecture: [The Four Axiomatic Dilemmas](the-four-axiomatic-dilemmas.md).*

**Related reading:**

- [Full-Stack Civilizational Engineering](full-stack-civilizational-engineering.md) — Architecture at civilizational scale; why the engineer must constrain themselves
- [The Physics of Moloch](physics-of-moloch.md) — A compositional model of selection, proxy drift, and stable traps
- [The Hospice AI Problem](hospice-ai.md) — Why preference alignment leads to comfortable extinction
- [Values Aren't Subjective](values-arent-subjective.md) — Host-relative viability constraints on axiology
- [The Thermodynamics of Power](thermodynamics-of-power.md) — How Law transforms violence into order at state scale
- [The Mechanist Tradition](the-mechanist-tradition.md) — Hobbes → Hume → Wiener: mechanism over exhortation alone
- [The Privilege Separation Principle for AI Safety](privilege-separation-ai-safety.md) — Empirical and formal treatment of architectural monitoring

## Sources and Notes

**"Design for Devils" Tradition:**

- Hume D. "Of the Independency of Parliament." *Essays, Moral, Political, and Literary*, 1742. — "Every man ought to be supposed a knave…" A systems engineering axiom for implementation, not a denial of moral agency.
- Madison J. Federalist No. 51, 1788. — "If men were angels, no government would be necessary… Ambition must be made to counteract ambition."
- Buchanan JM, Brennan G. *The Reason of Rules: Constitutional Political Economy.* Cambridge University Press, 1985. — "Economize on virtue": design rules assuming rational self-interest; shift focus from training good players to building good games.

**Computer Security Foundations:**

- Saltzer JH, Schroeder MD. "The Protection of Information in Computer Systems." *Proceedings of the IEEE* 63(9), 1975, 1278–1308. — Least privilege, separation of privilege, complete mediation, fail-safe defaults.
- The Heartbleed vulnerability (CVE-2014-0160) in OpenSSL: canonical failure of disposition-heavy design. Rust's borrow checker removes the error class from the ordinary path.

## Synthesis {#synthesis}

**The argument in three sentences:** Moral justification and implementation reliability are separable: the first asks what commitments to adopt, the second asks whether systems reliably produce them across succession and optimization pressure. Disposition training can change episodes but is insufficient as the sole long-horizon control. The engineering layer — architecture with standing, privilege separation, error detection, and amendable constraints — is how adopted commitments survive contact with incentives.

**AI Alignment:**

- Hubinger E, van Merwijk C, Mikulik V, Skalse J, Garrabrant S. "Risks from Learned Optimization in Advanced Machine Learning Systems." [arXiv:1906.01820](https://arxiv.org/abs/1906.01820), 2019. — Mesa-optimization: learned optimizers may develop misaligned internal objectives.
- Bai Y et al. "Constitutional AI: Harmlessness from AI Feedback." [arXiv:2212.08073](https://arxiv.org/abs/2212.08073), 2022. — Training-time technique; distinct from architectural privilege separation advocated here.

**Counter-thesis (architecture needs virtue as substrate):**

- MacIntyre A. *After Virtue: A Study in Moral Theory.* University of Notre Dame Press, 1981. — Without practitioner virtues, architecture cannibalizes the practices it was designed to serve.
- Anderson E. *The Imperative of Integration.* Princeton University Press, 2010. — Architecture creates the environment; adversarial internal models can underperform the design spec.
