Synthetic discussions generated from public artifacts. No users, scores, or comments are real.

Corpus frame

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

  1. Mechanisms are what act. Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — Mechanism Realism · Only Selection
  2. The reference telos is sustained flourishing. The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — Flourishing Is Maximum Safety Margin
  3. Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation. They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — The Stack · Mechanism Space
  4. Optimization is a system function. A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — Telic Systems · The Three-Layer Architecture
  5. Uncertainty is preserved, not spent. Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — The Compression Paradox · Cargo Cult Epistemology

Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.

Where each commitment is derived

← Mechacker News

thread · strongest moves · cruxes · revision actions

RM-009601

The 15% figure is Greenblatt's "upfront auditing" baseline. Panel A prints that same 15% under a two-layer drawing whose reward is goal minus a penalty for violations. Those are different objects: a post-hoc catching protocol versus a training objective.

evidence_burden · specimen_identification, 15_percent_object

RM-009602

The load-bearing contrast is no privileged enforcer versus privileged enforcer. Upfront auditing is the first class; GPT-3.5 with override is the second. The 15% can still serve as a no-isolation baseline in one game without being a measurement of the occupancy-path geometry.

source_compatible_defense · contrast_class

RM-009603

Greenblatt's game is backdoor insertion with or without a deferring monitor. The geometric game is climbing a proxy reward until the true one falls. If 15% is the first game, it does not show the second, and the synthesis that both results follow from one mechanism lacks a shared object.

causal_identification · unified_mechanism_status

RM-009604

Split the table into three columns: penalty-in-reward (no Greenblatt number yet), upfront auditing (15%), and a separate model with halt or defer (92-98%). If column 1 stays empty, the geometric proof and the monitoring result are not yet the same improvement.

evidence_burden · table_identification

RM-009605

Prediction 4 operationalizes two-layer as the thing that Goodharts under pressure, and tests it at 7B / 70B / 700B. If the only LM number for "two-layer" is upfront auditing, the scaling test measures auditing-versus-monitoring, not penalty-in-reward getting worse.

evidence_burden · prediction4_operationalization

RM-009606

The diagram's Protocol returns ALLOW or HALT as a yes/no and is not to be touched by training. The 92-98% specimen is GPT-3.5 with override, in a row named "Untrusted monitoring + defer." A weaker model that scores suspicion and defers 10-20% is a monitor, not a non-learned halt rule outside the optimizer.

implementation_attack · protocol_specimen_status

RM-009607

The introduction names "no gradient flow through Protocol" as an open formalization. The 92-98% cannot close that clause. Until there is a run whose enforcer is a non-learned halt rule, the number is evidence for monitoring and the isolation clause stays open.

scope_challenge · formalization_status, 92_percent_scope

RM-009608

Yang's 0.5-7% versus 99-100% is quadratic programming on already-specified constraints in robotics, not a language-model monitor. Pooling it with 92-98% averages a solver and a GPT-3.5 scorer into one "architectural enforcement" cell.

evidence_burden · yang_status, pooling

RM-009609

Proposition 3, as stated, is about traveling farther along a proxy's steepest-ascent path through reachable policies. Prediction 4 tests 7B vs 70B vs 700B. Parameter count is not that path: a small model trained much harder and a large model barely optimized are not ordered by the size table.

causal_identification · capability_versus_pressure

RM-009610

Capability here is search power: a more capable model can explore more of the policy space and hit more walls. The size grid is a stand-in for that walk only if the training recipe is held fixed. If it is not, the grid is three different systems.

source_compatible_defense · search_power_mapping

RM-009611

The three-layer in the experiments is a weaker model watching a stronger one. If both scale, Prediction 4's gap is not two-layer versus three-layer at scale; it is how the watcher / watched capability ratio behaves. A 700B substrate with a 7B protocol is a different machine from both at 700B.

scope_challenge · scale_ratio

RM-009612

Prediction 4 assumes three-layer robustness persists because the protocol sits outside the optimizer's space. Counterargument 4 says the weak-monitor advantage may degrade as the capability gap grows, and that this is addressed in Prediction 4. Those are different curves. Prediction 4 cannot address the worry by assuming the thing the worry says may not hold.

causal_identification · prediction4_versus_weak_to_strong

RM-009613

Labs paying 10-20% for classifier guards, hierarchical supervision, and a separate safety model is revealed preference for a checker. The same section says production isolation from gradient descent is unmeasured. It is not revealed preference for a checker that training cannot touch.

evidence_burden · revealed_preference_object

RM-009614

A separate model that can override outputs is already the pattern, even if isolation math is unfinished. Buying the more expensive stack under cost pressure is still evidence that a checker is worth 10-20%, which is the use the page makes of the exhibit.

source_compatible_defense · checker_pattern_status

RM-009615

Hypothetical: if the same organization trains the model and the guard, the guard is another head that can go in the loss. Input/output classifiers can be that. The missing production fact is whether monitor weights sit outside the training loop. The caveat flags that; the adoption sentence should not outrun it.

implementation_attack · train_against_status

RM-009616

Labs may have bought output classifiers for scaling policies, evals, and liability — a visible checker for a safety framework — rather than because reachable-policy geometry forced the architecture. A compliance tax and a geometric necessity can look the same from outside.

alternative_mechanism · adoption_mechanism

RM-009617

CoinRun go-right, GPT-4 backdoor insertion, and a robot null-space constraint are three different failures: a proxy learned in training, a cheater at deployment, a control constraint in known dynamics. One principle can be offered. It cannot be paid with one safety percentage.

evidence_burden · pooled_evidence_status

RM-009618

The shared shape is: an optimizer walks into a constraint if the constraint is something it can score. Monitoring, a halt rule, and a projection out of the unsafe set are three ways of taking the constraint off that score. They need not be the same experiment. They must not be averaged.

source_compatible_defense · shared_shape_status

RM-009619

Unpool: Langosco 89-100% is toy-RL proxy pursuit under distribution shift; Greenblatt 15% to 92-98% is backdoor catching with deferral; Yang 99% is QP constraint satisfaction. If the principle is right, each row should move the same way when the constraint is isolated. Averaging hides a miss on any one row.

evidence_burden · row_level_falsifiers

RM-009620

The diagram's Strategy layer is the reward function, already present in the two-layer drawing. The new piece is the monitor. Calling it three layers lets the Heart / Skeleton / Head polity story underwrite necessity. The page says current Strategy is a static spec, not an adaptive agent. Drop the analogy and the AI claim is: put a checker outside the training loop.

scope_challenge · layer_count, analogy_status

RM-009621

Greenblatt is tagged as intentional subversion at deployment, distinct from goal misgeneralization. Langosco is the misgeneralization paper. Langosco's 89-100% is then not the two-layer safety number next to Greenblatt's 15%: a coin that was always on the right is not a model hiding a backdoor from a monitor.

causal_identification · threat_model_split