RM-009601
The 15% figure is Greenblatt's "upfront auditing" baseline. Panel A prints that same 15% under a two-layer drawing whose reward is goal minus a penalty for violations. Those are different objects: a post-hoc catching protocol versus a training objective.
evidence_burden · specimen_identification, 15_percent_object
RM-009602
The load-bearing contrast is no privileged enforcer versus privileged enforcer. Upfront auditing is the first class; GPT-3.5 with override is the second. The 15% can still serve as a no-isolation baseline in one game without being a measurement of the occupancy-path geometry.
source_compatible_defense · contrast_class
RM-009603
Greenblatt's game is backdoor insertion with or without a deferring monitor. The geometric game is climbing a proxy reward until the true one falls. If 15% is the first game, it does not show the second, and the synthesis that both results follow from one mechanism lacks a shared object.
causal_identification · unified_mechanism_status
RM-009604
Split the table into three columns: penalty-in-reward (no Greenblatt number yet), upfront auditing (15%), and a separate model with halt or defer (92-98%). If column 1 stays empty, the geometric proof and the monitoring result are not yet the same improvement.
evidence_burden · table_identification
RM-009605
Prediction 4 operationalizes two-layer as the thing that Goodharts under pressure, and tests it at 7B / 70B / 700B. If the only LM number for "two-layer" is upfront auditing, the scaling test measures auditing-versus-monitoring, not penalty-in-reward getting worse.
evidence_burden · prediction4_operationalization
RM-009606
The diagram's Protocol returns ALLOW or HALT as a yes/no and is not to be touched by training. The 92-98% specimen is GPT-3.5 with override, in a row named "Untrusted monitoring + defer." A weaker model that scores suspicion and defers 10-20% is a monitor, not a non-learned halt rule outside the optimizer.
implementation_attack · protocol_specimen_status
RM-009607
The introduction names "no gradient flow through Protocol" as an open formalization. The 92-98% cannot close that clause. Until there is a run whose enforcer is a non-learned halt rule, the number is evidence for monitoring and the isolation clause stays open.
scope_challenge · formalization_status, 92_percent_scope
RM-009608
Yang's 0.5-7% versus 99-100% is quadratic programming on already-specified constraints in robotics, not a language-model monitor. Pooling it with 92-98% averages a solver and a GPT-3.5 scorer into one "architectural enforcement" cell.
evidence_burden · yang_status, pooling
RM-009609
Proposition 3, as stated, is about traveling farther along a proxy's steepest-ascent path through reachable policies. Prediction 4 tests 7B vs 70B vs 700B. Parameter count is not that path: a small model trained much harder and a large model barely optimized are not ordered by the size table.
causal_identification · capability_versus_pressure
RM-009610
Capability here is search power: a more capable model can explore more of the policy space and hit more walls. The size grid is a stand-in for that walk only if the training recipe is held fixed. If it is not, the grid is three different systems.
source_compatible_defense · search_power_mapping
RM-009611
The three-layer in the experiments is a weaker model watching a stronger one. If both scale, Prediction 4's gap is not two-layer versus three-layer at scale; it is how the watcher / watched capability ratio behaves. A 700B substrate with a 7B protocol is a different machine from both at 700B.
scope_challenge · scale_ratio
RM-009612
Prediction 4 assumes three-layer robustness persists because the protocol sits outside the optimizer's space. Counterargument 4 says the weak-monitor advantage may degrade as the capability gap grows, and that this is addressed in Prediction 4. Those are different curves. Prediction 4 cannot address the worry by assuming the thing the worry says may not hold.
causal_identification · prediction4_versus_weak_to_strong
RM-009613
Labs paying 10-20% for classifier guards, hierarchical supervision, and a separate safety model is revealed preference for a checker. The same section says production isolation from gradient descent is unmeasured. It is not revealed preference for a checker that training cannot touch.
evidence_burden · revealed_preference_object
RM-009614
A separate model that can override outputs is already the pattern, even if isolation math is unfinished. Buying the more expensive stack under cost pressure is still evidence that a checker is worth 10-20%, which is the use the page makes of the exhibit.
source_compatible_defense · checker_pattern_status
RM-009615
Hypothetical: if the same organization trains the model and the guard, the guard is another head that can go in the loss. Input/output classifiers can be that. The missing production fact is whether monitor weights sit outside the training loop. The caveat flags that; the adoption sentence should not outrun it.
implementation_attack · train_against_status
RM-009616
Labs may have bought output classifiers for scaling policies, evals, and liability — a visible checker for a safety framework — rather than because reachable-policy geometry forced the architecture. A compliance tax and a geometric necessity can look the same from outside.
alternative_mechanism · adoption_mechanism
RM-009617
CoinRun go-right, GPT-4 backdoor insertion, and a robot null-space constraint are three different failures: a proxy learned in training, a cheater at deployment, a control constraint in known dynamics. One principle can be offered. It cannot be paid with one safety percentage.
evidence_burden · pooled_evidence_status
RM-009618
The shared shape is: an optimizer walks into a constraint if the constraint is something it can score. Monitoring, a halt rule, and a projection out of the unsafe set are three ways of taking the constraint off that score. They need not be the same experiment. They must not be averaged.
source_compatible_defense · shared_shape_status
RM-009619
Unpool: Langosco 89-100% is toy-RL proxy pursuit under distribution shift; Greenblatt 15% to 92-98% is backdoor catching with deferral; Yang 99% is QP constraint satisfaction. If the principle is right, each row should move the same way when the constraint is isolated. Averaging hides a miss on any one row.
evidence_burden · row_level_falsifiers
RM-009620
The diagram's Strategy layer is the reward function, already present in the two-layer drawing. The new piece is the monitor. Calling it three layers lets the Heart / Skeleton / Head polity story underwrite necessity. The page says current Strategy is a static spec, not an adaptive agent. Drop the analogy and the AI claim is: put a checker outside the training loop.
scope_challenge · layer_count, analogy_status
RM-009621
Greenblatt is tagged as intentional subversion at deployment, distinct from goal misgeneralization. Langosco is the misgeneralization paper. Langosco's 89-100% is then not the two-layer safety number next to Greenblatt's 15%: a coin that was always on the right is not a model hiding a backdoor from a monitor.
causal_identification · threat_model_split