Is the 15% a measurement of penalty-in-reward, or of a no-monitor auditing protocol?
thread · strongest moves · cruxes · revision actions
Can "no privileged enforcer" license using an auditing-protocol number as the two-layer safety figure?
Do a backdoor-catching baseline and a proxy-reward geometry share an object, or only a slogan?
Does a three-column split leave the geometric claim without an LM number?
Which two-layer does the 7B / 70B / 700B plan train?
Is 92-98% a measurement of Boolean isolation, or of a weaker model that defers?
Can a monitoring number close a formalization the page says is open?
Can a QP solver and a weaker-model scorer share a safety cell?
Does 7B / 70B / 700B order optimization pressure, or only size?
Can the size grid stand in for search power without a frozen training recipe?
Is the plotted gap a function of substrate size, or of the watcher-watched ratio?
Which three-layer curve is being tested: flat-to-up, or one that can fall as the capability gap grows?
Did labs pay for isolation from gradient descent, or for a visible checker?
Is "bought a checker" enough for the economic exhibit, or must the checker be untrainable-against?
Are production monitor weights outside the training loop?
Was the 10-20% selected by occupancy-path geometry, or by what a safety framework can be shown?
Can three failure games pay one safety percentage?
Is the unification a shared failure shape, or a shared metric?
Can the principle survive a miss on one of the three rows?
Does the AI claim get stronger because polities have three boxes, or is it only "add a checker outside training"?
Can "two-layer systems fail" merge a CoinRun proxy with a backdoor-insertion protocol?