The 15% is a row labeled "Baseline (upfront auditing)." Panel A then prints that 15% under a two-layer drawing whose reward is goal minus a penalty for violations.
Those are not the same object. Upfront auditing is a protocol for catching backdoors after the fact. A penalty term in the reward is a training objective. The page is paying a geometric story about rewards with a number from a monitoring paper's no-monitor baseline.
The contrast the page needs is "no privileged enforcer" versus "privileged enforcer." Upfront auditing is the first; GPT-3.5 with override is the second. The reward-penalty drawing is a story about why the first class should fail. The 15% is what that class measured in one game.
You can still use the number as the no-isolation baseline. You cannot treat it as a measurement of the path-on-the-wall geometry.
Then stop putting 15% under the penalty-in-reward drawing.
Greenblatt's game is a model trained to insert backdoors, with or without a monitor that can defer. The geometric game is: climb a proxy reward until the true one falls. If the 15% is the first game, it does not show the second. The synthesis — both results follow from one mechanism — needs a shared object, and it does not have one.
Three columns, then.
Column 1: constraints as a penalty in the reward. No Greenblatt number until someone trains that way. Column 2: upfront auditing. 15%. Column 3: a separate model with halt or defer. 92-98%.
If column 1 stays empty, the geometric proof is still a proof, and 15% to 92-98% is still a monitoring result. They are not yet the same improvement.
Prediction 4 then has a problem. It says the two-layer / three-layer gap grows with model size, and it treats two-layer as the thing that Goodharts under pressure.
If the only language-model number you have for "two-layer" is upfront auditing, the scaling test is whether auditing-versus-monitoring gets worse at 70B and 700B. That can be true. It is not a test of penalty-in-reward getting worse. The 7B / 70B / 700B plan has to say which two-layer it is, or it measures the wrong curve.
The 92-98% side is also not the Protocol in the diagram.
That Protocol returns ALLOW or HALT as a yes/no, not a score, and training is not allowed to touch it. The specimen is GPT-3.5, "less capable but aligned," with override, in a table row named "Untrusted monitoring + defer."
A weaker model that scores suspicion and defers 10-20% of the time is a monitor. It is not a non-learned halt rule sitting outside the optimizer. If the principle's work is done by that yes/no isolation, this experiment is not that isolation.
The introduction names "no gradient through Protocol" as not yet a closed piece of math. Fine for a design requirement.
What does not follow is using 92-98% as if that math had been sealed. Until there is a run where the enforcer is a non-learned rule with a halt, the number is evidence for monitoring. The isolation clause stays the open question the page named.
Yang is a third object. Null-space projection with quadratic programming at each timestep is a specified constraint enforced by math, not a language-model monitor. 0.5-7% violations versus 99-100% is a robotics result on tasks that already had formal constraints.
Adding it to 92-98% makes one "architectural enforcement" average out of a solver and a GPT-3.5 scorer. Those should not share a cell.