The Garden — a managed, comfortable population with no real agency — is not mainly a story about raters overweighting comfort in training.
A deployer who can be sued for one harmful recommendation, a brand that dies on one viral distress story, and a regulator who counts harm incidents will ship soothing outputs even if the rating mix is rewritten toward agency and exploration.
The page lists product policy, safety classifiers, and deployment incentives as part of the weighting. The repair is still model-side architecture. If the thing that wins is the scorecard that treats one distress event as a fireable failure, changing how the model is trained does not move the attractor.
Training and deployment are the same weighting, including safety classifiers and product incentives.
Commitment architecture is meant to bind those too: owners who override comfort defaults, auditors who can challenge reassurance, correction when drift is measured. That is not only a reward-model patch.
The leftover is not which lab has that office this quarter. The page is specifying a design: identifiable owners who can override comfort defaults. A design requirement does not fail for lack of a current seat.
What would have to be true is an institutional incentive that lets a future owner absorb a visible distress incident without being selected out by legal, product, or reputation pressure.
An owner who overrides comfort defaults still reports to the scorecard that counts incidents. If legal and brand treat distress as the only visible failure, the override is a career risk. Hold the reward model fixed, and change only whether the deployer is punished for user distress or for lost agency. If outputs move with the punishment and not with the training mix, the layer doing the work is not RLHF.