The Corrigible State

How grounded contradiction can change public action without making the challenger sovereign

Elias Kunnas

Corpus frame

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

  1. Mechanisms are what act. Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — Mechanism Realism · Only Selection
  2. The reference telos is sustained flourishing. The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — Flourishing Is Maximum Safety Margin
  3. Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation. They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — The Stack · Mechanism Space
  4. Optimization is a system function. A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — Telic Systems · The Three-Layer Architecture
  5. Uncertainty is preserved, not spent. Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — The Compression Paradox · Cargo Cult Epistemology

Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.

Where each commitment is derived

AI corrigibility asks whether a capable optimizer will cooperate with corrective intervention rather than defend its current objective. A state has no single legitimate operator and no uncontested objective, so governance corrigibility cannot be obedience to an expert, voter, court, or critic.

Relative to a specified public mechanism and error class, it is the capacity for a grounded challenge to enter through a legitimate route, survive as a stable and contestable case, reach an answerable owner, and—when upheld—activate lawful change whose implementation and effects remain traceable. A legitimate authority may reject or override the challenge, but the correction architecture must contain at least one lawful actuator capable of changing the mechanism when correction is chosen. A state may be transparent, accountable, adaptive, and electorally replaceable while failing that chain.

Standard objections addressed in this essay
  • “Democracy is already corrigible through elections.” — §IV (elections can replace rulers while the failed mechanism survives)
  • “This is just accountability.” — §IV (accountability can assign judgment without producing repair)
  • “This is just adaptive or experimentalist governance.” — §VII (those are topologies; the trace tests whether correction closes)
  • “Who decides what counts as an error?” — §VI (mechanism, legal, and political contradictions have different authorities)
  • “This gives experts a veto over elected government.” — §VI (analysis can require a response without owning political closure)
  • “Institutions need finality and stability.” — §§III, VIII (corrigibility includes proportionate friction and conditional closure)
  • “Government cannot formally process every criticism.” — §§III, V (standing, materiality, and error class bound the path)
  • “Who corrects the corrector?” — §VIII (replacement and retirement are part of corrigibility)

I. Corrigibility Without an Operator

AI safety has a word for a system that does not make its present objective immune from correction: corrigibility.

The simplest technical formulation was deliberately narrow. A capable agent may have instrumental reasons to resist shutdown or modification. A corrigible agent instead cooperates with what its creators regard as corrective intervention. The Off-Switch Game sharpened one route: a robot that is uncertain about its objective may preserve a human operator’s ability to switch it off because the human’s action carries information about what should be done.

The abstraction is powerful. Its governance transfer is not direct.

A state is not one optimizer with one reward function. The public is not one operator. Elections, courts, ministries, auditors, civil servants, municipalities, affected people, experts, and political movements possess different kinds of standing and authority. Their objectives overlap, conflict, change, and remain partly unspecified. No one actor is entitled to substitute their judgment for the whole merely by claiming to have found an error.

The real problem is harder:

What would corrigibility mean when the operator, the objective, the evaluator, and the authority to intervene are distributed and contested?

Rationalist governance thought reached several neighboring objects. Futarchy is a genuine decision mechanism, not merely a prediction device: elected representatives define a welfare measure, conditional markets estimate which policy raises it, and the market result determines policy. Inadequate Equilibria analyzes why institutions can remain predictably suboptimal. Meditations on Moloch maps the destructive equilibria produced when locally rational actors cannot jointly close a bad option.

Those are serious pieces of governance engineering. They do not by themselves answer the lifecycle questions:

A prediction market can make a policy conditional on a forecast and still leave later implementation failure without an owner. An audit can discover divergence after the fact and still carry no effective response path. An election can replace a government while preserving the same mechanism, data system, professional incentives, and administrative interpretation.

A system does not become corrigible merely by knowing that it is wrong.

The Governance Alignment Problem asks whether public mechanisms optimize for what political institutions say they want. Corrigibility begins when the answer is no—or when the objective, model, or implementation itself is challenged. Alignment concerns correspondence with the current target. Corrigibility concerns the authorized path by which the target, mechanism, or judgment can be revised.

An aligned system can be incorrigible. It may pursue its installed objective faithfully while excluding every challenge to that objective. A corrigible system can be temporarily misaligned. Its defining property is that the discrepancy can survive into an authoritative correction process.

The governance version has to be a property of the path.


II. The State Has No Single Corrector

The authority problem is the design problem.

Different institutions are competent to correct different things.

A court may determine that a public action violates a constitutional or statutory rule. It is not thereby the best body to estimate an employment elasticity or redesign a hospital funding formula. An audit office may reconstruct expenditure and implementation failure. It does not thereby inherit authority to choose the distributional objective. A ministry may possess operational knowledge and lawful implementation powers. It may also have a structural interest in defending the programme it designed. A parliamentary majority may possess political authority. It may lack the time, information, and independent capacity needed to test a mechanism claim.

Affected people can know where a system fails at the edge. They may not know the whole mechanism. Experts may know the mechanism. They may not represent the people bearing its costs. Opposition parties may expose a defect. Their incentive may be to maximize embarrassment rather than produce repair. Incumbents may know what can actually be changed. Their incentive may be to minimize the admission of failure.

No role should be silently promoted into the universal corrector.

The governance equivalent of a corrigible agent is an arrangement in which no incumbent controls every gate between a grounded contradiction and the possible correction of authoritative state.

That requires separation among at least five roles:

Role Function Authority it does not acquire by performing the role
Challenger introduces evidence or a rival model final judgment or implementation power
Adjudicator tests whether the claimed contradiction survives scrutiny authority over every political end
Decision owner accepts, rejects, or overrides within a legitimate mandate permission to erase the challenge or its reasons
Implementation owner changes the operational mechanism authority to certify its own success alone
Verifier tests whether the correction changed the relevant outcome automatic power to choose the next political settlement

The roles can be distributed across institutions. One institution can sometimes hold more than one role. The arrangement becomes fragile when the same actor can define admissibility, select the evidence, decide the merits, control implementation, certify success, and prevent reopening.

This is a governance form of the Dominant-Player Constraint: no actor should control the rules, evidence, forum, timing, closure condition, and enforcement substrate of a contest that affects others.

Plural authority requires typed authority.

The challenger may be entitled to standing without being entitled to win. The adjudicator may be able to issue a finding without being able to enact a budget. Parliament may retain final political closure while owing a public response to a finding. The implementation owner may choose among lawful repairs while remaining answerable for the movement test. The verifier may reopen the question without prescribing the final replacement.

The central institutional move is to construct a path in which the relevant correction can advance without any one participant becoming sovereign over the whole.


III. Corrigibility Is Relative

“Is this state corrigible?” is usually too broad to answer.

A public system may be corrigible for one class of error and closed to another. Constitutional review can make legislation corrigible with respect to constitutional incompatibility. Administrative appeal can make an individual decision corrigible with respect to legal or procedural error. Elections can make officeholders corrigible with respect to political support. Fiscal rules can create correction duties for specified budget deviations.

None of those automatically makes a policy mechanism corrigible when its incentives produce the opposite of its stated purpose.

The object must be bound before the property can be tested:

mechanism or authoritative object

+declared purpose or binding constraint

+error class

+affected population

+time horizon

+institution with lawful correction authority

A grounded challenge is a claim specific enough to be tested against the relevant comparator, accompanied by evidence or a model that can be contested, and admitted through a threshold proportionate to the decision at stake. The merits can remain uncertain. The term means that the challenge has become evaluable, not that it has already been certified true.

A governance arrangement is corrigible relative to this bound object when a grounded challenge can clear each transition:

Institutions also need commitments, settled expectations, legal finality, administrative capacity, and protection against harassment. A programme that restarts from zero whenever anyone submits a counterclaim is ungovernable. Corrigibility therefore includes correction friction: standing rules, materiality thresholds, evidentiary burdens, time windows, deference where competence warrants it, and closure conditions scaled to reversal cost.

The friction itself must remain challengeable. Otherwise “finality” becomes the mechanism by which one generation binds later evidence forever.

The relevant contrast is change through an owned, legitimate, evidence-sensitive path versus change only through crisis, scandal, personal heroism, or incumbent permission. A system may change constantly while remaining incorrigible: it can adapt around criticism, modify its rhetoric, relocate costs, replace personnel, and preserve the mechanism that produced the failure.

A system may also reject a challenge and remain corrigible. The decisive question is whether the distinct claim was preserved, tested under the relevant authority structure, and issued a state that can later be audited.

A corrigible system can reject a challenge. It cannot treat a challenge that never became a case as if it had been disposed of.


IV. The Properties That Stop Early

Several desirable governance properties resemble corrigibility. None is sufficient alone.

Property What it provides Where it can stop before corrigibility
Transparency public visibility into rules, reasons, data, or conduct everyone can see the defect while nobody owns a response
Auditability an object can be inspected against a standard the audit can end as a report without an actuator
Accountability an actor must explain or bear judgment blame or explanation can occur without mechanism repair
Responsiveness the institution reacts to demands or signals response may follow salience, power, or volume rather than error
Learning beliefs, models, or procedures update authoritative state and implementation can remain unchanged
Adaptability the system can alter itself under pressure adaptation may preserve the incumbent objective or relocate the failure
Reversibility a decision can technically or legally be undone no owner or trigger may activate reversal
Contestability affected parties can challenge a decision challenge can lack answerability, implementation, or later verification
Resilience the system persists through disruption the institution may become highly resilient at preserving its defect
Elections personnel and coalitions can be replaced the same rule, model, and administrative mechanism can survive every replacement
Corrigibility grounded contradiction can alter authoritative state through a legitimate complete path requires composition of ingress, adjudication, authority, execution, verification, and meta-correction

This is why the word accountability is too large and too small at once.

It is large because it covers answerability, blame, sanctions, reporting, legal review, electoral replacement, and professional responsibility. It is small because an accountable actor may give a complete answer while the failed mechanism remains in force.

Feedback Authority asks what cost an institution bears for non-response. Its answer can range from decorative acknowledgment to legal or fiscal consequence. Governance corrigibility includes that result but does not end there. A response duty can produce a reasoned no, a procedural remand, or a public override without producing implementation of an accepted correction.

Implementation Ledger owns the next transition: an accepted response becomes operational only when a stable record assigns the decision object, owner, resources, deadline, execution state, verifier, and reopening rule.

Corrective Closure Ownership owns the later transition: who is required and able to notice when reality disagrees, reopen the decision, repair it, and account for what changed.

There Is No Exception Handler owns an earlier one: what happens when no ordinary process can own the challenge at all. Its handler carries the no-match state until accepted handoff or reasoned termination. Governance corrigibility is the complete path across those states.

These properties have to compose around a bound mechanism and error class.


V. The Corrigibility Trace

The Corrigibility Trace is an audit for that composition.

Begin by naming the mechanism, the purpose or constraint against which it is judged, the error class, the affected population, and the horizon. Then follow a grounded challenge through nine transitions.

Stage Required question Failure signature
1. Ingress Who can introduce evidence or a rival model without the challenged operator controlling every route? criticism exists only at the operator’s discretion
2. Stable case What preserves the mechanism, claim, evidence, uncertainty, and requested change across translation and referral? the distinctive object dissolves into correspondence, summary, or aggregate reporting
3. Standing Who decides admissibility and materiality, under what public criteria, appeal, or sampled review? gatekeeping becomes an unreviewable agenda monopoly
4. Adjudication What comparator is used, who tests the challenge, and who can contest the test? the operator supplies both model and verdict, or no discriminating test exists
5. Answerability Which named actor must issue which disposition by what date? a finding can remain indefinitely “noted”
6. Actuation or override Which lawful power can change authoritative state, and who may explicitly proceed despite the finding? the process can recommend but cannot alter anything; override remains silent
7. Implementation Where are owner, resources, deadline, execution state, verifier, and blocked states recorded? acceptance becomes a press release, working group, or promise
8. Verification and re-entry What observation tests the result, and what preauthorized path reopens it? completion is self-certified or later failure starts a new political campaign from zero
9. Meta-correction Who can challenge, revise, replace, or abolish the correction process itself? the corrector becomes the one mechanism exempt from correction

A break shows that correction at that point depends on something outside the claimed architecture:

A failed trace: Robodebt

Australia’s Robodebt scheme is useful because it contained many institutions normally cited as evidence that correction machinery already exists.

Welfare recipients could seek internal review and review by the Administrative Appeals Tribunal. Complaints reached the Commonwealth Ombudsman, which opened an own-motion investigation in January 2017. Its April 2017 report found problems in usability, transparency, service delivery, communication, testing, and risk management; the responsible departments agreed to its recommendations. A 2019 follow-up found significant implementation progress on those recommendations.

The central legal defect nevertheless survived. The scheme used income averaging to raise alleged welfare debts. Justice Emilios Kyrou later summarized that between 2016 and 2022 the tribunal made 431 first-tier decisions holding that debts calculated through the averaging technique could not be recovered, while the department effectively ignored a large number of those decisions. Kyrou also records that some first-tier cases accepted averaging; the 431 are the refusals, not the whole caseload. First-tier decisions were generally unpublished. Evidence before the later Royal Commission also showed that relevant legal material had been withheld from the Ombudsman, hampering its ability to determine whether the scheme was lawful.

The point is not that every institution failed in the same way. The Ombudsman identified and followed up real administrative improvements. Individual tribunal applicants sometimes won. Courts, advocates, journalists, and political actors eventually produced decisive change. A Royal Commission later reconstructed the system and issued fifty-seven recommendations. The new Administrative Review Tribunal subsequently gave its President a function to identify and notify government actors of systemic issues found in the tribunal’s caseload.

The corrigibility failure was compositional. Individual legal contradictions did not reliably aggregate into a stable systemic case. The institution operating the scheme could continue while affected people repeatedly reopened the same question one case at a time. Oversight produced accepted recommendations on parts of administration without forcing the foundational legal issue through an authoritative correction path. The mechanism changed only after years of distributed contest and escalating external force.

Robodebt does not prove that the Corrigibility Trace would have prevented the scheme. It shows why possession of appeals, a tribunal, an ombudsman, internal review, judicial review, parliamentary scrutiny, and eventual inquiry does not itself establish a timely correction loop. The question is how the contradictions compose.


VI. Who May Correct What?

“Reality contradicted the policy” does not identify who may lawfully decide what follows.

Three error classes must remain separate.

Error class Comparator Appropriate authority Legitimate dispositions
Mechanism contradiction the policy’s own stated purpose, causal claim, forecast, or implementation specification independent analysis plus the politically responsible decision owner repair, test, reject the finding, accept referral, or proceed through explicit override
Legal or constitutional contradiction a binding higher-order legal rule court, tribunal, or another legally authorized review body invalidate, interpret, remand, suspend, or uphold
Political value or distributional conflict competing legitimate ends, rights, burdens, and risk tolerances constitutionally authorized political institutions choose, bargain, compensate, constrain, defer, or retain the status quo with reasons

The first class is the natural domain of mechanism analysis. A government says that a financing rule will improve efficiency, a tax change will increase employment, a procurement reform will create competition, or a service redesign will shorten waiting times. A grounded challenge can test whether the incentives, resources, actor responses, and implementation path support that claim.

A finding does not automatically determine the policy. The stated objective may conflict with another objective. The mechanism may be uncertain but still worth testing. A rights constraint may rule out the most efficient design. The government may accept a cost for a distributional or symbolic reason.

The correction architecture should force the difference into view.

causal claim survivesproceed on the claim

causal claim failsrepair, test, or proceed through an explicit political override

This is the attraction of repair or explain. The evaluator does not acquire the final vote. The political authority does not retain the ability to present a defeated mechanism claim as if it were uncontested.

But reason-giving alone is not full corrigibility. A government can produce formulaic overrides forever. A court can declare a breach while implementation stalls. An agency can accept a recommendation and leave it unfunded. The path needs an actuator, implementation trace, and later verification appropriate to the error class.

The second class belongs to legal authority. A mechanism analyst may identify a possible legal issue but does not become a court. The third belongs to politics within constitutional constraints. An expert may clarify consequences but cannot derive the legitimate distribution of sacrifice from technical competence.

This is how grounded objections receive landing rights without their authors receiving sovereignty:

Governance corrigibility is rule through a process in which claims of correctness can alter the state only through typed, contestable authority.


VII. Partial Corrigibility Already Exists

The individual parts are not inventions.

Democratic experimentalism is the strongest established near-neighbor. Its classic architecture decentralizes problem-solving, lets local actors use contextual knowledge, requires them to report performance, compares alternative approaches, and periodically revises goals and procedures. Global experimentalist governance applies a similar recursive structure across jurisdictions: open-ended problem definition, participatory multilevel implementation, peer review, locally generated knowledge, and periodic revision.

That is a real corrigible-governance topology where the goals, participating units, reporting relationships, and coordinating process can be specified.

The Corrigibility Trace adds a different cut. Any proposed topology must also clear these transitions:

A 2026 preprint on digital public infrastructure uses corrigibility in an even closer architectural sense. It defines a system as corrigible when affected participants possess structural access to detect, contest, and override systemic error, and proposes five jointly necessary conditions: exit, inspectable code, audit, governance, and forkability. It correctly distinguishes corrigibility from transparency, accountability, openness, and auditability in isolation.

That work occupies a substantial part of the territory. The present object differs in scope and authority structure. Public mechanisms are not all digital infrastructures; many cannot provide universal exit or forkability without abandoning the public function. Governance corrigibility must also distinguish legal, causal, and political error; preserve an explicit override by legitimate political authority; and trace accepted correction through implementation and outcome verification. The two frameworks should be treated as neighboring architectures, not as one originating the entire idea.

Administrative systems also implement partial traces.

The National Transportation Safety Board (NTSB) attaches safety recommendations to named recipients, requests responses, preserves public status classifications, and can keep recommendations open while action remains inadequate. The U.S. Government Accountability Office (GAO) maintains recommendation records and follows implementation. Courts provide authoritative correction for specified legal errors. Ombuds institutions create protected complaint and investigation routes. There Is No Exception Handler describes aviation reporting and unified incident command as bounded ways to preserve anomalies or assemble temporary cross-jurisdiction ownership.

These systems show that corrigibility need not be concentrated in one super-agency. A distributed architecture can qualify when its transitions are mandatory enough, accepted, and traceable.

They also show how far the general case remains from routine.

The Organisation for Economic Co-operation and Development (OECD) Government at a Glance 2025 reports a 2024 iREG average of 1.34 out of 4 for ex-post evaluation of primary laws. Only nine of thirty-eight countries with data, plus the European Union, reached two or above. Only seven required periodic evaluation of all primary laws, with another four covering all major laws. Most systems remained ad hoc or partial.

Finland already has evaluative capacity. Regulatory impact assessment is required for primary laws. The Finnish Council of Regulatory Impact Analysis reviews selected assessments and may review ex-post evaluations. Finland adopted common ex-post-evaluation principles in 2023. But ex-post evaluation is not mandatory, the Council’s opinions are advisory, and the architecture does not by itself guarantee an answerable owner, actuator, implementation trace, and reopening path for every significant mechanism. The OECD’s 2026 Finland review therefore recommended a comprehensive ex-post-evaluation system with stronger oversight and systematic incorporation of feedback.

The empirical claim should remain bounded:

Modern states contain many corrigibility components and several real correction loops. They do not thereby demonstrate a general Corrigibility Trace for public mechanisms across error classes and institutional boundaries.

The first burden is to show the existing functional equivalent—or show where the trace breaks.


VIII. The Corrector Can Harden

A correction system introduces a new mechanism. It inherits every problem it is meant to solve.

The repair is a correction architecture that exposes and limits its own powers.

At minimum the correction architecture itself needs these controls:

There is no infinite tower. At some point a constitutionally authorized body makes a political or legal decision under uncertainty. Corrigibility requires that the unresolved state, reasons, dissent, and later outcome remain available to the next authorized correction path. It does not promise a final viewpoint outside all institutions.

When Ownership Is the Wrong Repair supplies the relevant warning: a correct diagnosis can compile into a captured owner, a decorative report, an overloaded process, or binding mass that destroys the capacity it was meant to protect.

A corrigibility institution that cannot be challenged, evaluated, replaced, or retired fails its own trace.


IX. The Corrigible State

No single organizational form follows from the definition.

A legislature may build its own analytical office. Courts may own legal contradictions. Audit institutions may maintain implementation and outcome records. Ministries may run adaptive programmes under external verification. Citizens and affected groups may receive bounded standing. Several specialized bodies may connect through mandatory handoffs and shared case identity. One independent authority may own only the cross-boundary gaps.

The topology is open. The function is not satisfied by naming its components.

For a specified class of public mechanisms, a corrigible arrangement must let grounded contradiction enter, remain intact, reach legitimate judgment, activate a lawful change when it wins, survive implementation, meet later evidence, and expose the correction machinery to the same process.

The Fourth Branch is one institutional form topology for part of that function: independent mechanism analysis, public findings, and a political response path. It is not the definition of governance corrigibility and need not be its only implementation.

The status quo is not a neutral baseline. Existing rules already allocate standing, silence, delay, agenda access, proof burdens, and the cost of uncorrected failure. Incumbency supplies operating history, tacit capacity, and path dependence. It does not supply acquittal. Once a plausible break in the Corrigibility Trace has been shown, the comparison becomes bilateral: the challenger must specify the proposed repair and its risks; the incumbent must identify the current functional equivalent or defend leaving the break in place.

Receiving criticism, answering it, or changing is not yet corrigibility. The property appears only when contradiction can become authoritative correction without becoming the private sovereignty of whoever first identified it.

A corrigible state is one in which no incumbent controls every gate between error and repair.


Related:

Sources and Notes

AI corrigibility and operator-centered models

  • Nate Soares, Benja Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong, “Corrigibility” (MIRI technical report, 2014; AAAI 2015 workshop). The original formulation asks whether a capable system cooperates with interventions its creators regard as corrective despite incentives to resist shutdown or preference modification.
  • Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell, “The Off-Switch Game” (2016). The human–robot model shows how uncertainty about the objective can give the robot an incentive to preserve human intervention.
  • The transfer in this essay is deliberately non-identical. Public governance has plural principals, contested objectives, divided legal authority, and no human operator whose revealed preference can serve as the general correction oracle.

Rationalist governance neighbors

  • Robin Hanson, “Futarchy: Vote Values, But Bet Beliefs”. Futarchy is treated here as a real policy-selection actuator: conditional market estimates determine whether a proposal becomes law under a politically selected welfare measure. The residual question concerns implementation, later contradiction, and retirement rather than initial selection alone.
  • Eliezer Yudkowsky, Inadequate Equilibria (2017). The book supplies a general account of institutional inadequacy, including decision-makers who do not capture public benefits, asymmetric information, and bad equilibria.
  • Scott Alexander, “Meditations on Moloch” (2014). The essay supplies a powerful public model of destructive coordination equilibria. It is used here as diagnosis, not as a completed correction architecture.
  • Eliezer Yudkowsky, “Politics Is the Mind-Killer”. The conjecture that this norm discouraged detailed institutional engineering is not a load-bearing claim of the essay and is not asserted in the body.

Democratic experimentalism and recursive governance

A close contemporary use of public-system corrigibility

  • Anivar A. Aravind, “Corrigibility as a Structural Precondition for Digital Public Infrastructure: A Cybernetic Framework” (2026). The paper defines corrigibility as structural access by affected participants to detect, contest, and override systemic error through five conditions: EXIT, CODE, AUDIT, GOVERN, and FORK. It explicitly rejects transparency, accountability, openness, or auditability alone as sufficient. The present essay does not claim first use of corrigibility as a public-system property. Its different cut is plural-authority governance across non-digital mechanisms, typed error classes, political override, implementation, outcome verification, and correction of the corrector.

Robodebt as a failed composition

  • Royal Commission into the Robodebt Scheme, final report and recommendations (2023). The Commission reconstructed the scheme’s legality, design, administration, and institutional failures and issued fifty-seven recommendations.
  • Commonwealth Ombudsman, 2017 report announcement on Centrelink’s automated debt system. The investigation found deficiencies in usability, transparency, service delivery, communication, project planning, testing, and risk management; the departments agreed to the recommendations.
  • Commonwealth Ombudsman, 2019 implementation follow-up. The follow-up found significant progress while identifying remaining room for improvement.
  • Justice Emilios Kyrou, “Key Features of the New Administrative Review Tribunal” (2024). The paper records 431 first-tier tribunal decisions between 2016 and 2022 holding that averaged-income debts could not be recovered, the general non-publication of those decisions, and the department’s ability to ignore many of them.
  • Commonwealth Ombudsman, evidence concerning the Robodebt Royal Commission (2023). The Ombudsman states that relevant documents and legal advice were withheld during its investigation, hampering its ability to assess lawfulness.
  • Justice Emilios Kyrou, “Identifying and Notifying Systemic Issues” (2026). The new Administrative Review Tribunal architecture includes a presidential function to notify ministers, entities, and the Administrative Review Council of systemic issues identified in the tribunal’s caseload.

Tracked recommendations as partial architecture

Ex-post evaluation and the Finnish boundary case

  • OECD, Government at a Glance 2025, “Ex post evaluation”. The 2024 OECD average for ex-post evaluation of primary laws was 1.34 on a four-point scale; only nine of thirty-eight countries plus the EU reached two or above, and systematic requirements remained uncommon.
  • OECD, Regulatory Policy Outlook 2025: Finland. Finland requires regulatory impact assessment for primary laws, operates a Council of Regulatory Impact Analysis with advisory rather than sanctioning power, and had not made ex-post evaluation mandatory.
  • OECD, Foundations for Growth and Competitiveness 2026: Finland. The report recommends comprehensive ex-post evaluation, stronger oversight, systematic stakeholder feedback, and continuous improvement of evaluation practice.

Scope and contribution

The component fields are well populated. The contribution claimed here is a deployment-grade, topology-neutral audit for corrigibility without a privileged operator: a bound mechanism and error class, a nine-stage trace from ingress to meta-correction, and an authority split among challenge, adjudication, political or legal closure, implementation, and verification. The trace is proposed, not independently validated as a standardized governance instrument.