The Game Theory of Cargo Cult Science

Integration is a public good. Authoritative fragments are private assets.

Elias Kunnas

Corpus frame

The corpus applies one lens to many domains: what mechanisms produce the outcome? It shares four methodological commitments and one explicit directional commitment. Each linked page argues for its part; the links are derivations and disputes, not evidence inherited by every page. The directional commitment does not by itself settle system boundary, distribution, sacrifice, or institutional authority.

  1. Mechanisms are what act. Incentive gradients, selection pressures, feedback loops, and capital stocks produce the distribution of outcomes. Intentions, labels, official categories, and stated values are evidence about mechanisms, or are themselves coordination mechanisms. They are not causal substitutes. — Mechanism Realism · Only Selection
  2. The reference telos is sustained flourishing. The broadest achievable adaptive safety margin over deep time — not the continuity of any incumbent state, coalition, institution, or doctrine. A mechanism's own stated goal can still serve as a local proof obligation — showing that its incentives defeat even the purpose it claims is a bounded finding — but meeting that goal establishes nothing about the margin. — Flourishing Is Maximum Safety Margin
  3. Law, rights, legitimacy, democracy, markets, and sovereignty are mechanisms under evaluation. They are constraints, carriers, or proxies inside the analysis. None is a terminal value or a boundary of what is real. Treating one as terminal ends the mechanism search before it starts. Evaluation carries current function, replacement cost, path dependence, uncertainty, capture risk, reversibility, and who bears model error into the ledger. — The Stack · Mechanism Space
  4. Optimization is a system function. A civilization has to build, exercise, and revise metamechanisms that search mechanism-space, discard dominated options, install, observe effects, and repair under uncertainty. Not running that loop leaves margin unrealized, and that is itself the failure. No single component — analyst, model, or institution — is presumed to contain a global optimum; the capacity is a property of the system. — From Telos to Policy · The Three-Layer Architecture
  5. Uncertainty is preserved, not spent. Partial orders, binding constraints, unknowns, and residuals stay explicit. An unmeasured effect is not a favorable default. — The Compression Paradox · Cargo Cult Epistemology

Each essay bears its own evidence. Links carry definitions, derivations, applications, and disputes; they do not transfer proof. Criticism is answered on its substance.

Where each commitment is derived

Integration is a public good. Authoritative fragments are private assets. A valid study of one slice can be spent as if the whole question had been settled. Cargo-cult science here is that substitution: public authority without maintained knowledge. The test is whether a study can be traced to a result that changes what is now believed, predicted, or decided, and whether it is being spent beyond what it established.

Standard objections addressed in this essay
  • “This is just publish or perish.” — §III (that explains distorted papers; the residual is composition even when papers are valid)
  • “This is just citation shopping or political think tanks.” — §IV (the stronger loop starts at question selection and accreditation; advocacy can be honest about its purpose)
  • “The literature itself is the maintained model.” — §V (an append-only record does not say what is current, rejected, or deprecated)
  • “Systematic reviews already do this.” — §VI (they are partial working cases; living maintenance remains a separate function)
  • “A canonical model would become a ministry of truth.” — §VII (state can be versioned, contested, and forked without being monopolized)
  • “This blames individual scientists.” — §III (local rationality is the mechanism; guilt is a separate question)
  • “Who maintains the maintainers?” — §VIII (the same test applies to the correction layer)

I. Every player did their job

A tax debate can return, year after year, to one binary while the technical literature already contains a richer object. One side says a higher rate will raise revenue. The other invokes the Laffer curve. An economist replies that the answer depends on behavioral response.

Researchers distinguish changes in labor supply from avoidance and income shifting, and from bargaining over compensation; they also study how the measured elasticity depends on the surrounding tax system.1 One paper estimates one margin under one reform. Another estimates another margin in another population. Each can be real work.

The researcher answers a tractable question. Reviewers assess the manuscript. A journal publishes it. The university records an output. The funder records a deliverable. An expert cites the relevant estimate. A politician cites whichever estimate bears on the argument at hand. The public discussion returns to the same binary.

No role has to fail its local checklist. Nothing in this chain requires anyone to produce the joined answer: the current response vector, its scope conditions, its unresolved branches, the claims that no longer survive, and the decisions that should now differ.

Suppose, to make the case harder, that every published estimate is methodologically sound and accurately describes its stated scope. The problem remains. A collection of true local claims does not maintain the relation among them, decide whether an apparent disagreement comes from different populations, bases, horizons, or channels, or tell the next public discussion which formulations have become indefensible.

The system can therefore be full of valid scientific transactions and empty of cumulative public state.

The obvious defense is that everything is working as designed. Researchers research; journals review; universities employ; funders fund; politicians choose. That defense identifies the game without establishing that the game performs the purpose from which science receives its authority.

II. Cargo cult at the level of composition

Richard Feynman called a practice cargo-cult science when it reproduced the visible forms of scientific investigation but omitted something necessary for the result: the runway looked right, yet no plane landed. His main repair was scientific integrity — reporting the facts that could defeat one’s preferred interpretation, publishing results regardless of direction, and making it possible for others to judge the work.2

That diagnosis applies inside a study. Cargo Cult Epistemology names a neighboring substitution at the level of speech and intellectual practice: epistemic forms used for tribal or instrumental work. The present failure is the clean-handed remainder one level higher. Participants may mean what they say, and each study may pass, while the composition still maintains no corrected state.

Within a bounded domain where research is offered as cumulative public knowledge or decision support, the plane is not the paper. The plane is a corrected operative model: what is now believed, with what scope and confidence, which alternatives remain live, and what predictions or decisions consequently change.

A system is piecewise scientific when its local operations — measurement, analysis, review, replication — can each satisfy real scientific standards. It is cargo-cult at the level of composition when those operations reproduce the authority and appearance of cumulative inquiry while no persistent state is required to register what the collection has learned.

The diagnostic is concrete. Before a study, what bounded model represented the question? Which claim, parameter, edge, or uncertainty did the study put at risk? After the study, who decided whether that object was changed, rejected, branched, or left unresolved? Which downstream claims and decisions inherited the result?

When these questions have no institutional answer, the study enters an archive. It may be correct, useful, and recoverable. It has not thereby become a maintained update.

This is a form of telic corruption. The scientific system borrows legitimacy, resources, and deference from one purpose — reliable cumulative correction toward truth — while its operative completion conditions can terminate at another: grant awarded, study conducted, paper accepted, citation counted, report delivered. The substitute objectives are not imaginary. They keep laboratories open and careers alive. The corruption lies in allowing their completion to stand in for the purpose that justified them.

The ritual may still produce salaries, prestige, trained specialists, and useful measurements. Cargo cult means the advertised causal chain is absent. The planes whose arrival legitimates the runway do not have to land.

III. The paper equilibrium

The composition failure is stable because the paper is a better local game object than the maintained model. The Reward Epidemic describes the general inversion in which a role becomes a distributable prize rather than a burden of function. Here the paper is both evidence and reward token: the countable object by which the research role is allocated, renewed, and ranked.

Player Locally rewarded completion Unowned remainder
Researcher Publishable study, citations, grant record, position Integrating rivals, maintaining scope, replication, deprecation
Journal Novel and citable papers, submissions, prestige A current state whose upkeep creates recurring cost and conflict
University or funder Countable outputs, placements, awards, programme deliverables Whether the domain model became more accurate or usable
Reviewer A result on one manuscript Reconciliation across the literature, usually with little credit
Policy user An admissible citation for a preferred action A model that constrains selective use and makes departures visible
Public Diffuse benefit from better knowledge Monitoring and rewarding the maintenance needed to obtain it

The payoff asymmetry is straightforward. A paper creates an attributable unit of credit. Maintaining the field’s current model creates a largely shared benefit and a concentrated cost. Everyone benefits if somebody does it; each participant has reasons to spend the next unit of effort on work that is easier to own.

For each actor, the captured benefit and credit from one unit of maintenance can remain below that actor’s cost, even when the total benefit to everyone exceeds the total cost. The first comparison governs individual action. The second exists only on the shared ledger.

Research on scientific incentives usually enters here by showing that the private return to publication can select bad methods. Models have shown how publication-based career competition can favor underpowered novelty, positive results, and high output, and how poor methods can proliferate without fraud or conscious misconduct.3 Publication bias can then make false claims accumulate toward apparent fact.4

Those are severe failures. They are not required for the paper equilibrium. Even perfecting every paper leaves four stabilizers in place.

No conspiracy is necessary. Each can sincerely want science to accumulate and still choose the action that secures their own project. Everyone can wish that somebody maintains the model and decline to become the somebody.

That is the clean game: the local products are privately credited; the integrated epistemic state is a public good. The equilibrium produces papers because papers close everyone’s local loop. Policy use adds a second game.

IV. The market for authoritative fragments

Authoritative fragments are private assets: a maintained model would constrain selective implication, expose omitted counterfactuals, and sometimes declare a familiar argument obsolete, while an append-only literature supplies a menu instead.

No fraud is required. A sponsor can select the question, comparator, population, horizon, outcome, and what counts as a cost. A study can accurately estimate jobs supported by an industry while never asking about displacement, incidence, or whether the disputed policy produces those jobs. The local answer can be true. Call that instrumental truth: a true bounded statement produced and circulated because it can be made to license a larger preferred conclusion.

A professorship, institute name, or expert appointment is also an asset that increases the policy force of claims. The credential does not prove the policy claim. It cheapens the crossing of legitimacy gates: a coalition wants X becomes research shows X. The Reward Epidemic owns the prize inversion; here the prize converts into standing. The Framing Machine names the containers that make the claim cheaper to trust before its content has been checked.

Player Locally valuable product What a maintained model would remove
Sponsor or coalition Accredited finding useful for an objective Implication beyond the study’s scope
Think tank or commissioned institute Funding, access, media demand Independence from comparison against the full state
Credentialing body or expert brand Prestige and policy access for its bearers Authority detached from a maintenance record
Journalist or politician Quotable expert support and blame insulation Ability to cite a fragment without declaring departure from the model

The public thinks it purchased reliable cognition. The mechanism may have produced accredited ammunition.

The same output can be honest advocacy at the producer, rigorous local research at the methods, and cargo-cult science at public composition. An institute may openly present evidence for an objective. The cargo-cult move is borrowing the authority of independent cumulative science, or a policy system treating the bounded analysis as an update to maintained knowledge.

Industry-sponsored drug and device studies more often report sponsor-favorable efficacy results and conclusions, a difference not explained by standard risk-of-bias tools.5 Ordinary method checks therefore do not exhaust sponsor influence. The harder remainder is a politically useful distribution of valid questions and fragments, consumed because no maintained model owns the composition.

Integration is a public good. Authoritative fragments are private assets.

V. The literature is an event log

A paper is naturally shaped like an event: under these conditions, using this method, these observations occurred and support this interpretation. The literature is an append-only record of such events. Information Is Not an Update states the interface distinction at smaller scale: arrival, retrieval, or quotation does not establish that a receiver’s model changed.

A record is not a current state.

Peer review, citation, and search can add, link, and retrieve events. None of them says what the model becomes after the event arrives.

A maintained state needs results an event log does not supply: narrower scope, changed parameters, branches versus contradictions, pending replication, deprecated public formulations, superseded models.

Without such states, contradiction becomes permanent inventory. “The literature is mixed” can remain a terminal sentence for decades. It does not open a work queue identifying which dimensions generate the mixture, what evidence would separate the branches, or who must resolve the public representation.

Systematic reviews and meta-analyses assemble the event log. They are often valuable, and commonly snapshots: another publication with a cutoff date. Cochrane treats currency as a distinct maintenance problem and describes living reviews that continually incorporate new evidence; resource limits make ordinary review maintenance difficult.6 That a special living form is needed reveals the default: synthesis itself can become another finished paper.

The paper is evidence. The knowledge product is the versioned state after the evidence has been received as an update.

Under that distinction, a paper is a proposed change. It should make visible the prior object, the proposed change, the conditions under which the change holds, the dependencies affected, and the result that would have supported another branch. The maintained state may reject the patch or leave it unresolved. It may not silently treat submission as incorporation.

VI. The planes that do land

Some domains already maintain an evaluated current state.

The Particle Data Group maintains the Review of Particle Physics as an evaluated current state, updated yearly and published every two years. For the 2024 edition it added 2,717 measurements from 869 new papers to 46,838 measurements retained from 12,909 earlier papers.7 The measurements do not merely coexist in a search result. A named collaboration maintains their evaluated relation.

Cochrane’s living systematic reviews continually monitor and incorporate relevant new evidence. Its prospective meta-analysis model can begin with the expectation that future studies will be integrated and work backward to coordinate the studies needed.6 That reverses the ordinary sequence: instead of producing studies and hoping synthesis occurs, the intended synthesis helps shape production.

Registered Reports alter another edge. Study protocols receive peer review before outcomes are known, and in-principle acceptance is not revoked merely because the result is negative or unexciting.8 The format makes methodological contribution more competitive with outcome-dependent publishability.

These mechanisms repair different layers:

Mechanism Layer repaired Remainder
Registered Reports Study selection and publication The result may still enter no maintained domain model
Living systematic review Currency of evidence synthesis The synthesis may still own no causal model or public decision
Particle Data Group Evaluated, versioned domain state Its success depends on a comparatively bounded ontology and organized maintainer community

Real cumulative maintenance exists. Institutions can be designed around it.

The transfer conditions matter. Particle physics has standardized measurements, strong shared questions, and consequences that eventually meet physical reality. Tax, education, institutional design, and public health contain more heterogeneous populations, contested objectives, long causal chains, and strategic users. Their models will branch more and close less.

That increases the need to represent the branches. It does not turn append-only publication into integration.

VII. A model that can receive a study

The repair begins with one bounded public question, not a central authority over truth.

For that question, a maintained model holds at least:

State What it records
Scope Population, mechanism, jurisdiction, horizon, counterfactual, and decision context
Claims Current propositions, parameter ranges, confidence, and exact provenance
Branches Competing explanations and whether disagreement is empirical, model-based, or normative
Dependencies Which conclusions and decisions rely on each claim
History Merges, rejections, scope changes, deprecations, reversals, and reasons

A new study then arrives as a proposed update. It names the prior object, the proposed change, the scope in which it applies, and the observations that would have supported another result.

The maintainers issue a state-changing response:

No result is entitled to merge because it was published. No result disappears because it was inconvenient. “Unresolved” is a valid state; silence is not.

The roles should remain separate. Contributors produce candidate updates. Maintainers preserve the current representation. Adversarial reviewers try to break proposed changes and the model itself. Decision users — ministries, clinicians, engineers, journalists, citizens — declare which version they used and where they departed from it.

Public iterability does not mean that anyone can overwrite the state. It means anyone can inspect provenance, submit a challenge, and fork the model when maintainers reject a live alternative. Authority comes from predictive performance, not merely from the office that hosts the record.

Maintenance must become a first-class scientific output. It needs named roles, budgets, career credit, succession, and a service level for processing updates. Otherwise the public-goods game simply reappears inside the new interface.

A paper can remain the unit of scholarly communication. It can no longer be the terminal unit of cumulative knowledge.

VIII. The model can become another cult

A maintained model creates a new attack surface. It can freeze a field’s ontology, suppress heterodox evidence, mistake consensus for truth, or optimize the appearance of currency while important challenges wait unprocessed.

The repair therefore needs its own correction architecture: provenance, reasons, named maintainers, visible challenges, possible forks, turnover, reopen triggers, and tests against predictions rather than internal agreement alone.

No maintainer chooses society’s terminal values by relabeling them as evidence. A maintained tax model can estimate revenue, incidence, output, administrative cost, and distribution under different assumptions. It cannot derive how those outcomes should be traded. It makes the empirical branches and the value choice separable enough that neither can hide inside the other.

Basic and exploratory research also remain legitimate. Correct Is Not Consequence separates a recoverable deposit from an intervention that claims present enactment. A measurement may be a deposit for a future question that does not yet exist. A mathematical result may have no current decision user. A surprising observation may be too early to place. The honest result is then “recoverable, presently unrouted,” not a fictional claim that the operative public model has already improved.

The test applies recursively here. A public-model institution can optimize update counts, consensus, prestige, funding, or the survival of its own ontology while retaining the language of correction. Its authority must remain conditional on whether it helps users make more accurate predictions, distinguish live alternatives, and revise consequential decisions.

This page is not exempt. It is a proposed update, not an enacted improvement. If it becomes one more essay cited as proof that integration is missing, while no maintained frame or model-update experiment is built, it will instantiate its own diagnosis.

IX. Change the terminal condition

A study has two different completion conditions.

The local study is complete when its question, method, evidence, and report are sound enough to stand as an artifact. Its cumulative contribution is complete when the artifact has received a result against the bounded model it claims to inform.

Conflating the two is the paper equilibrium. It lets “published” mean “knowledge advanced” without requiring anyone to identify the advance.

A compact audit exposes the missing layer. First the authority-use questions, then the update:

Authority-use

  1. Who selected and funded the question, and which nearby comparators, outcomes, and counterfactuals were excluded?
  2. What bounded proposition does the study establish, and what larger policy proposition is it being made to license?
  3. What does the expert’s credential add: evidence, or only standing?
  4. Who gains if the study remains a detachable citation rather than one constrained update?

Model-update

  1. What bounded model or public question was supposed to change?
  2. Which exact claim, parameter, edge, or uncertainty did the study put at risk?
  3. Who owns merge, rejection, branching, or deferral?
  4. What prior claim or public formulation no longer survives?
  5. Which prediction, experiment, or decision is now different?
  6. What later observation would reopen the update?

The test permits rejection, uncertainty, and multiple branches. It refuses two substitutions: counting publication as if the shared map changed, and spending a bounded true result as if it had settled the larger decision.

The scientific system need not be fraudulent to become cargo cult. It only has to reward every visible part before the invisible whole is corrected, and some of those parts need the missing whole so that unintegrated papers remain spendable as authority. Then everything can work perfectly: studies are conducted, journals filled, grants completed, careers advanced, experts cited, and evidence invoked.

The runway is busy. The missing plane is a public model that becomes more accurate because the studies happened.


Related:

Sources and Notes

1. Tax-response decomposition. Thomas Piketty, Emmanuel Saez, and Stefanie Stantcheva, “Optimal Taxation of Top Labor Incomes: A Tale of Three Elasticities”, American Economic Journal: Economic Policy 6(1), 2014, 230–271. See also Joel Slemrod and Wojciech Kopczuk, “The Optimal Elasticity of Taxable Income”, NBER Working Paper 7922, 2000. These papers motivate the opening as a composition problem: multiple estimated channels can exist while public argument returns to a binary. They do not document one named debate as a completed cargo-cult verdict.

2. Cargo Cult Science. Richard P. Feynman, “Cargo Cult Science”, Caltech commencement address, 1974. Feynman’s diagnosis is scientific form without the integrity needed for the result. This essay extends the test to composition among valid studies.

3. Selection for publishability. Brian A. Nosek, Jeffrey R. Spies, and Matt Motyl, “Scientific Utopia II”, Perspectives on Psychological Science 7(6), 2012. Andrew D. Higginson and Marcus R. Munafò, “Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions”, PLOS Biology 14(11), 2016. Paul E. Smaldino and Richard McElreath, “The Natural Selection of Bad Science”, Royal Society Open Science 3, 2016. These show how publication-based career competition can select distorted methods. They are neighboring mechanisms, not the residual: the paper equilibrium survives even if every paper is valid.

4. Publication bias and apparent fact. Silas Boye Nissen, Tali Magidson, Kevin Gross, and Carl T. Bergstrom, “Publication Bias and the Canonization of False Facts”, eLife 5:e21451, 2016.

5. Sponsorship and outcome. Andreas Lundh, Joel Lexchin, Barbara Mintzes, Jeppe Berg Schroll, and Lisa Bero, “Industry sponsorship and research outcome”, Cochrane Database of Systematic Reviews 2017, Issue 2, Art. No. MR000033. Seventy-five methodological studies: industry-sponsored drug and device studies more often reported sponsor-favorable efficacy (RR 1.27) and conclusions (RR 1.34); the difference was not explained by standard risk-of-bias tools except that industry studies reported satisfactory blinding more often. Neighboring evidence that method checks do not exhaust sponsor influence; not the definition of the composition failure.

6. Living evidence. James Thomas et al., “Prospective Approaches to Accumulating Evidence,” Chapter 22 in the Cochrane Handbook for Systematic Reviews of Interventions, version 6.5, 2024. Currency is treated as a distinct maintenance problem; living reviews and prospective meta-analysis are the special forms, which is evidence that ordinary synthesis defaults to another finished paper.

7. Maintained domain state. Particle Data Group, “About the Particle Data Group”, describing the 2024 edition of the Review of Particle Physics: 2,717 new measurements from 869 papers added to 46,838 measurements from 12,909 earlier papers. Journal publication: S. Navas et al. (Particle Data Group), Physical Review D 110, 030001 (2024). The PDG is a positive case of evaluated, versioned domain state, not a claim that every field can copy its ontology.

8. Outcome-independent publication. Center for Open Science, “Registered Reports”. In-principle acceptance is not revoked merely because results are negative or unexciting. The format repairs study selection; it does not by itself create a maintained domain model.

Claim boundary. The essay does not claim that all science is cargo cult, that all papers require immediate policy use, that one authority should impose a single model, or that every advocacy organization is fraudulent. An openly ideological research shop can be honest about producing detachable studies. The cargo-cult substitution is treating those studies as cumulative scientific authority in a system that maintains no joined model.