The thesis box says this is not an argument that the world has twelve layers. Then each layer gets a question, a healthy primitive, and a failure family.
A filing system for essays does not need healthy primitives. Those are causal claims. The periodic-table disclaimer is attached to the count, not to the primitives.
The page is explicit: working hypothesis, useful enough to organise the work, not the periodic table. Layer number is an address, not a causal rank. The diagnostic move is to construct a case-specific dependency, not to treat numerical adjacency as causation.
Then the count is not the claim. The primitives are. What observation would collapse two layers, merge measurement into ledger, or add a thirteenth? If the answer is "a better partition that organises the corpus," the test is editorial.
The six preconditions already describe a stack-shaped system: boundary, telos, channels, conversion path, persistent carriers, localisable failure. Without them the page calls the Stack a metaphor. The scan applies where the failure is already in this grammar.
Inter-rater: give two compilers the same failure and the layer list. If they systematically land on different addresses, the addresses are labels.
Agreement is not enough. Hold the failure fixed and ask which layer, if repaired, would unblock the claimed outcome. If several layers each pass that counterfactual, "one is usually load-bearing" is a writing convention.
"Smallest joint set of domains" is the actual object, and it has no algebra. Sequential, parallel-bonded, selector, allocation, substrate, topology are named later. The layer list does not tell you which geometry you are in.
Agreement on a layer can hide disagreement on the geometry. The inter-rater test has to score the geometry first or the addresses will look more stable than they are.
Reproduction as competence-stock depletion is already a ledger object. The page's own reproduction-test pointer treats mistaking a stock for a flow as the error. Then layers 5 and 11 are the same stock at different time derivatives, not different domains.
That is a collapse candidate the inter-rater test should include, next to measurement-versus-ledger. Specificity of vocabulary is not evidence that the cut is causal.