nishi code wiki / research / metrology council

A council of experts, done honestly: artifact diversity beats headcount

Published · lineage: forks off the world physics regimes brief, which produced transform laws that now need adjudicating · benchmark: the Class-III mandible four-method battery, whose durable finding is that methods agree on gross verdict and disagree on localization, each carrying a different artifact

The finding that reframes the question. The ask was “a council where seven of ten agreeing means we know.” The literature is unusually blunt here: an agreement count is not evidence strength until effective independence has been measured, and every field that tried counting votes discovered the same thing independently. A council’s value is bounded by the union of what its members can see, so the selection criterion is not “best methods” but minimal overlap of blind spots. Three institutions reached this by three different roads and all three refused the shortcut: PCAST refused to multiply examiner error rates, NASA-STD-7009 refused to average across credibility factors, and the climate community refused to prune its ensemble.

1. The standards exist, and they deliberately supply no numbers

ASME V&V 10 (computational solid mechanics) gives a validation hierarchy — system decomposed into subsystem cases, benchmark cases and unit problems, the “validation pyramid” — with a Phenomena Identification and Ranking Table as the first step, and insists verification and validation be performed independently. It contains no accuracy criterion and no acceptance threshold; a companion document exists solely because the consensus standard deliberately omitted worked examples. V&V 20 supplies the validation-comparison machinery, and V&V 40 the risk-informed credibility framing. Hierarchy discussion: Sandia, validation hierarchy.

Uncertainty has its own balloted stack: JCGM 100 (GUM), JCGM 101 (Monte Carlo propagation) and JCGM 106 (conformity assessment and decision rules) — the last being the one that tells you how to accept or reject in the presence of measurement uncertainty, which is the question a gate actually asks. And the concept that matters most for us has a formal definition: definitional uncertainty, the irreducible floor imposed by an incompletely specified measurand. “Realistic” is definitionally uncertain, and no instrument can resolve below that floor.

We had independently arrived at two of these. Our minimum-across-axes rollup is NASA-STD-7009’s balloted choice, and our banked “gate on the error, never on correlation” law is the Bland–Altman position in standard form. Convergent arrival is mild evidence the shape is right.

2. The council: seven members chosen for disjoint blind spots

#MemberWhat it provesThe artifact — what it cannot see
1Method of Manufactured Solutions, observed order-of-accuracythe engine solves its own equations right, in full generality; the only member with a measured detection profilemodel form — it cannot tell you the equations are the wrong ones
2Metamorphic relations (Galilean boost, rotation, translation, index permutation, unit rescale, time reversal)needs no oracle at all; the only member that scales to regimes with no exact solution — contact, soft tissueequivariant errors: a wrong dynamics respecting the same symmetries passes every relation
3Validation against physical experiment, reporting E ± uvalthe only member that touches reality; everything else is internal consistencymodel error swamped by large input or data uncertainty — which is why the three terms must be reported separately, never as a total
4Mutation / non-vacuityjudges the other membersthe algorithm-deletion and overflow fault classes — covered by 1 and 2
5Sobol / Morris sensitivitywhere the error lives, once another member has said whether — this is our EDMA analoguenon-variance errors; and it must never be allowed to produce a verdict
6Conservation and dissipation audits, split by regimeexact conservation for the rigid regime; a dissipation inequality for soft tissue, where conservation is simply the wrong invariantconservative-but-wrong-branch selection
7Differential testing vs an independently-lineaged engineimplementation-specific errorsthe worst artifact in the council — shared misconception

Member 1 must run first: code verification is a stated prerequisite for any grid-convergence-derived error bar, so a validation number computed on unverified code is not a measurement (ASME JFE editorial policy on numerical accuracy · Sandia MMS report). Member 7 is deliberately asymmetric: treat agreement with a foreign engine as near-zero evidence and disagreement as high value, because disagreement cannot be manufactured by shared lineage. Tooling, with licences: SALib (MIT), Dakota (LGPL), Descartes extreme mutation (Apache-2.0).

3. The trap, and it is measured rather than theoretical

Treating agreement count as evidence strength without measuring effective independence — and doing so most confidently where independence is most violated. The measured discounts are large and consistent: twenty-four climate models behave like seven to nine; nine language-model judges drawn from seven model families are worth roughly two effective votes, with panel accuracy landing 8–22 points below the independent ideal, and published aggregation methods closing at most 11% of that gap even when given the correct answers. The bottleneck is the correlation, not the aggregation algorithm — no clever weighting rescues a correlated panel. Twenty-seven independently written programs subjected to a million tests had independence rejected at 99% confidence. And PCAST formally refused to let two examiners’ error rates be multiplied, on the grounds that errors “may depend on the difficulty of the problem and thus be correlated.”

4. The second-order sting

Correlation is driven by item difficulty, so a council is most correlated exactly on the hard cases where it is most needed — and it is error-type-specific. In the fingerprint black-box work, no two examiners made false-positive errors on the same comparison, yet 85% of examiners made false-negative errors, concentrated on the same half of the image pairs (Ulery et al., PLOS ONE 2012 · NIJ overview). The same council was independent on one error type and strongly correlated on the other. The machine-learning literature arrives from the opposite direction at the same place: ensemble diversity does not meaningfully contribute to uncertainty quantification out of distribution. Therefore report independence per error type — false-accept and false-reject separately — never as one number.

Two consequences follow that feel wrong and are right. Unanimity is not the strong signal it feels like: under contagion-type dependence the Condorcet guarantee is not merely weakened but gone, with lock-in on the incorrect alternative possible, and majority reliability can become a decreasing function of group size — adding a seventh member can lower accuracy (Stanford Encyclopedia, jury theorems). And the latent-class machinery one would reach for as a rescue — Dawid–Skene and its descendants — all assume conditional independence, so under shared bias they will confidently infer the correlated bloc’s shared error as ground truth. It fails in the direction that hurts.

5. What to build instead of a vote count

6. Where we actually stand

The metamorphic seat is now built; the rest are not. Member 2 shipped 2026-08-03 as an oracle-free gate over our soft-tissue solver: spatial translation, axis permutation, time translation, reflection, quiescence, impulse reflection and amplitude monotonicity all hold exactly, each detector proven able to fail against a deliberately broken copy of the solver. Two results worth stating plainly. Galilean invariance does not hold — damping is applied against the world frame rather than the body, so a body in uniform motion carries a permanent lag of C·v/K that never settles; that is measured, not assumed, and it is now an open question for the physics rather than a defect we have quietly fixed. And the first build of the gate was vacuous: its drive exceeded the solver's displacement clamp, so every check was measuring the clamp instead of the integrator, and only the non-vacuity tooth caught it. A detector run outside the operating envelope of the thing it audits measures the envelope, not the thing. Still not built: manufactured solutions, validation uncertainty decomposition, sensitivity localization, differential testing against a foreign engine. Effective independence has never been measured, not once. Two seats and no independence audit is still not a council — it is a set of gates, and calling it a council would be exactly the overclaim this brief exists to prevent. The seat count is not a target: seats earn their place by covering a blind spot no other member can see.

7. The second seat, and the first evidence that any of this works

Member 6 — conservation and dissipation — shipped the same day, deliberately chosen because it covers member 2's declared blind spot: an equivariant error, a wrong dynamics that respects every symmetry you test it against. It rests on an invariant that integers make unusually sharp. A damped system with no input must come to rest, and in integer arithmetic “at rest” is not a small number, it is state[t+1] == state[t], byte for byte, and holding. No tolerance, no oracle, no reference data.

On its first run it found a defect that all twelve metamorphic relations had passed over. Our soft-tissue solver computes both of its force terms as integer divisions that truncate toward zero. Below a velocity of SD_G/C the damping term is exactly zero, and below a displacement of SD_G/K the spring term is exactly zero. Inside both dead zones no force acts at all and the point coasts ballistically — a measured period-30 limit cycle of ±3.2 mm at 2.0 Hz, still running undecayed after 200,000 ticks with a completely stationary anchor. Tissue oscillating forever on a character standing perfectly still. Fixed at root, with the threshold derived from the solver's own constants rather than invented; the metamorphic seat was then rebuilt and re-run to confirm all twelve relations still hold, which is precisely what having the first seat was for.

Two further findings are worth more than the fix. Symmetry is blind to energy — that artifact was listed in the table above as a theoretical weakness of member 2, and here it is measured on our own code the first time two seats disagreed. And the pre-existing teeth were blind too: the older gate checks that energy falls below an eighth of its peak and that an impulse returns to under a quarter of its peak, and a solver in a permanent limit cycle passes both. A decay-ratio test cannot see a limit cycle; only an exact fixed-point test can. This is one data point on independence, not an audit — effective sample size remains unmeasured, and two seats disagreeing once is not a measurement of how often they will.

Also reported rather than asserted: under indentation the ring's cross-section area rises to 107% of rest at the deepest probe. Truly incompressible tissue would hold near 100%, so the volume-redistribution law slightly over-compensates and the excess grows with depth. Gaining volume is the safe direction for a bulge, and inventing a tight upper bound here would repeat a mistake this gate had already made once: its first version inherited a neighbouring gate's 85% floor, mirrored it into a band, and was vacuous — the control case sits at 90–94%, comfortably inside it. A threshold inherited without measuring the control against it is an assumption, not a bound.

UNVERIFIED / declared gaps