nishi code wiki / research / metrology council
A council of experts, done honestly: artifact diversity beats headcount
Published · lineage: forks off the world physics regimes brief, which produced transform laws that now need adjudicating · benchmark: the Class-III mandible four-method battery, whose durable finding is that methods agree on gross verdict and disagree on localization, each carrying a different artifact
1. The standards exist, and they deliberately supply no numbers
ASME V&V 10 (computational solid mechanics) gives a validation hierarchy — system decomposed into subsystem cases, benchmark cases and unit problems, the “validation pyramid” — with a Phenomena Identification and Ranking Table as the first step, and insists verification and validation be performed independently. It contains no accuracy criterion and no acceptance threshold; a companion document exists solely because the consensus standard deliberately omitted worked examples. V&V 20 supplies the validation-comparison machinery, and V&V 40 the risk-informed credibility framing. Hierarchy discussion: Sandia, validation hierarchy.
Uncertainty has its own balloted stack: JCGM 100 (GUM), JCGM 101 (Monte Carlo propagation) and JCGM 106 (conformity assessment and decision rules) — the last being the one that tells you how to accept or reject in the presence of measurement uncertainty, which is the question a gate actually asks. And the concept that matters most for us has a formal definition: definitional uncertainty, the irreducible floor imposed by an incompletely specified measurand. “Realistic” is definitionally uncertain, and no instrument can resolve below that floor.
2. The council: seven members chosen for disjoint blind spots
| # | Member | What it proves | The artifact — what it cannot see |
|---|---|---|---|
| 1 | Method of Manufactured Solutions, observed order-of-accuracy | the engine solves its own equations right, in full generality; the only member with a measured detection profile | model form — it cannot tell you the equations are the wrong ones |
| 2 | Metamorphic relations (Galilean boost, rotation, translation, index permutation, unit rescale, time reversal) | needs no oracle at all; the only member that scales to regimes with no exact solution — contact, soft tissue | equivariant errors: a wrong dynamics respecting the same symmetries passes every relation |
| 3 | Validation against physical experiment, reporting E ± uval | the only member that touches reality; everything else is internal consistency | model error swamped by large input or data uncertainty — which is why the three terms must be reported separately, never as a total |
| 4 | Mutation / non-vacuity | judges the other members | the algorithm-deletion and overflow fault classes — covered by 1 and 2 |
| 5 | Sobol / Morris sensitivity | where the error lives, once another member has said whether — this is our EDMA analogue | non-variance errors; and it must never be allowed to produce a verdict |
| 6 | Conservation and dissipation audits, split by regime | exact conservation for the rigid regime; a dissipation inequality for soft tissue, where conservation is simply the wrong invariant | conservative-but-wrong-branch selection |
| 7 | Differential testing vs an independently-lineaged engine | implementation-specific errors | the worst artifact in the council — shared misconception |
Member 1 must run first: code verification is a stated prerequisite for any grid-convergence-derived error bar, so a validation number computed on unverified code is not a measurement (ASME JFE editorial policy on numerical accuracy · Sandia MMS report). Member 7 is deliberately asymmetric: treat agreement with a foreign engine as near-zero evidence and disagreement as high value, because disagreement cannot be manufactured by shared lineage. Tooling, with licences: SALib (MIT), Dakota (LGPL), Descartes extreme mutation (Apache-2.0).
3. The trap, and it is measured rather than theoretical
4. The second-order sting
Two consequences follow that feel wrong and are right. Unanimity is not the strong signal it feels like: under contagion-type dependence the Condorcet guarantee is not merely weakened but gone, with lock-in on the incorrect alternative possible, and majority reliability can become a decreasing function of group size — adding a seventh member can lower accuracy (Stanford Encyclopedia, jury theorems). And the latent-class machinery one would reach for as a rescue — Dawid–Skene and its descendants — all assume conditional independence, so under shared bias they will confidently infer the correlated bloc’s shared error as ground truth. It fails in the direction that hurts.
5. What to build instead of a vote count
- Measure effective sample size against a Condorcet null before trusting any aggregate. If seven methods are worth 1.8 votes, that is the headline finding.
- Run the two falsification diagnostics that ended the truth-plus-error paradigm in climate: does the council mean converge to the reference as members are added, and does pairwise member-error correlation average to zero? In climate both answered no. If both answer no, the council is not truth-centred and the spread represents collective uncertainty with the truth not necessarily at the centre.
- Enumerate correlation channels explicitly: shared code, shared lineage, shared tuning data, shared floating-point substrate, shared author, shared fixture construction. Two members with zero shared code can still be correlated through shared tuning data.
- Roll up by minimum, report a vector. Decompose disagreement with Mandel h versus k (location or spread?) and a Youden plot (systematic or random). Report Bland–Altman bias and spread in physical units — the only instrument in this entire sweep that refuses to invent a threshold.
- Never seat a member that has not passed a non-vacuity gate. A gate that cannot fail is a voter with error rate zero and information content zero: it inflates agreement while contributing nothing. Not hypothetical — one industrial study found 20% of first-run formal properties trivially valid, and another found 294 rotten green tests in 19,905.
- Expect kappa near zero at 95% raw agreement and do not report it as disagreement. That is the prevalence paradox produced by pass-skewed marginals — a fabrication manufactured by the statistic. Never report kappa without prevalence and bias indices beside it (Landis & Koch 1977 for the interpretation bands, which are themselves convention, not derivation).
6. Where we actually stand
7. The second seat, and the first evidence that any of this works
Member 6 — conservation and dissipation — shipped the same day, deliberately chosen because it covers member 2's declared blind spot: an equivariant error, a wrong dynamics that respects every symmetry you test it against. It rests on an invariant that integers make unusually sharp. A damped system with no input must come to rest, and in integer arithmetic “at rest” is not a small number, it is state[t+1] == state[t], byte for byte, and holding. No tolerance, no oracle, no reference data.
SD_G/C the damping term is exactly zero, and below a displacement of SD_G/K the spring term is exactly zero. Inside both dead zones no force acts at all and the point coasts ballistically — a measured period-30 limit cycle of ±3.2 mm at 2.0 Hz, still running undecayed after 200,000 ticks with a completely stationary anchor. Tissue oscillating forever on a character standing perfectly still. Fixed at root, with the threshold derived from the solver's own constants rather than invented; the metamorphic seat was then rebuilt and re-run to confirm all twelve relations still hold, which is precisely what having the first seat was for.Two further findings are worth more than the fix. Symmetry is blind to energy — that artifact was listed in the table above as a theoretical weakness of member 2, and here it is measured on our own code the first time two seats disagreed. And the pre-existing teeth were blind too: the older gate checks that energy falls below an eighth of its peak and that an impulse returns to under a quarter of its peak, and a solver in a permanent limit cycle passes both. A decay-ratio test cannot see a limit cycle; only an exact fixed-point test can. This is one data point on independence, not an audit — effective sample size remains unmeasured, and two seats disagreeing once is not a measurement of how often they will.
Also reported rather than asserted: under indentation the ring's cross-section area rises to 107% of rest at the deepest probe. Truly incompressible tissue would hold near 100%, so the volume-redistribution law slightly over-compensates and the excess grows with depth. Gaining volume is the safe direction for a bulge, and inventing a tight upper bound here would repeat a mistake this gate had already made once: its first version inherited a neighbouring gate's 85% floor, mirrored it into a band, and was vacuous — the control case sits at 90–94%, comfortably inside it. A threshold inherited without measuring the control against it is an assumption, not a bound.
UNVERIFIED / declared gaps
- The standards are paywalled. ASME and ISO documents were characterised from catalogue pages, abstracts and secondary summaries, not from the balloted text. Anything we adopt from them should be checked against a purchased copy before it becomes a shipped constant.
- Most agreement thresholds are convention, not derivation — the kappa interpretation bands, the %GRR limits, and the sensitivity-index cutoffs are all conventions, and at least one is admitted arbitrary in print. We adopt none of them as pass lines without saying so.
- No standard covers soft tissue. The solid-mechanics V&V standard contains no hyperelastic or soft-tissue benchmark extension.
- Several conflicts were found and are reported rather than averaged, including two different Bland–Altman limit conventions and four incompatible mutant-equivalence rates computed over different denominators.
- Not yet measured by us: everything in §5. This brief is the design of a council, not a report on one.