nishi code wiki / research / the council of instruments

The council of instruments: consensus metrology for verdicts we can trust

Published · lineage: forks off the measurement constitution and ruler-craft · operator's question: can we convene a council of measurement experts so that when 7 of 10 agree, we know?

Freshness declared. Verified to the author's May-2026 horizon plus the corpus; the May→July 2026 tail is UNVERIFIED (most volatile: learned perceptual metrics and identity-embedding leaders). Numeric bands are commonly-cited ranges pending the citation-archive pass.
The answer is yes — and international metrology already formalised it. But the first finding indicts our current panel: a council of CORRELATED experts is ONE expert wearing many hats. Cephalometric, Procrustes, EDMA and TPS all consume the SAME landmark detector. If that detector's anatomical prior is biased (we measured it resisting contradicting paint at ~30% per unit), all four inherit the bias and agree confidently and wrongly. Diversity must be in the SENSING, not only in the mathematics.

1. The independence audit: what our panel actually contains

familywhat it sensesindependent of the landmarker?
Cephalometric · Procrustes · EDMA · TPSthe same 468 detected pointsNO — one sensor, four maths. Counts as ONE council seat.
Silhouette / contour (Hausdorff on outlines, turning function, Fourier descriptors, visual hull bounds)segmentation masksYES — different sensor, and hull violations are geometric CERTAINTIES, not estimates
Dense depth / normals (metric monocular depth, normal estimators)a different network entirelyYES — and it sees the depth axis our frontal landmarks cannot
Identity embeddings (recognition nets, e.g. the ArcFace family)learned identity manifold from millions of facesYES — answers “same person?” with no landmarks at all
Learned perceptual distance (LPIPS, DISTS, CGVQM/CVVDP class)deep feature statistics, calibrated to human judgementYES — and it is the only seat trained on perception
Texture / frequency (ISO 25178 areal params, PSD, wavelet energy, GLCM)surface statistics, not shapeYES — the pore/stubble scale lives here
Colorimetry (CIE ΔE00, ITA°)calibrated colour spaceYES — a genuinely different instrument class
Point-process statistics (Ripley's K, pair-correlation, nearest-neighbour, Wasserstein on histograms)spatial distribution of featuresYES — the mole/tubercle/follicle seat
Biological plausibility (allometry, anthropometric canon, damage bounds)biology, not the reference imageYES — it can convict a render the reference cannot
Human panel (operator; psychophysics 2AFC, JND ladders)a personYES — the gold seat. Already registered in our loop.

Verdict: we have 1 seat filled and 9 available. Adding a tenth landmark metric buys nothing; adding a silhouette seat buys a genuine second opinion.

2. How real metrology combines disagreeing experts (this is the machinery the operator is asking for)

frameworkwhat it doesour use
Key comparisons (CIPM MRA): Key Comparison Reference Value + degrees of equivalenceNational institutes measure the same artefact and reconcile into ONE reference value, each lab reported as a deviation with its own uncertaintyThe exact template. Each instrument reports a value + uncertainty; the council publishes a consensus and every member's degree of equivalence to it.
Consensus estimators: weighted mean, DerSimonian–Laird random effects, Mandel–Paule, Vangel–Rukhin (as implemented in the NIST Consensus Builder)Combine discrepant measurements with unequal uncertainties; random-effects models admit that instruments differ systematically, not just noisilythe consensus value, with a between-instrument variance term that is itself informative
Birge ratio / χ² consistency testAsks whether the experts disagree MORE than their claimed uncertainties permit. Birge > 1 means someone understated their uncertaintyThis is the literal “do we know for sure” instrument. A council that agrees within its uncertainties is trustworthy; one that scatters wider is lying about precision — and the test names it.
Robust statistics (median, MAD, Huber — Algorithm A of the proficiency-testing standard)Consensus that ignores outlier experts instead of being dragged by themclosest to the operator's “7 of 10” intuition, and defensible: median + MAD, with outliers reported not deleted
Proficiency scoring (z-score, zeta-score, En number)Scores each participant against the consensus relative to uncertainty: |z| ≤ 2 satisfactory, > 3 unacceptableevery instrument gets a standing score; a chronically |z| > 3 seat is demoted, and that demotion is evidence
ISO 5725 repeatability vs reproducibilitySeparates within-method (r) from between-method (R) variationtells us whether our millimetres are limited by one gauge or by method choice
Gauge R&R / ANOVA variance componentsPartitions observed variation into part vs operator vs instrumentran a first pass 2026-08-03: crop-repeatability alone gives 1σ ≈ 0.59 mm, U(k=2) ≈ ±1.18 mm — so any “error” under ~1.2 mm was never real
Bland–Altman agreement analysis (Altman & Bland 1986)Bias + limits of agreement between two methods — the clinical standard, and it proves correlation is NOT agreementpairwise instrument audits; catches a seat that tracks the trend but sits offset
Latent-truth models (Dawid–Skene 1979, modern crowd-consensus descendants)Estimates each rater's reliability AND the hidden truth simultaneously — with no ground truth.the mature council: weights are LEARNED from agreement history rather than assigned by me
Bayesian model averaging / hierarchical consensusPosterior over the truth given multiple noisy instruments with learned reliabilitiesthe principled endgame once each seat has a calibrated noise model
Conformal predictionDistribution-free intervals with guaranteed coverage from ANY ensembleconverts council spread into an honest interval without assuming normality
Monte Carlo uncertainty propagation (GUM Supplement 1)Propagates uncertainty through non-linear models where the linear formula failsour IPD prior × projection × detector chain is non-linear — this is the correct propagation

3. The rules the council must obey (design, not decoration)

rulewhy
Independence over count. Seats are counted by SENSOR, not by metric. Four maths on one detector = one seat.correlated errors do not cancel; they masquerade as consensus
Every seat may ABSTAIN. An instrument that always returns a number is not measuring — it must be able to say “this quantity is unobservable to me” (our frontal landmarks on depth; single-view metrology with no vertical in frame).a forced vote is fabricated evidence
Disagreement is a FINDING, never an average to bury. Report the consensus AND the Birge ratio AND every degree of equivalence.the 08-03 jaw result: the anchored seat said “fine”, the anchor-free seat said 16.3 mm, the human seat said “gaunt” — 2 of 3 were right and the disagreement WAS the discovery
No verdict below the gauge floor. With U(k=2) ≈ ±1.2 mm today, a 0.9 mm “improvement” is noise.the honest floor on every claim we publish
Voting rules have known pathologies (Arrow's theorem): plain majority over incommensurable metrics is not sound.hence weighted/robust consensus with uncertainties, not a show of hands — the metrology answer to the “7 of 10” instinct
The human seat is not outvoted by machines. On presence/aliveness no automated metric exists; the operator's verdict is the reference, and machine consensus against it is a machine failure.declared limitation, already standing doctrine

4. Adoption order

#seat / actwhy now
C1Uncertainty budget per seat + the gauge floor published (extends the 0.59 mm first pass to all contributors)without it no seat can be scored and no verdict has a floor
C2Silhouette seat (contour Hausdorff + hull bound)cheapest genuinely independent sensor; our renders already emit exact silhouettes
C3Birge ratio + degrees of equivalence in the verdict outputturns “they agree” into a tested claim
C4Colorimetric seat (ITA°/ΔE00) and texture seat (ISO 25178)pure math on rendered pixels; retires eyeball colour/roughness verdicts
C5Identity-embedding seat and perceptual seat (LPIPS/CVVDP class)two independent learned opinions, one of them perception-calibrated
C6Dawid–Skene weighting once seats have historyreliability learned from agreement, not assigned
C7Conformal intervals on the published scoreboard numbersguaranteed coverage on every “best” claim — the forensic standard

UNVERIFIED / declared gaps