nishi code wiki / research / the council of instruments
The council of instruments: consensus metrology for verdicts we can trust
Published · lineage: forks off the measurement constitution and ruler-craft · operator's question: can we convene a council of measurement experts so that when 7 of 10 agree, we know?
Freshness declared. Verified to the author's May-2026 horizon plus the corpus; the May→July 2026 tail is UNVERIFIED (most volatile: learned perceptual metrics and identity-embedding leaders). Numeric bands are commonly-cited ranges pending the citation-archive pass.
The answer is yes — and international metrology already formalised it. But the first finding indicts our current panel: a council of CORRELATED experts is ONE expert wearing many hats. Cephalometric, Procrustes, EDMA and TPS all consume the SAME landmark detector. If that detector's anatomical prior is biased (we measured it resisting contradicting paint at ~30% per unit), all four inherit the bias and agree confidently and wrongly. Diversity must be in the SENSING, not only in the mathematics.
1. The independence audit: what our panel actually contains
| family | what it senses | independent of the landmarker? |
|---|---|---|
| Cephalometric · Procrustes · EDMA · TPS | the same 468 detected points | NO — one sensor, four maths. Counts as ONE council seat. |
| Silhouette / contour (Hausdorff on outlines, turning function, Fourier descriptors, visual hull bounds) | segmentation masks | YES — different sensor, and hull violations are geometric CERTAINTIES, not estimates |
| Dense depth / normals (metric monocular depth, normal estimators) | a different network entirely | YES — and it sees the depth axis our frontal landmarks cannot |
| Identity embeddings (recognition nets, e.g. the ArcFace family) | learned identity manifold from millions of faces | YES — answers “same person?” with no landmarks at all |
| Learned perceptual distance (LPIPS, DISTS, CGVQM/CVVDP class) | deep feature statistics, calibrated to human judgement | YES — and it is the only seat trained on perception |
| Texture / frequency (ISO 25178 areal params, PSD, wavelet energy, GLCM) | surface statistics, not shape | YES — the pore/stubble scale lives here |
| Colorimetry (CIE ΔE00, ITA°) | calibrated colour space | YES — a genuinely different instrument class |
| Point-process statistics (Ripley's K, pair-correlation, nearest-neighbour, Wasserstein on histograms) | spatial distribution of features | YES — the mole/tubercle/follicle seat |
| Biological plausibility (allometry, anthropometric canon, damage bounds) | biology, not the reference image | YES — it can convict a render the reference cannot |
| Human panel (operator; psychophysics 2AFC, JND ladders) | a person | YES — the gold seat. Already registered in our loop. |
Verdict: we have 1 seat filled and 9 available. Adding a tenth landmark metric buys nothing; adding a silhouette seat buys a genuine second opinion.
2. How real metrology combines disagreeing experts (this is the machinery the operator is asking for)
| framework | what it does | our use |
|---|---|---|
| Key comparisons (CIPM MRA): Key Comparison Reference Value + degrees of equivalence | National institutes measure the same artefact and reconcile into ONE reference value, each lab reported as a deviation with its own uncertainty | The exact template. Each instrument reports a value + uncertainty; the council publishes a consensus and every member's degree of equivalence to it. |
| Consensus estimators: weighted mean, DerSimonian–Laird random effects, Mandel–Paule, Vangel–Rukhin (as implemented in the NIST Consensus Builder) | Combine discrepant measurements with unequal uncertainties; random-effects models admit that instruments differ systematically, not just noisily | the consensus value, with a between-instrument variance term that is itself informative |
| Birge ratio / χ² consistency test | Asks whether the experts disagree MORE than their claimed uncertainties permit. Birge > 1 means someone understated their uncertainty | This is the literal “do we know for sure” instrument. A council that agrees within its uncertainties is trustworthy; one that scatters wider is lying about precision — and the test names it. |
| Robust statistics (median, MAD, Huber — Algorithm A of the proficiency-testing standard) | Consensus that ignores outlier experts instead of being dragged by them | closest to the operator's “7 of 10” intuition, and defensible: median + MAD, with outliers reported not deleted |
| Proficiency scoring (z-score, zeta-score, En number) | Scores each participant against the consensus relative to uncertainty: |z| ≤ 2 satisfactory, > 3 unacceptable | every instrument gets a standing score; a chronically |z| > 3 seat is demoted, and that demotion is evidence |
| ISO 5725 repeatability vs reproducibility | Separates within-method (r) from between-method (R) variation | tells us whether our millimetres are limited by one gauge or by method choice |
| Gauge R&R / ANOVA variance components | Partitions observed variation into part vs operator vs instrument | ran a first pass 2026-08-03: crop-repeatability alone gives 1σ ≈ 0.59 mm, U(k=2) ≈ ±1.18 mm — so any “error” under ~1.2 mm was never real |
| Bland–Altman agreement analysis (Altman & Bland 1986) | Bias + limits of agreement between two methods — the clinical standard, and it proves correlation is NOT agreement | pairwise instrument audits; catches a seat that tracks the trend but sits offset |
| Latent-truth models (Dawid–Skene 1979, modern crowd-consensus descendants) | Estimates each rater's reliability AND the hidden truth simultaneously — with no ground truth. | the mature council: weights are LEARNED from agreement history rather than assigned by me |
| Bayesian model averaging / hierarchical consensus | Posterior over the truth given multiple noisy instruments with learned reliabilities | the principled endgame once each seat has a calibrated noise model |
| Conformal prediction | Distribution-free intervals with guaranteed coverage from ANY ensemble | converts council spread into an honest interval without assuming normality |
| Monte Carlo uncertainty propagation (GUM Supplement 1) | Propagates uncertainty through non-linear models where the linear formula fails | our IPD prior × projection × detector chain is non-linear — this is the correct propagation |
3. The rules the council must obey (design, not decoration)
| rule | why |
|---|---|
| Independence over count. Seats are counted by SENSOR, not by metric. Four maths on one detector = one seat. | correlated errors do not cancel; they masquerade as consensus |
| Every seat may ABSTAIN. An instrument that always returns a number is not measuring — it must be able to say “this quantity is unobservable to me” (our frontal landmarks on depth; single-view metrology with no vertical in frame). | a forced vote is fabricated evidence |
| Disagreement is a FINDING, never an average to bury. Report the consensus AND the Birge ratio AND every degree of equivalence. | the 08-03 jaw result: the anchored seat said “fine”, the anchor-free seat said 16.3 mm, the human seat said “gaunt” — 2 of 3 were right and the disagreement WAS the discovery |
| No verdict below the gauge floor. With U(k=2) ≈ ±1.2 mm today, a 0.9 mm “improvement” is noise. | the honest floor on every claim we publish |
| Voting rules have known pathologies (Arrow's theorem): plain majority over incommensurable metrics is not sound. | hence weighted/robust consensus with uncertainties, not a show of hands — the metrology answer to the “7 of 10” instinct |
| The human seat is not outvoted by machines. On presence/aliveness no automated metric exists; the operator's verdict is the reference, and machine consensus against it is a machine failure. | declared limitation, already standing doctrine |
4. Adoption order
| # | seat / act | why now |
|---|---|---|
| C1 | Uncertainty budget per seat + the gauge floor published (extends the 0.59 mm first pass to all contributors) | without it no seat can be scored and no verdict has a floor |
| C2 | Silhouette seat (contour Hausdorff + hull bound) | cheapest genuinely independent sensor; our renders already emit exact silhouettes |
| C3 | Birge ratio + degrees of equivalence in the verdict output | turns “they agree” into a tested claim |
| C4 | Colorimetric seat (ITA°/ΔE00) and texture seat (ISO 25178) | pure math on rendered pixels; retires eyeball colour/roughness verdicts |
| C5 | Identity-embedding seat and perceptual seat (LPIPS/CVVDP class) | two independent learned opinions, one of them perception-calibrated |
| C6 | Dawid–Skene weighting once seats have history | reliability learned from agreement, not assigned |
| C7 | Conformal intervals on the published scoreboard numbers | guaranteed coverage on every “best” claim — the forensic standard |
UNVERIFIED / declared gaps
- May→July 2026 tail unverified (search budget). Most volatile: learned perceptual metric leaders, identity-embedding SOTA, crowd-consensus tooling.
- Seat independence is asserted from architecture, not yet measured. Two seats can share hidden training-data ancestry; the honest check is a correlation matrix of seat errors across many subjects, which requires the corpus intake first.
- Robust-consensus and random-effects estimators disagree when instruments are few and discrepant — with 2–3 seats the consensus is fragile by construction. Report the spread, do not launder it.
- Our gauge floor (±1.2 mm, k=2) is from crop-repeatability alone — render stochasticity, pose residual and the IPD prior are not yet in the budget, so the true floor is WIDER than published today.