Nishi FamilyCompare › Evaluation and Grading (Referee)

Nishi Compare · measured, not asserted

Evaluation and Grading (Referee)

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

Nishi vs OpenSSF Scorecard and openai-evals and EvalPlus and HELM

Layer 1 · Executive

Where we are. Measured 2026-08-19. The referee is the culture the roster answers to -- property-based oracles, liar-killed censuses, coaching that forced real corrections, abstaining safety grades, self-measurement on the public plane -- and it is narrow against the evaluation field: no scored qabench, an orphaned honesty grader, no intervals or paired tests, no standard suites, no contamination checks, no LLM or human-preference judging, no eval-over-time surface, no cost axis.

Where we need to go. Score what we already run, then add rigor (intervals, history), breadth (suites, contamination), judges (LLM and human preference under pre-declared agreement rules) and the supply-chain and cost axes -- so every claim the estate publishes carries an interval, a history and a cost.

The unit. 1 u = one measured session-leg. Calibration from landed rungs: the mangagen panel compositor went from existing substrate to shipped and live-verified in ONE leg (2026-08-13); the citations rung went from 3 to 55 domains in one leg across seven seats (2026-08-18); a greenfield engine with a bite-proven gate has measured 2 to 4 legs. Estimates recalibrate as rungs land and PR7 actuals write back.
Cost to score what we run: 1.5 u. Through M0: the qabench scorer and wiring an existing grader.
Cost to breadth with rigor: 7.5 u. Through M2.
Cost to the full referee: 13.5 u. Everything below.

9 of 22 capabilities measured|0 of them measured exceeds|13 open|coverage 409/1000|adoption 4 full / 5 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883517 domain=referee target_version=0.1 rungs=10 done=0 open=10 finish=0 ranker=nx_dr_ocm

#StageRungPriorityDerivation
#10.1Score qabench (R0) qa_score700v=7 m=1 c=10
#20.1Wire the honesty grader (R1) wc_honesty_gate200v=1 m=1 c=5
#3laterEval history series (R3) evh_series1000v=10 m=1 c=10
#4laterStandardized suites (R4) sb_run750v=15 m=1 c=20
#5laterContamination detection (R5) ct_scan600v=9 m=1 c=15
#6laterPaired statistics on every bench (R2) es_paired466v=7 m=1 c=15
#7laterLLM-as-judge (R6) lj_grade400v=8 m=1 c=20
#8laterCost and latency-aware eval (R9) ec_meter300v=3 m=1 c=10
#9laterHuman preference collection (R7) hp_collect200v=3 m=1 c=15
#10laterRepo supply-chain score (R8) rsc_score133v=2 m=1 c=15

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Score qabench (R0)qa_scoreExact-match plus a rubric score per question on the 251-question bar; the per-question rows are printed and the aggregate is a ratchet floor; the harness that exists finally gradesOrgan1 u
Wire the honesty grader (R1)wc_honesty_gateThe coach consumes nx_honesty_grader on every rung claim; a planted overclaim is refused and a measured claim passes (bite both ways)Organ0.5 u
Paired statistics on every bench (R2)es_pairedConfidence intervals and a paired test over repeated bench runs (the Stabilizer-validated layout-variance bench is the precedent); a difference inside the interval is reported as NOT DISTINGUISHABLE, never as a winOrgan1.5 u
Eval history series (R3)
after R2
evh_seriesEvery gate and bench verdict appended with its epoch to one series and rendered as a regression dashboard on /compare/referee; a regression is a downward step on the series, visible without re-runningOrgan1 u
Standardized suites (R4)
after R0
sb_runMMLU and HumanEval-class suites run from banked datasets through the same harness the modelwright lane names (eh_run); results carry the interval from R2Organ2 u
Contamination detection (R5)
after R4
ct_scann-gram and perplexity-based overlap between training corpus and eval sets printed per suite; a planted leaked item is caught; required before any trained-model score is publishedOrgan1.5 u
LLM-as-judge (R6)
after R2
lj_gradeThe sovereign seat grades text against a rubric; agreement with banked human labels is PRE-DECLARED; nx_dr_semjudge and nx_dr_densejudge are the lexical and dense control arms it must beat to be wiredLocal model2 u
Human preference collection (R7)
after R6
hp_collectPairwise preference collection on the survey engine (anonymous, k-floored) feeding an arena-style rating; ties and abstentions counted, never droppedOrgan1.5 u
Repo supply-chain score (R8)rsc_scoreScorecard-class checks over the estate's own repos and fetched sources: signed releases, pinned deps, branch protection where applicable; every check prints its evidenceOrgan1.5 u
Cost and latency-aware eval (R9)
after R3
ec_meterEvery bench row carries wall time, tokens and joules where measurable; a quality score without its cost is refused at publishOrgan1 u

Milestones

MilestoneRungsCumulative
M0 · Score what we already runR0,R11.5 u
M1 · RigorR2,R34 u
M2 · BreadthR4,R57.5 u
M3 · JudgesR6,R711 u
M4 · Supply chain and costR8,R913.5 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiScorecardopenai-evalsEvalPlusHELM
Property-based independent oracles (never author-supplied answers)Measured: INDEPENDENT oracle exists in runtime/nx_selfimprove_verified_gate.nx, verified at emit. Sort proven ordered AND a permutation, isqrt by bounds, pow by cross-method agreement; a deliberately BUGGY sort is CAUGHT (the verifier has teeth, gate GREEN 2026-07-06 era); EvalPlus test augmentation [evalplus23] is the partial peer [quickcheck00] Adoption: GATE:LIVE trial=GREEN — fully adopted (top of its ladder).
Liar-killed census generation (fabricated cells die)Measured: LIAR-KILL exists in runtime/nx_swcompare_matrix.nx, verified at emit. Every Nishi cell requires the implementing symbol on disk; neg-control + exceed-bounds + row-count checks fail the build -- 27 live comparisons carry it Adoption: LIVE — fully adopted (top of its ladder).
Anti-drift claim coaching (only COACH-OK claims a rung)Measured: wc_count exists in runtime/_hdl_build/nx_ws_coach.nx, verified at emit. U-axes, dated bars, claim-vs-proof ratio, freshness; it forced a real correction on the supervisor census 2026-07-10 (421 to 347 honest) Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
adoption BUILT-UNPROMOTED
Multi-profile industry quality gradingMeasured: nx_grade_profile_elm exists in runtime/nx_quality_grade.nx, verified at emit. Seven industry profiles (Elm, JPL, CERT...) with majority voting; Scorecard checks [scorecard-checks] and HELM multi-metrics [helm22] are partial peers Adoption: LIB-WIRED importers=5 nonval=4 — fully adopted (top of its ladder).
Safety-critical 8-axis grading that ABSTAINSMeasured: _sc_axis_bounded_loop exists in runtime/nx_safety_critical_grade.nx, verified at emit. JPL / IEC61508 / ISO26262 / DO-178C axes; abstains via UNMEASURED instead of guessing -- abstention as a grading verdict is rare in the field Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program.
adoption LIB-GATE-ONLY importers=1
Honesty and overclaim grading of proseMeasured: nx_honesty_arc_report exists in runtime/nx_honesty_grader.nx, verified at emit. Grades claims vs evidence verbs; HONEST FLAG: found orphaned 2026-07-09 (zero importers) -- wiring it to a consumer is a named rung Adoption: LIB-UNIMPORTED — PARTIAL: nothing imports it: wire it into a consumer or retire it.
adoption LIB-UNIMPORTED
QA benchmark harness with real question setsMeasured: qa_split_prose exists in runtime/_hdl_build/nx_qabench.nx, verified at emit. 251-question bar; SCORING is the named wall (2026-07-09); openai-evals registry harness is the bar [openai-evals] Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Census product-surface audit (U-axes enforcement)Measured: ua_find exists in runtime/_hdl_build/nx_census_uiaudit.nx, verified at emit. Flagged 83 of 160 censuses missing U-axes (2026-07-09) -- the instrument that audits the instruments Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Live-derived ecosystem grading (anti-staleness)Measured: em_domain_level exists in runtime/_hdl_build/nx_ecomat_lib.nx, verified at emit. Grades re-derive from evidence files on every read; Scorecard per-commit re-runs [scorecard-checks] are the partial peer Adoption: LIB-WIRED importers=21 nonval=13 — fully adopted (top of its ladder).
Public measured-results publicationOpen — no implementing organ is measured for this axis yet. The /compare hub publishes 27 liar-killed comparisons as REST+MCP; HELM public leaderboards are the bar [helm22]
LLM-as-judge evaluationOpen — watching runtime/nx_llm_judge.nx : lj_grade, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. openai-evals model-graded evals are the bar [openai-evals] [llmjudge23]; our reverse-AI-judge exists for IMAGES (nx_natstat, photoreal domain) but no text judge
watching lj_grade
Standardized benchmark suites (MMLU HumanEval class)Open — watching runtime/nx_std_bench.nx : sb_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. HELM living-benchmark breadth is the bar [helm22]; our benches are domain-organ KATs
watching sb_run
Contamination and memorization detectionOpen — watching runtime/nx_contamination.nx : ct_scan, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Critical for the coming TRAIN-our-own-model arc [contamination23]; EvalPlus [evalplus23] and HELM carry partial answers
watching ct_scan
Supply-chain security scoring of reposOpen — watching runtime/nx_repo_score.nx : rsc_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Scorecard IS this [scorecard-checks]; ties to the janitor census SBOM gap (mom 1864)
watching rsc_score
Statistical rigor (confidence intervals, paired tests)Open — watching runtime/nx_eval_stats.nx : es_paired, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The Stabilizer-validated paired bench [stabilizer13] (2026-05-20, layout variance) set the precedent but no organ generalizes it
watching es_paired
Human-preference evaluationOpen — watching runtime/nx_human_pref.nx : hp_collect, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Arena-style preference collection [chatbotarena24] absent; HELM instruction-following slices partial
watching hp_collect
Longitudinal regression dashboardsOpen — watching runtime/nx_eval_history.nx : evh_series, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Ratchet floors track direction but no eval-over-time surface exists
watching evh_series
Cost and latency-aware evaluationOpen — watching runtime/nx_eval_cost.nx : ec_meter, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. HELM efficiency metrics are the bar [helm22]; our zero-token loops make cost moot internally but externally-facing evals need it
watching ec_meter
Cheat-proof grading BY CONSTRUCTIONOpen — no implementing organ is measured for this axis yet. Held-out scoring + property oracles [quickcheck00] + teeth-tested (the buggy-sort catch) + abstention verdicts: the grader CANNOT be gamed by the graded, in-substrate; the field's harnesses trust dataset answers and model judges
The evaluator grades ITSELF on the public planeOpen — no implementing organ is measured for this axis yet. The maturity rollup grades the ecosystem INCLUDING the grading fleet, the coach gates the coach's own censuses, and this very page is liar-killed by the organ it measures -- self-measurement as culture, not feature
QA bench scoring (the named wall)Open — watching runtime/_hdl_build/nx_qabench.nx : qa_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The 251-question bar has a harness and no scorer; exact-match plus a held-out graded score per question, printed per row
watching qa_score
Honesty grader wired to a consumerOpen — watching runtime/_hdl_build/nx_ws_coach.nx : wc_honesty_gate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. nx_honesty_grader is orphaned (zero importers, 2026-07-09); the coach consuming it turns overclaim grading from a file into a gate
watching wc_honesty_gate
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The Nishi referee is the culture the rest of the roster answers to: graders with TEETH. Capabilities are proven only by INDEPENDENT oracles (defining properties or cross-method agreement -- a deliberately buggy sort MUST be caught), censuses are liar-killed so fabricated cells die at generation, the coach gates every SOTA claim, the safety grader ABSTAINS as UNMEASURED rather than guess, and the whole apparatus publishes its own measured results on the public compare plane it uses to grade everything else. Against the evaluation field it is honestly narrow: no LLM-as-judge, no standardized benchmark suites, no contamination detection, no confidence intervals or paired statistics, no human-preference evaluation, no longitudinal regression dashboards -- HELM and openai-evals own breadth and statistics. The climb: score qabench (scoring is the named wall, bar 251), wire the orphaned honesty grader to a consumer, contamination checks for the coming trained models, then statistical rigor on every bench.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/referee.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain referee, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [scorecard-checks] OpenSSF Scorecard: Check Documentation (docs/checks.md) -- the automated repository security checks (Branch-Protection, CI-Tests, Code-Review, Maintained, SAST, SBOM, Signed-Releases, Vulnerabilities and others). publisher · read in our library knowledge/fetched/cmp_instrument_scorecard-checks.html · pin h0c8e0ca6826693f3b80261f35d8f4430187dd9fd687089dd9eefaf2e9c8979bb · accessed 2026-08-18 · vendor-docGrounds: The Scorecard column and the "Supply-chain security scoring of repos" row (Scorecard IS this, graded 2; Nishi _ABSENT_), plus the partial-peer codes on "Multi-profile industry quality grading" and "Live-derived ecosystem grading" (per-commit re-runs). Mirror shared with instrument.refs.
  2. [openai-evals] OpenAI. evals -- a framework for evaluating LLMs and LLM systems, with an open-source registry of benchmarks and model-graded evals (GitHub repository README). publisher · read in our library knowledge/fetched/cmp_referee_openai-evals.html · pin h4b287d1c163c5b01ad84f0d6e7183118a6a1cf1f134134f422f118e40cb9460d · accessed 2026-08-18 · vendor-docGrounds: The openai-evals column: "QA benchmark harness with real question sets" (openai-evals registry harness is the bar, graded 2) and "LLM-as-judge evaluation" (model-graded evals are the bar, graded 2; Nishi _ABSENT_ for text).
  3. [evalplus23] Liu, Xia, Wang, Zhang. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation (EvalPlus). arXiv:2305.01210, 2023. publisher · read in our library knowledge/fetched/cmp_referee_evalplus23.html · pin h570b56cda244c325ef61ef5e3630513a83c4564c05c9ba4d8a4c40cd4294091e · accessed 2026-08-18 · published-paperGrounds: The EvalPlus column: "Property-based independent oracles (never author-supplied answers)" where EvalPlus test augmentation (HumanEval extended 80x by automated test generation) is the partial peer (3) to our defining-property oracles, and "Contamination and memorization detection" (partial, 3).
  4. [helm22] Liang, Bommasani, Lee et al. Holistic Evaluation of Language Models (HELM). Transactions on Machine Learning Research, 2023 (arXiv:2211.09110). publisher · read in our library knowledge/fetched/cmp_referee_helm22.html · pin h3d95aeb51499159d0652860c626f21c231f4150477148d51dec2dd284870e0d5 · accessed 2026-08-18 · published-paperGrounds: The HELM column: "Standardized benchmark suites (MMLU HumanEval class)" (HELM living-benchmark breadth is the bar), "Public measured-results publication" (public leaderboards, graded 2), "Cost and latency-aware evaluation" (efficiency metrics) and the multi-metric partial peer on "Multi-profile industry quality grading".
  5. [llmjudge23] Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv:2306.05685). publisher · read in our library knowledge/fetched/cmp_instrument_llmjudge23.html · pin h77854867787231a84724d8825cb1754bf14a29b8a93629cf9fc00d76a0734123 · accessed 2026-08-18 · published-paperGrounds: The "LLM-as-judge evaluation" row (Nishi _ABSENT_ for text; the reverse-AI-judge exists only for images): the agreement-with-humans result and the position, verbosity and self-enhancement biases any sovereign text judge must carry as neg-controls. Mirror shared with instrument.refs.
  6. [contamination23] Golchin, Surdeanu. Time Travel in LLMs: Tracing Data Contamination in Large Language Models. ICLR 2024 spotlight (arXiv:2308.08493). publisher · read in our library knowledge/fetched/cmp_referee_contamination23.html · pin h80bbf7ee2b9e6fe89a701f87d1446e0bf70baac8f4597c0e7c63c3bb6e9a10d1 · accessed 2026-08-18 · published-paperGrounds: The "Contamination and memorization detection" row (Nishi _ABSENT_, critical for the TRAIN-our-own-model arc): a published guided-instruction detection method (92-100 percent accuracy over seven datasets) that the coming contamination rung can adopt as its bar.
  7. [stabilizer13] Curtsinger, Berger. STABILIZER: Statistically Sound Performance Evaluation. ASPLOS 2013. publisher · read in our library knowledge/fetched/cmp_referee_stabilizer13.pdf · pin h819c930cc8f51a65a24cdc46452a29ec2c872391724974d21589a8729dac9c49 · accessed 2026-08-18 · published-paperGrounds: The "Statistical rigor (confidence intervals, paired tests)" row: the Stabilizer-validated paired bench of 2026-05-20 (layout variance) that row cites as the precedent is this paper's method -- re-randomized layout so layout effects are Gaussian and ANOVA-testable; no organ yet generalizes it.
  8. [chatbotarena24] Chiang, Zheng, Sheng et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132, 2024. publisher · read in our library knowledge/fetched/cmp_referee_chatbotarena24.html · pin h4f458a9e882f2a76eb50dfb3f0410a385de550f475a90439cb42e6a67c891cf7 · accessed 2026-08-18 · published-paperGrounds: The "Human-preference evaluation" row (Arena-style preference collection absent, Nishi _ABSENT_): the pairwise crowdsourced preference platform with Bradley-Terry ranking that the row's Arena-style bar names.
  9. [quickcheck00] Claessen, Hughes. QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs. ICFP 2000 (author PDF, UPenn course mirror). publisher · read in our library knowledge/fetched/cmp_referee_quickcheck00.pdf · pin he7c8837a1ef48a2acbca01f21b7b12fd636b56f01e2e0f212c973c0dc74b6c46 · accessed 2026-08-18 · published-paperGrounds: The "Property-based independent oracles (never author-supplied answers)" row and the "Cheat-proof grading BY CONSTRUCTION" EXCEED: sort proven ordered AND a permutation, isqrt by bounds, is property-based testing in this paper's sense -- the defining property is the oracle, never an author-supplied expected value.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/referee.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers