Nishi Family › Compare › Evaluation and Grading (Referee)
Nishi Compare · measured, not asserted
Evaluation and Grading (Referee)
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs OpenSSF Scorecard and openai-evals and EvalPlus and HELM
Where we are. Measured 2026-08-19. The referee is the culture the roster answers to -- property-based oracles, liar-killed censuses, coaching that forced real corrections, abstaining safety grades, self-measurement on the public plane -- and it is narrow against the evaluation field: no scored qabench, an orphaned honesty grader, no intervals or paired tests, no standard suites, no contamination checks, no LLM or human-preference judging, no eval-over-time surface, no cost axis.
Where we need to go. Score what we already run, then add rigor (intervals, history), breadth (suites, contamination), judges (LLM and human preference under pre-declared agreement rules) and the supply-chain and cost axes -- so every claim the estate publishes carries an interval, a history and a cost.
9 of 22 capabilities measured|0 of them measured exceeds|13 open|coverage 409/1000|adoption 4 full / 5 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883517 domain=referee target_version=0.1 rungs=10 done=0 open=10 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | 0.1 | Score qabench (R0) qa_score | 700 | v=7 m=1 c=10 |
| #2 | 0.1 | Wire the honesty grader (R1) wc_honesty_gate | 200 | v=1 m=1 c=5 |
| #3 | later | Eval history series (R3) evh_series | 1000 | v=10 m=1 c=10 |
| #4 | later | Standardized suites (R4) sb_run | 750 | v=15 m=1 c=20 |
| #5 | later | Contamination detection (R5) ct_scan | 600 | v=9 m=1 c=15 |
| #6 | later | Paired statistics on every bench (R2) es_paired | 466 | v=7 m=1 c=15 |
| #7 | later | LLM-as-judge (R6) lj_grade | 400 | v=8 m=1 c=20 |
| #8 | later | Cost and latency-aware eval (R9) ec_meter | 300 | v=3 m=1 c=10 |
| #9 | later | Human preference collection (R7) hp_collect | 200 | v=3 m=1 c=15 |
| #10 | later | Repo supply-chain score (R8) rsc_score | 133 | v=2 m=1 c=15 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Score qabench (R0) | qa_score | Exact-match plus a rubric score per question on the 251-question bar; the per-question rows are printed and the aggregate is a ratchet floor; the harness that exists finally grades | Organ | 1 u |
| Wire the honesty grader (R1) | wc_honesty_gate | The coach consumes nx_honesty_grader on every rung claim; a planted overclaim is refused and a measured claim passes (bite both ways) | Organ | 0.5 u |
| Paired statistics on every bench (R2) | es_paired | Confidence intervals and a paired test over repeated bench runs (the Stabilizer-validated layout-variance bench is the precedent); a difference inside the interval is reported as NOT DISTINGUISHABLE, never as a win | Organ | 1.5 u |
| Eval history series (R3) after R2 | evh_series | Every gate and bench verdict appended with its epoch to one series and rendered as a regression dashboard on /compare/referee; a regression is a downward step on the series, visible without re-running | Organ | 1 u |
| Standardized suites (R4) after R0 | sb_run | MMLU and HumanEval-class suites run from banked datasets through the same harness the modelwright lane names (eh_run); results carry the interval from R2 | Organ | 2 u |
| Contamination detection (R5) after R4 | ct_scan | n-gram and perplexity-based overlap between training corpus and eval sets printed per suite; a planted leaked item is caught; required before any trained-model score is published | Organ | 1.5 u |
| LLM-as-judge (R6) after R2 | lj_grade | The sovereign seat grades text against a rubric; agreement with banked human labels is PRE-DECLARED; nx_dr_semjudge and nx_dr_densejudge are the lexical and dense control arms it must beat to be wired | Local model | 2 u |
| Human preference collection (R7) after R6 | hp_collect | Pairwise preference collection on the survey engine (anonymous, k-floored) feeding an arena-style rating; ties and abstentions counted, never dropped | Organ | 1.5 u |
| Repo supply-chain score (R8) | rsc_score | Scorecard-class checks over the estate's own repos and fetched sources: signed releases, pinned deps, branch protection where applicable; every check prints its evidence | Organ | 1.5 u |
| Cost and latency-aware eval (R9) after R3 | ec_meter | Every bench row carries wall time, tokens and joules where measurable; a quality score without its cost is refused at publish | Organ | 1 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M0 · Score what we already run | R0,R1 | 1.5 u |
| M1 · Rigor | R2,R3 | 4 u |
| M2 · Breadth | R4,R5 | 7.5 u |
| M3 · Judges | R6,R7 | 11 u |
| M4 · Supply chain and cost | R8,R9 | 13.5 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | Scorecard | openai-evals | EvalPlus | HELM |
|---|---|---|---|---|---|
Property-based independent oracles (never author-supplied answers)Measured:INDEPENDENT oracle exists in runtime/nx_selfimprove_verified_gate.nx, verified at emit. Sort proven ordered AND a permutation, isqrt by bounds, pow by cross-method agreement; a deliberately BUGGY sort is CAUGHT (the verifier has teeth, gate GREEN 2026-07-06 era); EvalPlus test augmentation [evalplus23] is the partial peer [quickcheck00] Adoption: GATE:LIVE trial=GREEN — fully adopted (top of its ladder). | ● | ○ | ◐ | ◐ | ○ |
Liar-killed census generation (fabricated cells die)Measured:LIAR-KILL exists in runtime/nx_swcompare_matrix.nx, verified at emit. Every Nishi cell requires the implementing symbol on disk; neg-control + exceed-bounds + row-count checks fail the build -- 27 live comparisons carry it Adoption: LIVE — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Anti-drift claim coaching (only COACH-OK claims a rung)Measured:wc_count exists in runtime/_hdl_build/nx_ws_coach.nx, verified at emit. U-axes, dated bars, claim-vs-proof ratio, freshness; it forced a real correction on the supervisor census 2026-07-10 (421 to 347 honest) Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ○ | ○ | ○ | ○ |
Multi-profile industry quality gradingMeasured:nx_grade_profile_elm exists in runtime/nx_quality_grade.nx, verified at emit. Seven industry profiles (Elm, JPL, CERT...) with majority voting; Scorecard checks [scorecard-checks] and HELM multi-metrics [helm22] are partial peers Adoption: LIB-WIRED importers=5 nonval=4 — fully adopted (top of its ladder). | ● | ◐ | ○ | ○ | ◐ |
Safety-critical 8-axis grading that ABSTAINSMeasured:_sc_axis_bounded_loop exists in runtime/nx_safety_critical_grade.nx, verified at emit. JPL / IEC61508 / ISO26262 / DO-178C axes; abstains via UNMEASURED instead of guessing -- abstention as a grading verdict is rare in the field Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program. | ● | ○ | ○ | ○ | ○ |
Honesty and overclaim grading of proseMeasured:nx_honesty_arc_report exists in runtime/nx_honesty_grader.nx, verified at emit. Grades claims vs evidence verbs; HONEST FLAG: found orphaned 2026-07-09 (zero importers) -- wiring it to a consumer is a named rung Adoption: LIB-UNIMPORTED — PARTIAL: nothing imports it: wire it into a consumer or retire it. | ● | ○ | ○ | ○ | ○ |
QA benchmark harness with real question setsMeasured:qa_split_prose exists in runtime/_hdl_build/nx_qabench.nx, verified at emit. 251-question bar; SCORING is the named wall (2026-07-09); openai-evals registry harness is the bar [openai-evals] Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ○ | ◉ | ● | ● |
Census product-surface audit (U-axes enforcement)Measured:ua_find exists in runtime/_hdl_build/nx_census_uiaudit.nx, verified at emit. Flagged 83 of 160 censuses missing U-axes (2026-07-09) -- the instrument that audits the instruments Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ○ | ○ | ○ | ○ |
Live-derived ecosystem grading (anti-staleness)Measured:em_domain_level exists in runtime/_hdl_build/nx_ecomat_lib.nx, verified at emit. Grades re-derive from evidence files on every read; Scorecard per-commit re-runs [scorecard-checks] are the partial peer Adoption: LIB-WIRED importers=21 nonval=13 — fully adopted (top of its ladder). | ● | ◐ | ○ | ○ | ○ |
Public measured-results publicationOpen — no implementing organ is measured for this axis yet. The /compare hub publishes 27 liar-killed comparisons as REST+MCP; HELM public leaderboards are the bar [helm22] | ○ | ◐ | ◐ | ◐ | ◉ |
LLM-as-judge evaluationOpen — watchingruntime/nx_llm_judge.nx : lj_grade, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. openai-evals model-graded evals are the bar [openai-evals] [llmjudge23]; our reverse-AI-judge exists for IMAGES (nx_natstat, photoreal domain) but no text judge | ○ | ○ | ◉ | ○ | ◐ |
Standardized benchmark suites (MMLU HumanEval class)Open — watchingruntime/nx_std_bench.nx : sb_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. HELM living-benchmark breadth is the bar [helm22]; our benches are domain-organ KATs | ○ | ○ | ◐ | ◐ | ◉ |
Contamination and memorization detectionOpen — watchingruntime/nx_contamination.nx : ct_scan, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Critical for the coming TRAIN-our-own-model arc [contamination23]; EvalPlus [evalplus23] and HELM carry partial answers | ○ | ○ | ○ | ◐ | ◐ |
Supply-chain security scoring of reposOpen — watchingruntime/nx_repo_score.nx : rsc_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Scorecard IS this [scorecard-checks]; ties to the janitor census SBOM gap (mom 1864) | ○ | ◉ | ○ | ○ | ○ |
Statistical rigor (confidence intervals, paired tests)Open — watchingruntime/nx_eval_stats.nx : es_paired, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The Stabilizer-validated paired bench [stabilizer13] (2026-05-20, layout variance) set the precedent but no organ generalizes it | ○ | ○ | ○ | ○ | ◐ |
Human-preference evaluationOpen — watchingruntime/nx_human_pref.nx : hp_collect, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Arena-style preference collection [chatbotarena24] absent; HELM instruction-following slices partial | ○ | ○ | ○ | ○ | ◐ |
Longitudinal regression dashboardsOpen — watchingruntime/nx_eval_history.nx : evh_series, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Ratchet floors track direction but no eval-over-time surface exists | ○ | ◐ | ○ | ○ | ◐ |
Cost and latency-aware evaluationOpen — watchingruntime/nx_eval_cost.nx : ec_meter, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. HELM efficiency metrics are the bar [helm22]; our zero-token loops make cost moot internally but externally-facing evals need it | ○ | ○ | ○ | ○ | ◐ |
Cheat-proof grading BY CONSTRUCTIONOpen — no implementing organ is measured for this axis yet. Held-out scoring + property oracles [quickcheck00] + teeth-tested (the buggy-sort catch) + abstention verdicts: the grader CANNOT be gamed by the graded, in-substrate; the field's harnesses trust dataset answers and model judges | ○ | ○ | ○ | ○ | ○ |
The evaluator grades ITSELF on the public planeOpen — no implementing organ is measured for this axis yet. The maturity rollup grades the ecosystem INCLUDING the grading fleet, the coach gates the coach's own censuses, and this very page is liar-killed by the organ it measures -- self-measurement as culture, not feature | ○ | ○ | ○ | ○ | ○ |
QA bench scoring (the named wall)Open — watchingruntime/_hdl_build/nx_qabench.nx : qa_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The 251-question bar has a harness and no scorer; exact-match plus a held-out graded score per question, printed per row | ○ | ○ | ◉ | ● | ● |
Honesty grader wired to a consumerOpen — watchingruntime/_hdl_build/nx_ws_coach.nx : wc_honesty_gate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. nx_honesty_grader is orphaned (zero importers, 2026-07-09); the coach consuming it turns overclaim grading from a file into a gate | ○ | ○ | ○ | ○ | ○ |
Person · product · place — not yet measured for this domain
knowledge/compare/referee.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain referee, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [scorecard-checks] OpenSSF Scorecard: Check Documentation (docs/checks.md) -- the automated repository security checks (Branch-Protection, CI-Tests, Code-Review, Maintained, SAST, SBOM, Signed-Releases, Vulnerabilities and others). publisher · read in our library
knowledge/fetched/cmp_instrument_scorecard-checks.html· pinh0c8e0ca6826693f3b80261f35d8f4430187dd9fd687089dd9eefaf2e9c8979bb· accessed 2026-08-18 · vendor-docGrounds: The Scorecard column and the "Supply-chain security scoring of repos" row (Scorecard IS this, graded 2; Nishi _ABSENT_), plus the partial-peer codes on "Multi-profile industry quality grading" and "Live-derived ecosystem grading" (per-commit re-runs). Mirror shared with instrument.refs. - [openai-evals] OpenAI. evals -- a framework for evaluating LLMs and LLM systems, with an open-source registry of benchmarks and model-graded evals (GitHub repository README). publisher · read in our library
knowledge/fetched/cmp_referee_openai-evals.html· pinh4b287d1c163c5b01ad84f0d6e7183118a6a1cf1f134134f422f118e40cb9460d· accessed 2026-08-18 · vendor-docGrounds: The openai-evals column: "QA benchmark harness with real question sets" (openai-evals registry harness is the bar, graded 2) and "LLM-as-judge evaluation" (model-graded evals are the bar, graded 2; Nishi _ABSENT_ for text). - [evalplus23] Liu, Xia, Wang, Zhang. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation (EvalPlus). arXiv:2305.01210, 2023. publisher · read in our library
knowledge/fetched/cmp_referee_evalplus23.html· pinh570b56cda244c325ef61ef5e3630513a83c4564c05c9ba4d8a4c40cd4294091e· accessed 2026-08-18 · published-paperGrounds: The EvalPlus column: "Property-based independent oracles (never author-supplied answers)" where EvalPlus test augmentation (HumanEval extended 80x by automated test generation) is the partial peer (3) to our defining-property oracles, and "Contamination and memorization detection" (partial, 3). - [helm22] Liang, Bommasani, Lee et al. Holistic Evaluation of Language Models (HELM). Transactions on Machine Learning Research, 2023 (arXiv:2211.09110). publisher · read in our library
knowledge/fetched/cmp_referee_helm22.html· pinh3d95aeb51499159d0652860c626f21c231f4150477148d51dec2dd284870e0d5· accessed 2026-08-18 · published-paperGrounds: The HELM column: "Standardized benchmark suites (MMLU HumanEval class)" (HELM living-benchmark breadth is the bar), "Public measured-results publication" (public leaderboards, graded 2), "Cost and latency-aware evaluation" (efficiency metrics) and the multi-metric partial peer on "Multi-profile industry quality grading". - [llmjudge23] Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv:2306.05685). publisher · read in our library
knowledge/fetched/cmp_instrument_llmjudge23.html· pinh77854867787231a84724d8825cb1754bf14a29b8a93629cf9fc00d76a0734123· accessed 2026-08-18 · published-paperGrounds: The "LLM-as-judge evaluation" row (Nishi _ABSENT_ for text; the reverse-AI-judge exists only for images): the agreement-with-humans result and the position, verbosity and self-enhancement biases any sovereign text judge must carry as neg-controls. Mirror shared with instrument.refs. - [contamination23] Golchin, Surdeanu. Time Travel in LLMs: Tracing Data Contamination in Large Language Models. ICLR 2024 spotlight (arXiv:2308.08493). publisher · read in our library
knowledge/fetched/cmp_referee_contamination23.html· pinh80bbf7ee2b9e6fe89a701f87d1446e0bf70baac8f4597c0e7c63c3bb6e9a10d1· accessed 2026-08-18 · published-paperGrounds: The "Contamination and memorization detection" row (Nishi _ABSENT_, critical for the TRAIN-our-own-model arc): a published guided-instruction detection method (92-100 percent accuracy over seven datasets) that the coming contamination rung can adopt as its bar. - [stabilizer13] Curtsinger, Berger. STABILIZER: Statistically Sound Performance Evaluation. ASPLOS 2013. publisher · read in our library
knowledge/fetched/cmp_referee_stabilizer13.pdf· pinh819c930cc8f51a65a24cdc46452a29ec2c872391724974d21589a8729dac9c49· accessed 2026-08-18 · published-paperGrounds: The "Statistical rigor (confidence intervals, paired tests)" row: the Stabilizer-validated paired bench of 2026-05-20 (layout variance) that row cites as the precedent is this paper's method -- re-randomized layout so layout effects are Gaussian and ANOVA-testable; no organ yet generalizes it. - [chatbotarena24] Chiang, Zheng, Sheng et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132, 2024. publisher · read in our library
knowledge/fetched/cmp_referee_chatbotarena24.html· pinh4f458a9e882f2a76eb50dfb3f0410a385de550f475a90439cb42e6a67c891cf7· accessed 2026-08-18 · published-paperGrounds: The "Human-preference evaluation" row (Arena-style preference collection absent, Nishi _ABSENT_): the pairwise crowdsourced preference platform with Bradley-Terry ranking that the row's Arena-style bar names. - [quickcheck00] Claessen, Hughes. QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs. ICFP 2000 (author PDF, UPenn course mirror). publisher · read in our library
knowledge/fetched/cmp_referee_quickcheck00.pdf· pinhe7c8837a1ef48a2acbca01f21b7b12fd636b56f01e2e0f212c973c0dc74b6c46· accessed 2026-08-18 · published-paperGrounds: The "Property-based independent oracles (never author-supplied answers)" row and the "Cheat-proof grading BY CONSTRUCTION" EXCEED: sort proven ordered AND a permutation, isqrt by bounds, is property-based testing in this paper's sense -- the defining property is the oracle, never an author-supplied expected value.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/referee.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers