Nishi FamilyCompare › Evaluation Instrument

Nishi Compare · measured, not asserted

Evaluation Instrument

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

How Nishi evaluates a program, product, app, capability, ecosystem or suite -- from its git, docs, schematics and papers -- measured against the field's own 2025/26 evaluation tools: OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-bench (LLM-judge/G-Eval/SPDX noted in cells)

Overview

The purpose, declared position and evidence coverage of this domain. Source presence and completed acceptance are different measures.

Where we are. The evaluation instrument is the organ every /compare domain inherits: frontier banking keyless at zero token cost (exceed), field-derived axes by cross-product corroboration, repo layer analysis, Scorecard and CHAOSS-class health from banked evidence, harness discipline with neg-controls and determinism floors, the self-auditing census that flagged 83 false-SOTA signals (exceed), and the coach that refuses any claim without paired evidence (exceed). Named next rungs, carried on the matrix since 2026-07-10: Lighthouse-class product audits wired census-wide, a sovereign LLM-judge for subjective axes, a scheduled refresh cadence, and SPDX-class composition analysis.

Where we need to go. One evaluator, not four: product audits and a sovereign judge folded into the same census that already banks bars and gates claims, refreshed on the beat so no domain page can outlive its evidence, and composition analysis of the suites we evaluate -- with the instrument still auditing itself.

The unit. 1 u = one measured session-leg (estate calibration: graphics R21 in one leg 2026-08-15). Local evidence: R0-R3 and the repo-health gate each landed in one leg on 2026-07-09 and 07-10; estimates are relative to those.
Where we are: 4 open rungs. The banking, gating and self-audit ring present and exceeding; the product-audit, judge, cadence and composition rungs open. Counts measured at emit below this line.
Cost to a living instrument: 1 u. IN1 the refresh cadence -- the freshness gate exists, putting it on the beat is the cheapest rung and the one that makes every other number trustworthy.
Cost to product-grade audits: 2 u. IN2 the ui-judge, contrast and exceed graders wired into every product-surface census.
Cost to a sovereign judge and composition: 4.5 u. IN3 the G-Eval-class judge over the no-float model with a calibration set, IN4 SPDX-class composition of evaluated suites.

Research bar. OpenSSF Scorecard is measured on 18 automated repo checks. Theirs: the security-posture bar. Ours: already matched on the subset; IN1 keeps it fresh.

Research bar. Lighthouse is measured on automated perf, a11y and UX audits. Theirs: the product-audit bar. Ours: IN2.

Research bar. G-Eval (arXiv 2303.16634) is measured on LLM-as-judge with chain-of-thought form filling. Theirs: the subjective-axis bar. Ours: IN3.

Research bar. SPDX is measured on software bill of materials interchange. Theirs: the composition bar. Ours: IN4.

Latest recorded release

No valid dated release entry is recorded for this domain.

Release entries describe recorded changes; they do not establish that every capability passed evaluation.

9 of 13 capabilities measured|3 of them measured exceeds|4 open|coverage 692/1000|adoption 1 full / 8 partial

Evidence profile — what the gaps on this board actually are

Measured by nx_swcompare_evidence, read back by nx_evprofile_lib. Every figure is a count with its denominator — there is deliberately no score, no grade and no percentage anywhere in this band, because a stored scalar is a field a seat can edit and a counted partition is not.

evidence|grounded 9/9|unsupported 0|gates green 7/7|proven able to fail 6/7|never bitten 1|green at 0/0 4|open gaps 4|of them unnamed 0|of them proof withheld 0|flips ready 0

proven able to fail counts the gates that have a RECORDED RED — nx_gate_bite mutated the gate subject, rebuilt it, watched the gate go red, and that record is inside the shared TTL. never bitten is its complement over the same denominator: those gates ran and were green, and nothing has ever shown them able to detect anything, so their green is a statement about this run and not about the gate. green at 0/0 is a separate and much weaker observation — the gate printed GREEN on a zero denominator, so its own tooth counter says it examined nothing. A gate can be green, non-zero, and still never bitten; that is the common case and it is now visible instead of implied.

partition: grounded + unsupported = 9 vs present 9 · named + unnamed + withheld = 4 vs open 4 · both reconcile

liar-kill conj=GPQN · all four conjuncts held

graded document: BUILDROOT tree, 7727 bytes · gates map: PRIMARY · stamped 0d 21h ago · source ../knowledge/status/evstamp_instrument.verdict

Gap classWhat it is, and the work it names
VACUOUS-GATEA gate reported GREEN on a ZERO denominator. A gate with real teeth and a broken counter, and a gate with no teeth at all, are indistinguishable from outside; that indistinguishability is the finding, so it is reported here and never convicted.
This band reports the referee counts. The per-axis worklist rows — which axis is unsupported, which watch is ready to flip, which gap is unnamed — are printed by nx_swcompare_evidence instrument itself and are not carried on the stamp, so this page names the classes and the producer names the rows. That split is stated rather than hidden: a count without a worklist is not actionable, and this band is honest about which half of that it is.

Production map

Follow the dependencies, declared acceptance criteria and recorded priorities. Inspect source binding before treating a rank as executable work.

Ranking source binding: PLAN_MATRIX_BOUND_ONLY. Recorded priorities require current acceptance evidence and resource checks before execution.

Ranking matches the captured plan and matrix only. Latest execution outcome, research freshness, accepted delivery and investment return are unverified.

Recorded priority estimates

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: listed before new work by this heuristic. Priority is not measured delivery cost or execution readiness. Stamp: # asof=1789541578 domain=instrument target_version=0.1 rungs=7 done=2 open=5 finish=2 ranker=nx_dr_ocm

#StageRungPriorityDerivation
FFINISHAnti-drift coach (IR0) COACH-OKBUILT-UNPROMOTEDcompiled, never promoted to the serving root: /api/promote it
FFINISHSelf-audit of every census (IR2) FLAGGEDSOURCE-ONLYsource exists, never compiled: /api/build it
#10.1Scheduled refresh cadence (IN1) wsc_refresh_beat200v=1 m=2 c=10
#2laterProduct audits census-wide (IN2) ev_product_audit666v=2 m=2 c=6
#3laterSovereign model-as-judge (IN3) ev_llm_judge200v=1 m=2 c=10
#4laterComposition analysis (SPDX class) (IN4) ev_sbom_compose200v=1 m=1 c=5
#5laterRepo health (Scorecard and CHAOSS class) (IR1) Scorecard10v=1 m=1 c=100

Declared roadmap — contract, acceptance, executor, effort

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Anti-drift coach (IR0)COACH-OKRefuses any claim without paired evidence -- LANDED gated 3/3Organ0 u
Repo health (Scorecard and CHAOSS class) (IR1)ScorecardCI presence, maintained-ness, contributors, community health from banked evidence -- LANDED gated 7/7Organ0 u
Self-audit of every census (IR2)FLAGGED83 of 162 censuses flagged for false-SOTA signals -- LANDEDOrgan0 u
Scheduled refresh cadence (IN1)
after IR0
wsc_refresh_beatThe coach's freshness gate on the clock plane: every domain's banked bars re-checked on a declared cadence, a stale bar turns the domain's claim AMBER and names the as-of; gate proves a bar aged past the window flips AMBER and a fresh one stays GREENOrgan1 u
Product audits census-wide (IN2)
after IR2
ev_product_auditnx_ui_judge, contrast and exceed graders run on every product-surface census with a Lighthouse-class score per surface; gate proves a surface with a planted contrast failure scores below the bar and the certified nx_wflow_ui console scores at its known 1000 permilOrgan2 u
Sovereign model-as-judge (IN3)
after IR0
ev_llm_judgeThe no-float model grades subjective axes by form-filling against a rubric, CALIBRATED first: agreement with a labelled set published before it may return a verdict, and every verdict carries its rubric; gate proves the judge separates a known-good from a known-bad fixture and abstains when the rubric is missingOrgan3 u
Composition analysis (SPDX class) (IN4)
after IR1
ev_sbom_composeAn SPDX-shaped bill of materials for any evaluated suite derived from its banked repo tree, joined to license tags; gate proves a fixture suite with a known third-party dependency lists it and our own stack lists zero third-party deps (the trivially complete SBOM)Organ1.5 u

Milestones

MilestoneRungsCumulative
M1 · LivingIN11 u
M2 · Product-gradeIN23 u
M3 · Judge and compositionIN3,IN47.5 u
Ladder verdict. NOT DECLARED. This board names no dated best-in-class or frontier target and no rung roles (sotatarget and rungrole rows on its plan); the ranker labels it NO-LADDER until it does, and until then its rungs climb toward a target nobody has written down. Bars: 0 (fresh 0, attested 0, stale 0, unattested 0), verdict NO-BAR against the current month 2026-09.

Inspect a rung and its prerequisites

Declared nodes 7. Rank input binding: PLAN_MATRIX_BOUND_ONLY. Dependency order is authored. Implementation, acceptance evidence, authority and resource readiness are unverified. No action is recommended or dispatched here.

Plan SHA-256 035c74955ce5c52cd4bac15ac6bc1b973c64ed9f09d8b36d45895b94c068ca31. Target rows 0; role rows 0. Existing risks and release worklog retain their own scope; no node completion is inferred.

Use Enter or Space on a rung to inspect its contract. Prerequisite links locate another rung in this list; open its summary to inspect it. Estimates are authored effort, not forecasts.

  1. IR0 — Anti-drift coach

    Prerequisites: None declared; this does not establish execution eligibility.

    Contract: COACH-OK

    Acceptance: Refuses any claim without paired evidence -- LANDED gated 3/3

    Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  2. IR1 — Repo health (Scorecard and CHAOSS class)

    Prerequisites: None declared; this does not establish execution eligibility.

    Contract: Scorecard

    Acceptance: CI presence, maintained-ness, contributors, community health from banked evidence -- LANDED gated 7/7

    Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  3. IR2 — Self-audit of every census

    Prerequisites: None declared; this does not establish execution eligibility.

    Contract: FLAGGED

    Acceptance: 83 of 162 censuses flagged for false-SOTA signals -- LANDED

    Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  4. IN1 — Scheduled refresh cadence

    Prerequisites: IR0 (acceptance unverified)

    Contract: wsc_refresh_beat

    Acceptance: The coach's freshness gate on the clock plane: every domain's banked bars re-checked on a declared cadence, a stale bar turns the domain's claim AMBER and names the as-of; gate proves a bar aged past the window flips AMBER and a fresh one stays GREEN

    Authored effort: 1. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  5. IN2 — Product audits census-wide

    Prerequisites: IR2 (acceptance unverified)

    Contract: ev_product_audit

    Acceptance: nx_ui_judge, contrast and exceed graders run on every product-surface census with a Lighthouse-class score per surface; gate proves a surface with a planted contrast failure scores below the bar and the certified nx_wflow_ui console scores at its known 1000 permil

    Authored effort: 2. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  6. IN3 — Sovereign model-as-judge

    Prerequisites: IR0 (acceptance unverified)

    Contract: ev_llm_judge

    Acceptance: The no-float model grades subjective axes by form-filling against a rubric, CALIBRATED first: agreement with a labelled set published before it may return a verdict, and every verdict carries its rubric; gate proves the judge separates a known-good from a known-bad fixture and abstains when the rubric is missing

    Authored effort: 3. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

  7. IN4 — Composition analysis (SPDX class)

    Prerequisites: IR1 (acceptance unverified)

    Contract: ev_sbom_compose

    Acceptance: An SPDX-shaped bill of materials for any evaluated suite derived from its banked repo tree, joined to license tags; gate proves a fixture suite with a known third-party dependency lists it and our own stack lists zero third-party deps (the trivially complete SBOM)

    Authored effort: 1.5. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.

Learning and practice paths

No structured learning path is declared for this plan. Existing research, roadmap and worklog remain available above.

Capability comparisons

Compare the field, search individual capabilities and open their source and adoption evidence. Documented presence does not establish comparative quality.

Position map — centrality and distinctiveness

The four-quadrant map the field uses for brand strategy (Dawar and Bagga, HBR June 2015), re-derived from this matrix on every publish. Centrality is the share of the category's feature mass a player covers, each feature weighted by how many hold it; distinctiveness is the average lead over each rival on the rows the player holds; breadth is the depth-weighted share of the whole matrix (the bubble); depth is how deeply the rows held are held; momentum is the day-over-day move off the spine (green rising, red falling, grey until day two); the dashed path runs first day → previous day → today. Dividers are the category means. Axes are fitted to the field of play, so read the tick numerals, not the frame. Rival marks are documented presence, so a rival's position reads the record, never its quality. The picture grades its own readability below; the table beside it is the same data for a screen reader or a second method.

Views. 2D cut: centrality, distinctiveness · 3D cube: centrality, distinctiveness, breadth · axes registered: centrality, distinctiveness, breadth, depth, momentum · a board picks its own in knowledge/compare/instrument.cdmap (cut|x|y, cube|x|y|z); absent = the HBR defaults
Position map, two-dimensional cutOne bubble per player. Bubble area is breadth, the ring colour is momentum, the dashed lines are the category means, and both axes are fitted to the field of play with their tick numerals shown. Every value is repeated in the table that follows. 0 200 400 600 800 1000 400 500 600 700 800 900 1000 centrality (permil, fitted 0–1000) distinctiveness (permil, fitted 400–1000) mean 424 mean 633 Unconventional Aspirational Peripheral Mainstream Nishi, centrality 937, distinctiveness 666, breadth 538, depth 777, momentum 0 OpenSSF Scorecard, centrality 375, distinctiveness 500, breadth 128, depth 833, momentum 0 CHAOSS, centrality 375, distinctiveness 500, breadth 128, depth 833, momentum 0 Lighthouse, centrality 62, distinctiveness 1000, breadth 76, depth 1000, momentum 0 SWE-bench, centrality 375, distinctiveness 500, breadth 128, depth 833, momentum 0 Nishi OpenSSF Scorecard CHAOSS Lighthouse SWE-bench

bubble area = breadth · ring = momentum (green rising, red falling, grey until day two) · dashed = category means · axes fitted to the field of play: centrality 0–1000, distinctiveness 400–1000 of 0–1000 permil (the full range put every player in one corner)
readability of the 2D cut, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 3 (a fact of the data: two players that close are that close) · data spread 854 permil of the plot · quadrant words unseated 0
text contrast, measured with wcag2-ratio (floors from contrast.conf), light theme: labels 16.24 (floor 4.50) · notes 16.24 (floor 4.50) · quadrant words 5.89 (floor 4.50) · tick numerals 5.89 (floor 4.50) · axis titles 5.59 (floor 4.50) · dark theme: labels 13.78 (floor 4.50) · quadrant words 5.55 (floor 4.50) · axis titles 5.97 (floor 4.50) · dark classes under their floor 0 (one figure serves both themes: a dark shortfall is a token to fix, never a class to hide) · classes refused under their floor 0 (a refused class is not drawn; the scale classes are measured, never hidden) · export: SVG PNG (receipt, rendered by the estate's own rasteriser from this page)
Position map, three-dimensional cubeThe same players in an isometric cube. Each axis is fitted to its own field of play and carries numerals on the floor grid and the vertical axis; each bubble drops a dotted line to its floor shadow so height reads as height. Values are in the table that follows. 0 200 400 600 800 1000 400 500 600 700 800 900 1000 0 100 200 300 400 500 600 centrality distinctiveness breadth OpenSSF Scorecard: centrality 375, distinctiveness 500, breadth 128 CHAOSS: centrality 375, distinctiveness 500, breadth 128 SWE-bench: centrality 375, distinctiveness 500, breadth 128 Lighthouse: centrality 62, distinctiveness 1000, breadth 76 Nishi: centrality 937, distinctiveness 666, breadth 538 Nishi OpenSSF Scorecard CHAOSS Lighthouse SWE-bench
axes fitted: centrality 0–1000 · distinctiveness 400–1000 · breadth 0–600 permil · farther bubbles are painted first, nearer ones over them · the floor shadow is each bubble's (x, y) at height 0
readability of the 3D cube, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 3 (a fact of the data: two players that close are that close) · data spread 510 permil of the plot
Position map, small multiplesEvery pair of registered axes as one small cut. Each panel fits both axes to the field of play, the dashed lines are the category means, bubble area is breadth and the ring colour is momentum; the legend below names the colours. 0 1000 400 1000 distinctiveness vs centrality Nishi: centrality 937, distinctiveness 666 OpenSSF Scorecard: centrality 375, distinctiveness 500 CHAOSS: centrality 375, distinctiveness 500 Lighthouse: centrality 62, distinctiveness 1000 SWE-bench: centrality 375, distinctiveness 500 0 1000 0 600 breadth vs centrality Nishi: centrality 937, breadth 538 OpenSSF Scorecard: centrality 375, breadth 128 CHAOSS: centrality 375, breadth 128 Lighthouse: centrality 62, breadth 76 SWE-bench: centrality 375, breadth 128 0 1000 750 1000 depth vs centrality Nishi: centrality 937, depth 777 OpenSSF Scorecard: centrality 375, depth 833 CHAOSS: centrality 375, depth 833 Lighthouse: centrality 62, depth 1000 SWE-bench: centrality 375, depth 833 0 1000 480 520 momentum vs centrality Nishi: centrality 937, momentum 500 OpenSSF Scorecard: centrality 375, momentum 500 CHAOSS: centrality 375, momentum 500 Lighthouse: centrality 62, momentum 500 SWE-bench: centrality 375, momentum 500 400 1000 0 600 breadth vs distinctiveness Nishi: distinctiveness 666, breadth 538 OpenSSF Scorecard: distinctiveness 500, breadth 128 CHAOSS: distinctiveness 500, breadth 128 Lighthouse: distinctiveness 1000, breadth 76 SWE-bench: distinctiveness 500, breadth 128 400 1000 750 1000 depth vs distinctiveness Nishi: distinctiveness 666, depth 777 OpenSSF Scorecard: distinctiveness 500, depth 833 CHAOSS: distinctiveness 500, depth 833 Lighthouse: distinctiveness 1000, depth 1000 SWE-bench: distinctiveness 500, depth 833 400 1000 480 520 momentum vs distinctiveness Nishi: distinctiveness 666, momentum 500 OpenSSF Scorecard: distinctiveness 500, momentum 500 CHAOSS: distinctiveness 500, momentum 500 Lighthouse: distinctiveness 1000, momentum 500 SWE-bench: distinctiveness 500, momentum 500 0 600 750 1000 depth vs breadth Nishi: breadth 538, depth 777 OpenSSF Scorecard: breadth 128, depth 833 CHAOSS: breadth 128, depth 833 Lighthouse: breadth 76, depth 1000 SWE-bench: breadth 128, depth 833 0 600 480 520 momentum vs breadth Nishi: breadth 538, momentum 500 OpenSSF Scorecard: breadth 128, momentum 500 CHAOSS: breadth 128, momentum 500 Lighthouse: breadth 76, momentum 500 SWE-bench: breadth 128, momentum 500 750 1000 480 520 momentum vs depth Nishi: depth 777, momentum 500 OpenSSF Scorecard: depth 833, momentum 500 CHAOSS: depth 833, momentum 500 Lighthouse: depth 1000, momentum 500 SWE-bench: depth 833, momentum 500
Nishi OpenSSF Scorecard CHAOSS Lighthouse SWE-bench
10 panels over 5 registered axes, every pair once (the lower axis on x, the higher on y) · each panel fitted to its own field of play, first and last tick numerals shown · no labels in a small cut, the legend names the colours; the grade below reports the mark terms only (MM summed over the panels, spread averaged)
readability of the small multiples, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 30 (a fact of the data: two players that close are that close) · data spread 675 permil of the plot
PlayerQuadrantcentralitydistinctivenessbreadthdepthmomentumRows heldDays on spineFirst seen
NishiAspirational9376665387770 since 2026-09-15922026-09-15
OpenSSF ScorecardPeripheral3755001288330 since 2026-09-15222026-09-15
CHAOSSPeripheral3755001288330 since 2026-09-15222026-09-15
LighthouseUnconventional6210007610000 since 2026-09-15122026-09-15
SWE-benchPeripheral3755001288330 since 2026-09-15222026-09-15
DayMatrix rows reviewedPlayers recorded
2026-09-15135
2026-09-16135

players 5|matrix rows 13|feature mass 16|centrality mean 424|distinctiveness mean 633|axes 5|spine days 2 (shown 2)|rows written today 0|readability defects 2D 0 cube 0|spine knowledge/status/cdmap/instrument.spine

How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. A dark tag means the organ file EXISTS but does not declare the contracted symbol: something shipped there under another name, and until the contract is repointed to the real entry point (the plan rung and this row) or the function is renamed, that capability is invisible to this board — a build lost to darkness, named so it is not. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1789497476, gate census asof 1789498715 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiOpenSSF ScorecardCHAOSSLighthouseSWE-bench
Papers and frontier banking (git/docs/white-papers/research)Measured exceed: rf_fetch_bank in runtime/nx_research_engine.nx, verified at emit. RE-KEYED 2026-08-26 and it was UNGROUNDED, not absent: the row named rf_fetch at runtime/nx_swcompare_research.nx, where that symbol occurs ZERO times, so nx_swcompare_evidence read the row as ungrounded and the domain as RED. The capability is real -- rf_fetch_bank is DECLARED at nx_research_engine.nx line 277 (corpus_complete 1) -- and nx_swcompare_research is a 3,127-byte thin DRIVER that IMPORTS it and declares only sr_read and main. A ROW KEYED TO THE FILE THAT USES A CAPABILITY INSTEAD OF THE FILE THAT DECLARES IT READS AS AN ABSENCE, and the fix is to name the declaration, never to add a stub where the row was pointing. The sibling deepresearch row already keyed this exact capability correctly as nx_research_engine plus rf_fetch_bank, so this is agreement with an existing grounded row rather than a new claim. Nishi banks the live 2025/26 frontier keyless over sovereign TLS at ZERO token cost (0 vs 107 LLM agents/query); none of the eval tools ingest the research frontier -- reuse: nx_swcompare_research + a per-domain .q spec -- ANY workstream banks its field bars this way, zero new code Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
not adopted: compiled, never promoted to the serving root: /api/promote it
Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it
Field-derived comparison axes (not hand-picked)Measured: ad_pop exists in runtime/_hdl_build/nx_axis_discover.nx, verified at emit. axes derived from banked product corpora by cross-product corroboration (782 terms, private vocab excluded); the field hand-authors metric lists -- reuse: nx_axis_discover derives a domains axes from its banked bars; reuse for any new census Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
not adopted: compiled, never promoted to the serving root: /api/promote it
Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it
Repo layer and full-stack structure analysisMeasured: LAYER exists in runtime/_hdl_build/nx_product_bar_gate.nx, verified at emit. Nishi banks the UI/backend/infra/mobile layer fingerprint from the keyless GitHub API; Scorecard/CHAOSS/SWE-bench also read repo trees [scorecard-paper23] Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.
not adopted: compiled, never promoted to the serving root: /api/promote it, then roster it
Pros Nishi has it, measured on disk; ahead of LighthouseCons not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it
Security and quality posture checksMeasured: rh_product exists in runtime/_hdl_build/nx_repo_health_gate.nx, verified at emit. OpenSSF Scorecard leads (18 automated checks) [scorecard-checks] [scorecard-paper23]; Nishi runs the CI-presence and maintained-ness subset from banked evidence, liar-killed (unbanked = NO-EVIDENCE, never healthy) -- reuse: nx_repo_health_gate scores any repo Scorecard/CHAOSS-class from banked GitHub evidence Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.
not adopted: compiled, never promoted to the serving root: /api/promote it, then roster it
Pros Nishi has it, measured on disk; ahead of CHAOSS, Lighthouse, SWE-benchCons behind OpenSSF Scorecard (leads); not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it
Community and ecosystem health metricsMeasured: rh_report exists in runtime/_hdl_build/nx_repo_health_gate.nx, verified at emit. CHAOSS leads [chaoss-metrics]; Nishi reads contributor base and GitHub community health_percentage from banked profiles Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.
not adopted: compiled, never promoted to the serving root: /api/promote it, then roster it
Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, Lighthouse, SWE-benchCons behind CHAOSS (leads); not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it
Benchmark harness discipline (gates, neg-controls, determinism)Measured: INSTRUMENT exists in runtime/_hdl_build/nx_evalsota_census.nx, verified at emit. SWE-bench is the harness-rigor reference [swebench23]; Nishi matches on its own domains -- every rung gated with neg-controls and determinism floors Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
not adopted: source exists, never compiled: /api/build it
Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, LighthouseCons behind SWE-bench (leads); not adopted yet: source exists, never compiled: /api/build it
Meta-evaluation -- the evaluator audits itselfMeasured exceed: FLAGGED in runtime/_hdl_build/nx_census_uiaudit.nx, verified at emit. Nishi scans its own 162 censuses and FLAGS the 83 giving false SOTA signals; the field's tools rarely self-audit -- reuse: nx_census_uiaudit flags any product-surface census missing U-axes; run ecosystem-wide Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
not adopted: source exists, never compiled: /api/build it
Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: source exists, never compiled: /api/build it
Drift and claim/proof gating (a claim needs its evidence)Measured exceed: COACH-OK in runtime/_hdl_build/nx_ws_coach.nx, verified at emit. the coach REFUSES any parity/exceed claim without a paired gate/live-URL token; no eval tool enforces this on its own verdicts -- reuse: nx_ws_coach gates every workstream rung (claim/proof + freshness); run before claiming SOTA Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
not adopted: compiled, never promoted to the serving root: /api/promote it
Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it
Automated product audits (perf, a11y, UX -- Lighthouse-class)DARK — runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_product_audit: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. Lighthouse leads automated product audits [lighthouse-overview]; Nishi has ui-judge/contrast/exceed graders but they are not yet wired census-wide
watching ev_product_audit
Pros none measured yetCons behind Lighthouse (leads); open contract, nothing on disk yet
Model-as-judge for subjective axes (G-Eval class)DARK — runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_llm_judge: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. domain judges are gated; a general sovereign LLM-judge [geval23] [llmjudge23] is buildable now that the no-float model generates faithfully -- named next rung
watching ev_llm_judge
Pros none measured yetCons open contract, nothing on disk yet
Living refresh and scheduled cadenceDARK — runtime/_hdl_build/nx_ws_coach.nx EXISTS but does not declare wsc_refresh_beat: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. the freshness gate exists (current-quarter check); automated re-run cadence on the bg census infra is the next rung
watching wsc_refresh_beat
Pros none measured yetCons open contract, nothing on disk yet
Composition and SBOM analysis (SPDX class)DARK — runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_sbom_compose: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. our own stack is a trivially-complete SBOM (zero third-party deps); SPDX-interop [spdx-spec] composition analysis of evaluated suites is absent
watching ev_sbom_compose
Pros none measured yetCons open contract, nothing on disk yet
Capability claims measured by a declaration ruler on diskMeasured: sd_present exists in runtime/nx_symdecl_lib.nx, verified at emit. ONE ruler decides every Nishi cell on every /compare board: the claimed symbol must be a top-level declaration in the named organ, with the rule chosen from the shape of organ and symbol and printed beside the cell, so a mention in a comment or a call site can no longer carry a capability claim. The field's evaluators grade repositories, community activity, pages and patches; none of them re-derives a product's own capability table from its source on every publish Adoption: LIB-WIRED importers=20 nonval=18 — fully adopted (top of its ladder).
Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-bench; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth)
Rival-claim provenance. Every Yes, Best or Part mark in a rival column is a claim about someone else's product. rows with rival marks 5 · cited 5 · uncited 0. A cited row carries a reference mark that resolves to a pinned mirror (the refs gate measures that); an uncited row is an observation read and is badged in its evidence until a reference lands. The badge is the receipt, never the proof: a mark proves a mirror exists, not that the mirror supports the code.

Delivery and evidence

Inspect risks, technical debt, rendered observations, experiments and references. Read scope and limitations alongside every result.

Risk register

RiskLikelihood x impactMitigation
A model-as-judge that is not calibrated is an opinion with an authoritative namelikely x highIN3 done-rule is agreement on a labelled set BEFORE any verdict; uncalibrated it reports numbers, never verdicts.
Audits wired census-wide hammer the box on the beatpossible x mediumIN2 rides the existing census infrastructure and asks nx_spendgate; one census owner per instrument per day.
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The coverage above is capability presence measured against source — not depth, scale, or polish, where mature rivals may lead. Exceeds are claimed only where a mechanism backs them. Every open gap is a watch contract: it names the organ and symbol that closes it, and this page flips the cell itself when that workstream ships.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/instrument.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain instrument, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

The field — discovered, not chosen

Rows written by nx_field_discover from instrument.seeds: the industry's own lists (Wikipedia wikitext, GitHub topics, awesome lists) read mechanically, every candidate counted across seeds. The matrix columns above are a SEAT'S pick; this band is the population they were picked from, and the stats line measures one against the other. A rival here is a lead, never a verdict — it earns a column when its capabilities are read and pinned.

field|seeds=4|scanned=4|fetched=4|reused=0|failed=0|named=4|candidates=4|mentions=4|capped=0 rival|OpenSSF Scorecard|1|1|col0-openssf|https://github.com/ossf/scorecard/blob/main/docs/checks.md|named rival|CHAOSS|1|1|col1-chaoss|https://chaoss.community/kb-metrics-and-metrics-models/|named rival|Lighthouse|1|1|col2-lighthouse|https://developer.chrome.com/docs/lighthouse/overview|named rival|SWE-bench|1|1|col3-swe|https://arxiv.org/abs/2310.06770|named
RankRivalSeedsMentionsFirst seedKindLink
1OpenSSF Scorecard11col0-openssfnamedhttps://github.com/ossf/scorecard/blob/main/docs/checks.md
2CHAOSS11col1-chaossnamedhttps://chaoss.community/kb-metrics-and-metrics-models/
3Lighthouse11col2-lighthousenamedhttps://developer.chrome.com/docs/lighthouse/overview
4SWE-bench11col3-swenamedhttps://arxiv.org/abs/2310.06770

field candidates 4|shown 4 of 4|matrix columns in the field 4 of 4|discovered rivals with no column 0|malformed rows 0 (counted, never rendered)|read-capped 0

Gaps from the record — what the estate does that no board carries

The record census (nx_goalmap record) reads the invoked-tool population and every plan queue row and files each organ or directive that NO matrix, plan or gates row names. A row here is a callout the boards missed: adjudicate it onto a board or declare it infrastructure. Census state BLIND (age 74432 s), sources read 4 of 7 declared — a BLIND census is a FLOOR: unread sources can only add rows.

kindnameboardsourceevidence

rows shown 0|this board's directives 0|estate-wide un-boarded organs 492 (listed in full on /compare/ecosystem)|census rows 611|malformed 0 (counted, never rendered)

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [scorecard-checks] OpenSSF Scorecard: Check Documentation (docs/checks.md) -- the automated repository security checks (Binary-Artifacts, Branch-Protection, CI-Tests, Code-Review, Maintained, SAST, SBOM, Signed-Releases, Vulnerabilities and others). publisher · read in our library knowledge/fetched/cmp_instrument_scorecard-checks.html · pin h0c8e0ca6826693f3b80261f35d8f4430187dd9fd687089dd9eefaf2e9c8979bb · accessed 2026-08-18 · vendor-docGrounds: The OpenSSF Scorecard column and the "Security and quality posture checks" row (Scorecard leads, graded 2): the check list itself -- the matrix note says 18 automated checks as banked; the live document read 2026-08-18 lists about 20, so the count on the page is a floor, not the current total. Nishi runs the CI-presence and maintained-ness subset from banked evidence.
  2. [scorecard-paper23] Zahan, Kanakiya, Hambleton, Shohan, Williams. OpenSSF Scorecard: On the Path Toward Ecosystem-wide Automated Security Metrics. arXiv:2208.03412 (2022); DOI 10.1109/MSEC.2023.3279773 (2023). publisher · read in our library knowledge/fetched/cmp_instrument_scorecard-paper23.html · pin hc4cbea8d2eb5a9f0cec321a242c5b343975cc989c403977ce2d80e4caa35c7d5 · accessed 2026-08-18 · published-paperGrounds: The Scorecard column at large and the "Repo layer and full-stack structure analysis" row: the peer-reviewed account of ecosystem-wide automated repository scoring that nx_repo_health_gate is measured against (Scorecard/CHAOSS-class, liar-killed).
  3. [chaoss-metrics] CHAOSS (Community Health Analytics in Open Source Software): Metrics and Metrics Models knowledge base -- single-question metrics and multi-metric models of community health. publisher · read in our library knowledge/fetched/cmp_instrument_chaoss-metrics.html · pin hae7133efef9b089a51ac6d34a172ffccf57135b926c88b16a390b0544901e900 · accessed 2026-08-18 · vendor-docGrounds: The CHAOSS column and the "Community and ecosystem health metrics" row (CHAOSS leads, graded 2): the metric definitions the row's contributor-base and community-health reads are graded against.
  4. [lighthouse-overview] Chrome for Developers: Introduction to Lighthouse -- open-source automated audits for performance, accessibility, SEO and best practices of web pages. publisher · read in our library knowledge/fetched/cmp_instrument_lighthouse-overview.html · pin h31972223ccea41cfc07f5b70abde2ca1d3ae4a1cc4a0232929a7e3d433f37fbd · accessed 2026-08-20 · vendor-docGrounds: The Lighthouse column and the "Automated product audits (perf, a11y, UX -- Lighthouse-class)" row (Lighthouse leads, Nishi _ABSENT_): what an automated product audit is in the field; our ui-judge/contrast/exceed graders are not yet wired census-wide.
  5. [swebench23] Jimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024 (arXiv:2310.06770). publisher · read in our library knowledge/fetched/cmp_coding_swebench23.html · pin hcd170cab3a56f891ee94070d5748d923685f2a84a91201d6336a221436fd0f9f · accessed 2026-08-18 · published-paperGrounds: The SWE-bench column and the "Benchmark harness discipline (gates, neg-controls, determinism)" row (SWE-bench is the harness-rigor reference, graded 2): the execution-based, containerized, held-out-test harness design our domain gates match on their own subjects. Mirror shared with coding.refs (same bytes, same pin).
  6. [geval23] Liu, Iter, Xu, Wang, Xu, Zhu. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv:2303.16634, 2023. publisher · read in our library knowledge/fetched/cmp_instrument_geval23.html · pin h663d4662d59ef070020d412d18c10095cd299776cc005a1fa934110273995d48 · accessed 2026-08-18 · published-paperGrounds: The "Model-as-judge for subjective axes (G-Eval class)" row and the NEXT rung "sovereign LLM-judge for subjective axes": the chain-of-thought form-filling judge recipe the row names as buildable now that the no-float model generates faithfully.
  7. [llmjudge23] Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv:2306.05685). publisher · read in our library knowledge/fetched/cmp_instrument_llmjudge23.html · pin h77854867787231a84724d8825cb1754bf14a29b8a93629cf9fc00d76a0734123 · accessed 2026-08-18 · published-paperGrounds: The same "Model-as-judge for subjective axes (G-Eval class)" row: the agreement-with-humans and known-bias analysis (position, verbosity, self-enhancement) any sovereign judge rung must carry as neg-controls before it grades a subjective axis.
  8. [spdx-spec] SPDX (System Package Data Exchange) Specification 3.0.1, The Linux Foundation: the software bill of materials data model and serialization. publisher · read in our library knowledge/fetched/cmp_devguardrails_spdx-spec.html · pin ha905ca11b54700955ba92b4dbdf9d35553852b82bfddf2fabc06cabc194e120a · accessed 2026-08-18 · published-standardGrounds: The "Composition and SBOM analysis (SPDX class)" row (Nishi _ABSENT_): the interchange format an SPDX-interop composition analysis of evaluated suites would emit; our own stack is a trivially-complete SBOM but nothing states it in this shape. Mirror shared with devguardrails.refs.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/instrument.matrix · source checks show implementation presence; runtime and user-outcome evidence are reported separately · JavaScript supports page controls