Nishi Family › Compare › Evaluation Instrument
Nishi Compare · measured, not asserted
Evaluation Instrument
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
How Nishi evaluates a program, product, app, capability, ecosystem or suite -- from its git, docs, schematics and papers -- measured against the field's own 2025/26 evaluation tools: OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-bench (LLM-judge/G-Eval/SPDX noted in cells)
Overview
The purpose, declared position and evidence coverage of this domain. Source presence and completed acceptance are different measures.
Where we are. The evaluation instrument is the organ every /compare domain inherits: frontier banking keyless at zero token cost (exceed), field-derived axes by cross-product corroboration, repo layer analysis, Scorecard and CHAOSS-class health from banked evidence, harness discipline with neg-controls and determinism floors, the self-auditing census that flagged 83 false-SOTA signals (exceed), and the coach that refuses any claim without paired evidence (exceed). Named next rungs, carried on the matrix since 2026-07-10: Lighthouse-class product audits wired census-wide, a sovereign LLM-judge for subjective axes, a scheduled refresh cadence, and SPDX-class composition analysis.
Where we need to go. One evaluator, not four: product audits and a sovereign judge folded into the same census that already banks bars and gates claims, refreshed on the beat so no domain page can outlive its evidence, and composition analysis of the suites we evaluate -- with the instrument still auditing itself.
Research bar. OpenSSF Scorecard is measured on 18 automated repo checks. Theirs: the security-posture bar. Ours: already matched on the subset; IN1 keeps it fresh.
Research bar. Lighthouse is measured on automated perf, a11y and UX audits. Theirs: the product-audit bar. Ours: IN2.
Research bar. G-Eval (arXiv 2303.16634) is measured on LLM-as-judge with chain-of-thought form filling. Theirs: the subjective-axis bar. Ours: IN3.
Research bar. SPDX is measured on software bill of materials interchange. Theirs: the composition bar. Ours: IN4.
Latest recorded release
No valid dated release entry is recorded for this domain.
Release entries describe recorded changes; they do not establish that every capability passed evaluation.
9 of 13 capabilities measured|3 of them measured exceeds|4 open|coverage 692/1000|adoption 1 full / 8 partial
Evidence profile — what the gaps on this board actually are
Measured by nx_swcompare_evidence, read back by nx_evprofile_lib. Every figure is a count with its denominator — there is deliberately no score, no grade and no percentage anywhere in this band, because a stored scalar is a field a seat can edit and a counted partition is not.
evidence|grounded 9/9|unsupported 0|gates green 7/7|proven able to fail 6/7|never bitten 1|green at 0/0 4|open gaps 4|of them unnamed 0|of them proof withheld 0|flips ready 0
proven able to fail counts the gates that have a RECORDED RED — nx_gate_bite mutated the gate subject, rebuilt it, watched the gate go red, and that record is inside the shared TTL. never bitten is its complement over the same denominator: those gates ran and were green, and nothing has ever shown them able to detect anything, so their green is a statement about this run and not about the gate. green at 0/0 is a separate and much weaker observation — the gate printed GREEN on a zero denominator, so its own tooth counter says it examined nothing. A gate can be green, non-zero, and still never bitten; that is the common case and it is now visible instead of implied.
partition: grounded + unsupported = 9 vs present 9 · named + unnamed + withheld = 4 vs open 4 · both reconcile
liar-kill conj=GPQN · all four conjuncts held
graded document: BUILDROOT tree, 7727 bytes · gates map: PRIMARY · stamped 0d 21h ago · source ../knowledge/status/evstamp_instrument.verdict
| Gap class | What it is, and the work it names |
|---|---|
| VACUOUS-GATE | A gate reported GREEN on a ZERO denominator. A gate with real teeth and a broken counter, and a gate with no teeth at all, are indistinguishable from outside; that indistinguishability is the finding, so it is reported here and never convicted. |
nx_swcompare_evidence instrument itself and are not carried on the stamp, so this page names the classes and the producer names the rows. That split is stated rather than hidden: a count without a worklist is not actionable, and this band is honest about which half of that it is.Production map
Follow the dependencies, declared acceptance criteria and recorded priorities. Inspect source binding before treating a rank as executable work.
Ranking source binding: PLAN_MATRIX_BOUND_ONLY. Recorded priorities require current acceptance evidence and resource checks before execution.
Ranking matches the captured plan and matrix only. Latest execution outcome, research freshness, accepted delivery and investment return are unverified.
Recorded priority estimates
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: listed before new work by this heuristic. Priority is not measured delivery cost or execution readiness. Stamp: # asof=1789541578 domain=instrument target_version=0.1 rungs=7 done=2 open=5 finish=2 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| F | FINISH | Anti-drift coach (IR0) COACH-OK | BUILT-UNPROMOTED | compiled, never promoted to the serving root: /api/promote it |
| F | FINISH | Self-audit of every census (IR2) FLAGGED | SOURCE-ONLY | source exists, never compiled: /api/build it |
| #1 | 0.1 | Scheduled refresh cadence (IN1) wsc_refresh_beat | 200 | v=1 m=2 c=10 |
| #2 | later | Product audits census-wide (IN2) ev_product_audit | 666 | v=2 m=2 c=6 |
| #3 | later | Sovereign model-as-judge (IN3) ev_llm_judge | 200 | v=1 m=2 c=10 |
| #4 | later | Composition analysis (SPDX class) (IN4) ev_sbom_compose | 200 | v=1 m=1 c=5 |
| #5 | later | Repo health (Scorecard and CHAOSS class) (IR1) Scorecard | 10 | v=1 m=1 c=100 |
Declared roadmap — contract, acceptance, executor, effort
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Anti-drift coach (IR0) | COACH-OK | Refuses any claim without paired evidence -- LANDED gated 3/3 | Organ | 0 u |
| Repo health (Scorecard and CHAOSS class) (IR1) | Scorecard | CI presence, maintained-ness, contributors, community health from banked evidence -- LANDED gated 7/7 | Organ | 0 u |
| Self-audit of every census (IR2) | FLAGGED | 83 of 162 censuses flagged for false-SOTA signals -- LANDED | Organ | 0 u |
| Scheduled refresh cadence (IN1) after IR0 | wsc_refresh_beat | The coach's freshness gate on the clock plane: every domain's banked bars re-checked on a declared cadence, a stale bar turns the domain's claim AMBER and names the as-of; gate proves a bar aged past the window flips AMBER and a fresh one stays GREEN | Organ | 1 u |
| Product audits census-wide (IN2) after IR2 | ev_product_audit | nx_ui_judge, contrast and exceed graders run on every product-surface census with a Lighthouse-class score per surface; gate proves a surface with a planted contrast failure scores below the bar and the certified nx_wflow_ui console scores at its known 1000 permil | Organ | 2 u |
| Sovereign model-as-judge (IN3) after IR0 | ev_llm_judge | The no-float model grades subjective axes by form-filling against a rubric, CALIBRATED first: agreement with a labelled set published before it may return a verdict, and every verdict carries its rubric; gate proves the judge separates a known-good from a known-bad fixture and abstains when the rubric is missing | Organ | 3 u |
| Composition analysis (SPDX class) (IN4) after IR1 | ev_sbom_compose | An SPDX-shaped bill of materials for any evaluated suite derived from its banked repo tree, joined to license tags; gate proves a fixture suite with a known third-party dependency lists it and our own stack lists zero third-party deps (the trivially complete SBOM) | Organ | 1.5 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M1 · Living | IN1 | 1 u |
| M2 · Product-grade | IN2 | 3 u |
| M3 · Judge and composition | IN3,IN4 | 7.5 u |
Inspect a rung and its prerequisites
Declared nodes 7. Rank input binding: PLAN_MATRIX_BOUND_ONLY. Dependency order is authored. Implementation, acceptance evidence, authority and resource readiness are unverified. No action is recommended or dispatched here.
Plan SHA-256 035c74955ce5c52cd4bac15ac6bc1b973c64ed9f09d8b36d45895b94c068ca31. Target rows 0; role rows 0. Existing risks and release worklog retain their own scope; no node completion is inferred.
Use Enter or Space on a rung to inspect its contract. Prerequisite links locate another rung in this list; open its summary to inspect it. Estimates are authored effort, not forecasts.
IR0 — Anti-drift coach
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
COACH-OKAcceptance: Refuses any claim without paired evidence -- LANDED gated 3/3
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IR1 — Repo health (Scorecard and CHAOSS class)
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
ScorecardAcceptance: CI presence, maintained-ness, contributors, community health from banked evidence -- LANDED gated 7/7
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IR2 — Self-audit of every census
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
FLAGGEDAcceptance: 83 of 162 censuses flagged for false-SOTA signals -- LANDED
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IN1 — Scheduled refresh cadence
Prerequisites: IR0 (acceptance unverified)
Contract:
wsc_refresh_beatAcceptance: The coach's freshness gate on the clock plane: every domain's banked bars re-checked on a declared cadence, a stale bar turns the domain's claim AMBER and names the as-of; gate proves a bar aged past the window flips AMBER and a fresh one stays GREEN
Authored effort: 1. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IN2 — Product audits census-wide
Prerequisites: IR2 (acceptance unverified)
Contract:
ev_product_auditAcceptance: nx_ui_judge, contrast and exceed graders run on every product-surface census with a Lighthouse-class score per surface; gate proves a surface with a planted contrast failure scores below the bar and the certified nx_wflow_ui console scores at its known 1000 permil
Authored effort: 2. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IN3 — Sovereign model-as-judge
Prerequisites: IR0 (acceptance unverified)
Contract:
ev_llm_judgeAcceptance: The no-float model grades subjective axes by form-filling against a rubric, CALIBRATED first: agreement with a labelled set published before it may return a verdict, and every verdict carries its rubric; gate proves the judge separates a known-good from a known-bad fixture and abstains when the rubric is missing
Authored effort: 3. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
IN4 — Composition analysis (SPDX class)
Prerequisites: IR1 (acceptance unverified)
Contract:
ev_sbom_composeAcceptance: An SPDX-shaped bill of materials for any evaluated suite derived from its banked repo tree, joined to license tags; gate proves a fixture suite with a known third-party dependency lists it and our own stack lists zero third-party deps (the trivially complete SBOM)
Authored effort: 1.5. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
Learning and practice paths
No structured learning path is declared for this plan. Existing research, roadmap and worklog remain available above.
Capability comparisons
Compare the field, search individual capabilities and open their source and adoption evidence. Documented presence does not establish comparative quality.
Position map — centrality and distinctiveness
The four-quadrant map the field uses for brand strategy (Dawar and Bagga, HBR June 2015), re-derived from this matrix on every publish. Centrality is the share of the category's feature mass a player covers, each feature weighted by how many hold it; distinctiveness is the average lead over each rival on the rows the player holds; breadth is the depth-weighted share of the whole matrix (the bubble); depth is how deeply the rows held are held; momentum is the day-over-day move off the spine (green rising, red falling, grey until day two); the dashed path runs first day → previous day → today. Dividers are the category means. Axes are fitted to the field of play, so read the tick numerals, not the frame. Rival marks are documented presence, so a rival's position reads the record, never its quality. The picture grades its own readability below; the table beside it is the same data for a screen reader or a second method.
knowledge/compare/instrument.cdmap (cut|x|y, cube|x|y|z); absent = the HBR defaultsbubble area = breadth · ring = momentum (green rising, red falling, grey until day two) · dashed = category means · axes fitted to the field of play: centrality 0–1000, distinctiveness 400–1000 of 0–1000 permil (the full range put every player in one corner)
readability of the 2D cut, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 3 (a fact of the data: two players that close are that close) · data spread 854 permil of the plot · quadrant words unseated 0
text contrast, measured with wcag2-ratio (floors from contrast.conf), light theme: labels 16.24 (floor 4.50) · notes 16.24 (floor 4.50) · quadrant words 5.89 (floor 4.50) · tick numerals 5.89 (floor 4.50) · axis titles 5.59 (floor 4.50) · dark theme: labels 13.78 (floor 4.50) · quadrant words 5.55 (floor 4.50) · axis titles 5.97 (floor 4.50) · dark classes under their floor 0 (one figure serves both themes: a dark shortfall is a token to fix, never a class to hide) · classes refused under their floor 0 (a refused class is not drawn; the scale classes are measured, never hidden) · export: SVG PNG (receipt, rendered by the estate's own rasteriser from this page)
readability of the 3D cube, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 3 (a fact of the data: two players that close are that close) · data spread 510 permil of the plot
10 panels over 5 registered axes, every pair once (the lower axis on x, the higher on y) · each panel fitted to its own field of play, first and last tick numerals shown · no labels in a small cut, the legend names the colours; the grade below reports the mark terms only (MM summed over the panels, spread averaged)
readability of the small multiples, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 30 (a fact of the data: two players that close are that close) · data spread 675 permil of the plot
| Player | Quadrant | centrality | distinctiveness | breadth | depth | momentum | Rows held | Days on spine | First seen |
|---|---|---|---|---|---|---|---|---|---|
| Nishi | Aspirational | 937 | 666 | 538 | 777 | 0 since 2026-09-15 | 9 | 2 | 2026-09-15 |
| OpenSSF Scorecard | Peripheral | 375 | 500 | 128 | 833 | 0 since 2026-09-15 | 2 | 2 | 2026-09-15 |
| CHAOSS | Peripheral | 375 | 500 | 128 | 833 | 0 since 2026-09-15 | 2 | 2 | 2026-09-15 |
| Lighthouse | Unconventional | 62 | 1000 | 76 | 1000 | 0 since 2026-09-15 | 1 | 2 | 2026-09-15 |
| SWE-bench | Peripheral | 375 | 500 | 128 | 833 | 0 since 2026-09-15 | 2 | 2 | 2026-09-15 |
| Day | Matrix rows reviewed | Players recorded |
|---|---|---|
| 2026-09-15 | 13 | 5 |
| 2026-09-16 | 13 | 5 |
players 5|matrix rows 13|feature mass 16|centrality mean 424|distinctiveness mean 633|axes 5|spine days 2 (shown 2)|rows written today 0|readability defects 2D 0 cube 0|spine knowledge/status/cdmap/instrument.spine
comparewatch- plane row flips with it. A dark tag means the organ file EXISTS but does not declare the contracted symbol: something shipped there under another name, and until the contract is repointed to the real entry point (the plan rung and this row) or the function is renamed, that capability is invisible to this board — a build lost to darkness, named so it is not. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1789497476, gate census asof 1789498715 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | OpenSSF Scorecard | CHAOSS | Lighthouse | SWE-bench |
|---|---|---|---|---|---|
Papers and frontier banking (git/docs/white-papers/research)Measured exceed:rf_fetch_bank in runtime/nx_research_engine.nx, verified at emit. RE-KEYED 2026-08-26 and it was UNGROUNDED, not absent: the row named rf_fetch at runtime/nx_swcompare_research.nx, where that symbol occurs ZERO times, so nx_swcompare_evidence read the row as ungrounded and the domain as RED. The capability is real -- rf_fetch_bank is DECLARED at nx_research_engine.nx line 277 (corpus_complete 1) -- and nx_swcompare_research is a 3,127-byte thin DRIVER that IMPORTS it and declares only sr_read and main. A ROW KEYED TO THE FILE THAT USES A CAPABILITY INSTEAD OF THE FILE THAT DECLARES IT READS AS AN ABSENCE, and the fix is to name the declaration, never to add a stub where the row was pointing. The sibling deepresearch row already keyed this exact capability correctly as nx_research_engine plus rf_fetch_bank, so this is agreement with an existing grounded row rather than a new claim. Nishi banks the live 2025/26 frontier keyless over sovereign TLS at ZERO token cost (0 vs 107 LLM agents/query); none of the eval tools ingest the research frontier -- reuse: nx_swcompare_research + a per-domain .q spec -- ANY workstream banks its field bars this way, zero new code Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it | ◉ | ○ | ○ | ○ | ○ |
Field-derived comparison axes (not hand-picked)Measured:ad_pop exists in runtime/_hdl_build/nx_axis_discover.nx, verified at emit. axes derived from banked product corpora by cross-product corroboration (782 terms, private vocab excluded); the field hand-authors metric lists -- reuse: nx_axis_discover derives a domains axes from its banked bars; reuse for any new census Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it | ● | ○ | ○ | ○ | ○ |
Repo layer and full-stack structure analysisMeasured:LAYER exists in runtime/_hdl_build/nx_product_bar_gate.nx, verified at emit. Nishi banks the UI/backend/infra/mobile layer fingerprint from the keyless GitHub API; Scorecard/CHAOSS/SWE-bench also read repo trees [scorecard-paper23] Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.Pros Nishi has it, measured on disk; ahead of LighthouseCons not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it | ● | ● | ● | ○ | ● |
Security and quality posture checksMeasured:rh_product exists in runtime/_hdl_build/nx_repo_health_gate.nx, verified at emit. OpenSSF Scorecard leads (18 automated checks) [scorecard-checks] [scorecard-paper23]; Nishi runs the CI-presence and maintained-ness subset from banked evidence, liar-killed (unbanked = NO-EVIDENCE, never healthy) -- reuse: nx_repo_health_gate scores any repo Scorecard/CHAOSS-class from banked GitHub evidence Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.Pros Nishi has it, measured on disk; ahead of CHAOSS, Lighthouse, SWE-benchCons behind OpenSSF Scorecard (leads); not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it | ● | ◉ | ○ | ○ | ○ |
Community and ecosystem health metricsMeasured:rh_report exists in runtime/_hdl_build/nx_repo_health_gate.nx, verified at emit. CHAOSS leads [chaoss-metrics]; Nishi reads contributor base and GitHub community health_percentage from banked profiles Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, Lighthouse, SWE-benchCons behind CHAOSS (leads); not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it | ● | ○ | ◉ | ○ | ○ |
Benchmark harness discipline (gates, neg-controls, determinism)Measured:INSTRUMENT exists in runtime/_hdl_build/nx_evalsota_census.nx, verified at emit. SWE-bench is the harness-rigor reference [swebench23]; Nishi matches on its own domains -- every rung gated with neg-controls and determinism floors Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, LighthouseCons behind SWE-bench (leads); not adopted yet: source exists, never compiled: /api/build it | ● | ○ | ○ | ○ | ◉ |
Meta-evaluation -- the evaluator audits itselfMeasured exceed:FLAGGED in runtime/_hdl_build/nx_census_uiaudit.nx, verified at emit. Nishi scans its own 162 censuses and FLAGS the 83 giving false SOTA signals; the field's tools rarely self-audit -- reuse: nx_census_uiaudit flags any product-surface census missing U-axes; run ecosystem-wide Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: source exists, never compiled: /api/build it | ◉ | ○ | ○ | ○ | ○ |
Drift and claim/proof gating (a claim needs its evidence)Measured exceed:COACH-OK in runtime/_hdl_build/nx_ws_coach.nx, verified at emit. the coach REFUSES any parity/exceed claim without a paired gate/live-URL token; no eval tool enforces this on its own verdicts -- reuse: nx_ws_coach gates every workstream rung (claim/proof + freshness); run before claiming SOTA Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.Pros Nishi leads, a measured exceed; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-benchCons not adopted yet: compiled, never promoted to the serving root: /api/promote it | ◉ | ○ | ○ | ○ | ○ |
Automated product audits (perf, a11y, UX -- Lighthouse-class)DARK —runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_product_audit: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. Lighthouse leads automated product audits [lighthouse-overview]; Nishi has ui-judge/contrast/exceed graders but they are not yet wired census-widePros none measured yetCons behind Lighthouse (leads); open contract, nothing on disk yet | ○ | ○ | ○ | ◉ | ○ |
Model-as-judge for subjective axes (G-Eval class)DARK —runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_llm_judge: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. domain judges are gated; a general sovereign LLM-judge [geval23] [llmjudge23] is buildable now that the no-float model generates faithfully -- named next rungPros none measured yetCons open contract, nothing on disk yet | ○ | ○ | ○ | ○ | ○ |
Living refresh and scheduled cadenceDARK —runtime/_hdl_build/nx_ws_coach.nx EXISTS but does not declare wsc_refresh_beat: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. the freshness gate exists (current-quarter check); automated re-run cadence on the bg census infra is the next rungPros none measured yetCons open contract, nothing on disk yet | ○ | ○ | ○ | ○ | ○ |
Composition and SBOM analysis (SPDX class)DARK —runtime/_hdl_build/nx_evalsota_census.nx EXISTS but does not declare ev_sbom_compose: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. our own stack is a trivially-complete SBOM (zero third-party deps); SPDX-interop [spdx-spec] composition analysis of evaluated suites is absentPros none measured yetCons open contract, nothing on disk yet | ○ | ○ | ○ | ○ | ○ |
Capability claims measured by a declaration ruler on diskMeasured:sd_present exists in runtime/nx_symdecl_lib.nx, verified at emit. ONE ruler decides every Nishi cell on every /compare board: the claimed symbol must be a top-level declaration in the named organ, with the rule chosen from the shape of organ and symbol and printed beside the cell, so a mention in a comment or a call site can no longer carry a capability claim. The field's evaluators grade repositories, community activity, pages and patches; none of them re-derives a product's own capability table from its source on every publish Adoption: LIB-WIRED importers=20 nonval=18 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; ahead of OpenSSF Scorecard, CHAOSS, Lighthouse, SWE-bench; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ● | ○ | ○ | ○ | ○ |
Delivery and evidence
Inspect risks, technical debt, rendered observations, experiments and references. Read scope and limitations alongside every result.
Risk register
| Risk | Likelihood x impact | Mitigation |
|---|---|---|
| A model-as-judge that is not calibrated is an opinion with an authoritative name | likely x high | IN3 done-rule is agreement on a labelled set BEFORE any verdict; uncalibrated it reports numbers, never verdicts. |
| Audits wired census-wide hammer the box on the beat | possible x medium | IN2 rides the existing census infrastructure and asks nx_spendgate; one census owner per instrument per day. |
Person · product · place — not yet measured for this domain
knowledge/compare/instrument.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain instrument, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).The field — discovered, not chosen
Rows written by nx_field_discover from instrument.seeds: the industry's own lists (Wikipedia wikitext, GitHub topics, awesome lists) read mechanically, every candidate counted across seeds. The matrix columns above are a SEAT'S pick; this band is the population they were picked from, and the stats line measures one against the other. A rival here is a lead, never a verdict — it earns a column when its capabilities are read and pinned.
field|seeds=4|scanned=4|fetched=4|reused=0|failed=0|named=4|candidates=4|mentions=4|capped=0
rival|OpenSSF Scorecard|1|1|col0-openssf|https://github.com/ossf/scorecard/blob/main/docs/checks.md|named
rival|CHAOSS|1|1|col1-chaoss|https://chaoss.community/kb-metrics-and-metrics-models/|named
rival|Lighthouse|1|1|col2-lighthouse|https://developer.chrome.com/docs/lighthouse/overview|named
rival|SWE-bench|1|1|col3-swe|https://arxiv.org/abs/2310.06770|named
| Rank | Rival | Seeds | Mentions | First seed | Kind | Link |
|---|---|---|---|---|---|---|
| 1 | OpenSSF Scorecard | 1 | 1 | col0-openssf | named | https://github.com/ossf/scorecard/blob/main/docs/checks.md |
| 2 | CHAOSS | 1 | 1 | col1-chaoss | named | https://chaoss.community/kb-metrics-and-metrics-models/ |
| 3 | Lighthouse | 1 | 1 | col2-lighthouse | named | https://developer.chrome.com/docs/lighthouse/overview |
| 4 | SWE-bench | 1 | 1 | col3-swe | named | https://arxiv.org/abs/2310.06770 |
field candidates 4|shown 4 of 4|matrix columns in the field 4 of 4|discovered rivals with no column 0|malformed rows 0 (counted, never rendered)|read-capped 0
Gaps from the record — what the estate does that no board carries
The record census (nx_goalmap record) reads the invoked-tool population and every plan queue row and files each organ or directive that NO matrix, plan or gates row names. A row here is a callout the boards missed: adjudicate it onto a board or declare it infrastructure. Census state BLIND (age 74432 s), sources read 4 of 7 declared — a BLIND census is a FLOOR: unread sources can only add rows.
| kind | name | board | source | evidence |
|---|
rows shown 0|this board's directives 0|estate-wide un-boarded organs 492 (listed in full on /compare/ecosystem)|census rows 611|malformed 0 (counted, never rendered)
References
- [scorecard-checks] OpenSSF Scorecard: Check Documentation (docs/checks.md) -- the automated repository security checks (Binary-Artifacts, Branch-Protection, CI-Tests, Code-Review, Maintained, SAST, SBOM, Signed-Releases, Vulnerabilities and others). publisher · read in our library
knowledge/fetched/cmp_instrument_scorecard-checks.html· pinh0c8e0ca6826693f3b80261f35d8f4430187dd9fd687089dd9eefaf2e9c8979bb· accessed 2026-08-18 · vendor-docGrounds: The OpenSSF Scorecard column and the "Security and quality posture checks" row (Scorecard leads, graded 2): the check list itself -- the matrix note says 18 automated checks as banked; the live document read 2026-08-18 lists about 20, so the count on the page is a floor, not the current total. Nishi runs the CI-presence and maintained-ness subset from banked evidence. - [scorecard-paper23] Zahan, Kanakiya, Hambleton, Shohan, Williams. OpenSSF Scorecard: On the Path Toward Ecosystem-wide Automated Security Metrics. arXiv:2208.03412 (2022); DOI 10.1109/MSEC.2023.3279773 (2023). publisher · read in our library
knowledge/fetched/cmp_instrument_scorecard-paper23.html· pinhc4cbea8d2eb5a9f0cec321a242c5b343975cc989c403977ce2d80e4caa35c7d5· accessed 2026-08-18 · published-paperGrounds: The Scorecard column at large and the "Repo layer and full-stack structure analysis" row: the peer-reviewed account of ecosystem-wide automated repository scoring that nx_repo_health_gate is measured against (Scorecard/CHAOSS-class, liar-killed). - [chaoss-metrics] CHAOSS (Community Health Analytics in Open Source Software): Metrics and Metrics Models knowledge base -- single-question metrics and multi-metric models of community health. publisher · read in our library
knowledge/fetched/cmp_instrument_chaoss-metrics.html· pinhae7133efef9b089a51ac6d34a172ffccf57135b926c88b16a390b0544901e900· accessed 2026-08-18 · vendor-docGrounds: The CHAOSS column and the "Community and ecosystem health metrics" row (CHAOSS leads, graded 2): the metric definitions the row's contributor-base and community-health reads are graded against. - [lighthouse-overview] Chrome for Developers: Introduction to Lighthouse -- open-source automated audits for performance, accessibility, SEO and best practices of web pages. publisher · read in our library
knowledge/fetched/cmp_instrument_lighthouse-overview.html· pinh31972223ccea41cfc07f5b70abde2ca1d3ae4a1cc4a0232929a7e3d433f37fbd· accessed 2026-08-20 · vendor-docGrounds: The Lighthouse column and the "Automated product audits (perf, a11y, UX -- Lighthouse-class)" row (Lighthouse leads, Nishi _ABSENT_): what an automated product audit is in the field; our ui-judge/contrast/exceed graders are not yet wired census-wide. - [swebench23] Jimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024 (arXiv:2310.06770). publisher · read in our library
knowledge/fetched/cmp_coding_swebench23.html· pinhcd170cab3a56f891ee94070d5748d923685f2a84a91201d6336a221436fd0f9f· accessed 2026-08-18 · published-paperGrounds: The SWE-bench column and the "Benchmark harness discipline (gates, neg-controls, determinism)" row (SWE-bench is the harness-rigor reference, graded 2): the execution-based, containerized, held-out-test harness design our domain gates match on their own subjects. Mirror shared with coding.refs (same bytes, same pin). - [geval23] Liu, Iter, Xu, Wang, Xu, Zhu. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv:2303.16634, 2023. publisher · read in our library
knowledge/fetched/cmp_instrument_geval23.html· pinh663d4662d59ef070020d412d18c10095cd299776cc005a1fa934110273995d48· accessed 2026-08-18 · published-paperGrounds: The "Model-as-judge for subjective axes (G-Eval class)" row and the NEXT rung "sovereign LLM-judge for subjective axes": the chain-of-thought form-filling judge recipe the row names as buildable now that the no-float model generates faithfully. - [llmjudge23] Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv:2306.05685). publisher · read in our library
knowledge/fetched/cmp_instrument_llmjudge23.html· pinh77854867787231a84724d8825cb1754bf14a29b8a93629cf9fc00d76a0734123· accessed 2026-08-18 · published-paperGrounds: The same "Model-as-judge for subjective axes (G-Eval class)" row: the agreement-with-humans and known-bias analysis (position, verbosity, self-enhancement) any sovereign judge rung must carry as neg-controls before it grades a subjective axis. - [spdx-spec] SPDX (System Package Data Exchange) Specification 3.0.1, The Linux Foundation: the software bill of materials data model and serialization. publisher · read in our library
knowledge/fetched/cmp_devguardrails_spdx-spec.html· pinha905ca11b54700955ba92b4dbdf9d35553852b82bfddf2fabc06cabc194e120a· accessed 2026-08-18 · published-standardGrounds: The "Composition and SBOM analysis (SPDX class)" row (Nishi _ABSENT_): the interchange format an SPDX-interop composition analysis of evaluated suites would emit; our own stack is a trivially-complete SBOM but nothing states it in this shape. Mirror shared with devguardrails.refs.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/instrument.matrix · source checks show implementation presence; runtime and user-outcome evidence are reported separately · JavaScript supports page controls