Nishi Family › Compare › Evaluation Instrument
666
/1000 coverage
Established
Evaluation Instrument
The sovereign tooling that evaluates any program, product, app, capability, ecosystem or suite from its git, docs, schematics and papers — measured against the field's own 2025/26 evaluation tools, at zero token cost, with drift and false-SOTA liar-killed.
3 sovereign exceeds · 8/12 capabilities present · bar: OpenSSF Scorecard|CHAOSS|Lighthouse|SWE-bench
The climb — development arc
DONE 2026-07-09 R0 anti-drift coach
nx_ws_coach GREEN 3/3
gate→census→coach: only COACH-OK claims a rung; refuses any parity/exceed claim without a paired evidence token
DONE 2026-07-09 R1 product-bar researcher
nx_product_bar_gate GREEN 6/6
banks entire OSS products (full-stack: ui/backend/infra/mobile) keyless
DONE 2026-07-09 R2 field-derived axes
nx_axis_discover GREEN 5/5
census axes come from cross-product corroboration of banked corpora, not hand-picked
DONE 2026-07-10 R3 structured claim/evidence
nx_ws_coach @claim/@ev
machine-checkable claim-proof pairing kills vocabulary false-drifts
DONE 2026-07-10 E2/E3 repo health (Scorecard/CHAOSS-class)
nx_repo_health_gate GREEN 7/7
CI + maintained-ness + contributors + community health from banked evidence, liar-killed
DONE 2026-07-10 LIVE + self-assembling registry
nx_compare_registry_assemble + publish_hub.sh
integrated into /compare, non-regressive publish, sovereign-verified
NOW 2026-07-10 unified comparison view
nx_compare_unified
consolidate matrix+frontier+hub+api into one high-level→granular workstream-arc view (this page)
NEXT sovereign LLM-judge for subjective axes
-
the no-float model now generates faithfully → a G-Eval-class judge is buildable
NEXT converge with nx_ecosystem_maturity_rollup
-
fold instrument signals into the one ecosystem SOTA critic; ONE evaluator not four
Capabilities — measured vs the field
| Capability | Nishi |
OpenSSF Scorecard | CHAOSS | Lighthouse | SWE-bench |
| Papers and frontier banking (git/docs/white-papers/research) | Best | No | No | No | No |
| Field-derived comparison axes (not hand-picked) | Yes | No | No | No | No |
| Repo layer and full-stack structure analysis | Yes | Yes | Yes | No | Yes |
| Security and quality posture checks | Yes | Best | No | No | No |
| Community and ecosystem health metrics | Yes | No | Best | No | No |
| Benchmark harness discipline (gates, neg-controls, determinism) | Yes | No | No | No | Best |
| Meta-evaluation -- the evaluator audits itself | Best | No | No | No | No |
| Drift and claim/proof gating (a claim needs its evidence) | Best | No | No | No | No |
| Automated product audits (perf, a11y, UX -- Lighthouse-class) | No | No | No | Best | No |
| Model-as-judge for subjective axes (G-Eval class) | No | No | No | No | No |
| Living refresh and scheduled cadence | No | No | No | No | No |
| Composition and SBOM analysis (SPDX class) | No | No | No | No | No |
coverage 666/1000 · Best 3 · Yes 5 · No 4 · every Nishi cell measured by implementing-symbol on disk (no self-grading)
Evidence & reuse — so workstreams build on, not rebuild
Papers and frontier banking (git/docs/white-papers/research) — Best
Nishi banks the live 2025/26 frontier keyless over sovereign TLS at ZERO token cost (0 vs 107 LLM agents/query); none of the eval tools ingest the research frontier
organ: runtime/nx_swcompare_research.nx · verified-symbol: rf_fetch
REUSE: reuse: nx_swcompare_research + a per-domain .q spec -- ANY workstream banks its field bars this way, zero new code
Field-derived comparison axes (not hand-picked) — Yes
axes derived from banked product corpora by cross-product corroboration (782 terms, private vocab excluded); the field hand-authors metric lists
organ: runtime/_hdl_build/nx_axis_discover.nx · verified-symbol: corroborat
REUSE: reuse: nx_axis_discover derives a domains axes from its banked bars; reuse for any new census
Repo layer and full-stack structure analysis — Yes
Nishi banks the UI/backend/infra/mobile layer fingerprint from the keyless GitHub API; Scorecard/CHAOSS/SWE-bench also read repo trees
organ: runtime/_hdl_build/nx_product_bar_gate.nx · verified-symbol: LAYER
Security and quality posture checks — Yes
OpenSSF Scorecard leads (18 automated checks); Nishi runs the CI-presence and maintained-ness subset from banked evidence, liar-killed (unbanked = NO-EVIDENCE, never healthy)
organ: runtime/_hdl_build/nx_repo_health_gate.nx · verified-symbol: Scorecard
REUSE: reuse: nx_repo_health_gate scores any repo Scorecard/CHAOSS-class from banked GitHub evidence
Community and ecosystem health metrics — Yes
CHAOSS leads; Nishi reads contributor base and GitHub community health_percentage from banked profiles
organ: runtime/_hdl_build/nx_repo_health_gate.nx · verified-symbol: community
Benchmark harness discipline (gates, neg-controls, determinism) — Yes
SWE-bench is the harness-rigor reference; Nishi matches on its own domains -- every rung gated with neg-controls and determinism floors
organ: runtime/_hdl_build/nx_evalsota_census.nx · verified-symbol: INSTRUMENT
Meta-evaluation -- the evaluator audits itself — Best
Nishi scans its own 162 censuses and FLAGS the 83 giving false SOTA signals; the field's tools rarely self-audit
organ: runtime/_hdl_build/nx_census_uiaudit.nx · verified-symbol: FLAGGED
REUSE: reuse: nx_census_uiaudit flags any product-surface census missing U-axes; run ecosystem-wide
Drift and claim/proof gating (a claim needs its evidence) — Best
the coach REFUSES any parity/exceed claim without a paired gate/live-URL token; no eval tool enforces this on its own verdicts
organ: runtime/_hdl_build/nx_ws_coach.nx · verified-symbol: COACH-OK
REUSE: reuse: nx_ws_coach gates every workstream rung (claim/proof + freshness); run before claiming SOTA
Automated product audits (perf, a11y, UX -- Lighthouse-class) — No
Lighthouse leads automated product audits; Nishi has ui-judge/contrast/exceed graders but they are not yet wired census-wide
organ: runtime/_hdl_build/nx_evalsota_census.nx · verified-symbol: _ABSENT_
Model-as-judge for subjective axes (G-Eval class) — No
domain judges are gated; a general sovereign LLM-judge is buildable now that the no-float model generates faithfully -- named next rung
organ: runtime/_hdl_build/nx_evalsota_census.nx · verified-symbol: _ABSENT_
Living refresh and scheduled cadence — No
the freshness gate exists (current-quarter check); automated re-run cadence on the bg census infra is the next rung
organ: runtime/_hdl_build/nx_ws_coach.nx · verified-symbol: _ABSENT_
Composition and SBOM analysis (SPDX class) — No
our own stack is a trivially-complete SBOM (zero third-party deps); SPDX-interop composition analysis of evaluated suites is absent
organ: runtime/_hdl_build/nx_evalsota_census.nx · verified-symbol: _ABSENT_