{"v":2,"api":"nishi-compare-unified","domain":"instrument","title":"Evaluation Instrument","summary":"The sovereign tooling that evaluates any program, product, app, capability, ecosystem or suite from its git, docs, schematics and papers &mdash; measured against the field's own 2025/26 evaluation tools, at zero token cost, with drift and false-SOTA liar-killed.","coverage":666,"band":"Established","exceeds":3,"present":5,"absent":4,"total":12,"arc":[{"status":"DONE 2026-07-09 ","label":" R0 anti-drift coach ","evidence":" nx_ws_coach GREEN 3/3 "},{"status":"DONE 2026-07-09 ","label":" R1 product-bar researcher ","evidence":" nx_product_bar_gate GREEN 6/6 "},{"status":"DONE 2026-07-09 ","label":" R2 field-derived axes ","evidence":" nx_axis_discover GREEN 5/5 "},{"status":"DONE 2026-07-10 ","label":" R3 structured claim/evidence ","evidence":" nx_ws_coach @claim/@ev "},{"status":"DONE 2026-07-10 ","label":" E2/E3 repo health (Scorecard/CHAOSS-class) ","evidence":" nx_repo_health_gate GREEN 7/7 "},{"status":"DONE 2026-07-10 ","label":" LIVE + self-assembling registry ","evidence":" nx_compare_registry_assemble + publish_hub.sh "},{"status":"NOW 2026-07-10 ","label":" unified comparison view ","evidence":" nx_compare_unified "},{"status":"NEXT ","label":" sovereign LLM-judge for subjective axes ","evidence":" - "},{"status":"NEXT ","label":" converge with nx_ecosystem_maturity_rollup ","evidence":" - "}],"axes":[{"label":"Papers and frontier banking (git/docs/white-papers/research)","nishi":2,"organ":"runtime/nx_swcompare_research.nx","note":"Nishi banks the live 2025/26 frontier keyless over sovereign TLS at ZERO token cost (0 vs 107 LLM agents/query); none of the eval tools ingest the research frontier","reuse":"reuse: nx_swcompare_research + a per-domain .q spec -- ANY workstream banks its field bars this way, zero new code"},{"label":"Field-derived comparison axes (not hand-picked)","nishi":1,"organ":"runtime/_hdl_build/nx_axis_discover.nx","note":"axes derived from banked product corpora by cross-product corroboration (782 terms, private vocab excluded); the field hand-authors metric lists","reuse":"reuse: nx_axis_discover derives a domains axes from its banked bars; reuse for any new census"},{"label":"Repo layer and full-stack structure analysis","nishi":1,"organ":"runtime/_hdl_build/nx_product_bar_gate.nx","note":"Nishi banks the UI/backend/infra/mobile layer fingerprint from the keyless GitHub API; Scorecard/CHAOSS/SWE-bench also read repo trees","reuse":""},{"label":"Security and quality posture checks","nishi":1,"organ":"runtime/_hdl_build/nx_repo_health_gate.nx","note":"OpenSSF Scorecard leads (18 automated checks); Nishi runs the CI-presence and maintained-ness subset from banked evidence, liar-killed (unbanked = NO-EVIDENCE, never healthy)","reuse":"reuse: nx_repo_health_gate scores any repo Scorecard/CHAOSS-class from banked GitHub evidence"},{"label":"Community and ecosystem health metrics","nishi":1,"organ":"runtime/_hdl_build/nx_repo_health_gate.nx","note":"CHAOSS leads; Nishi reads contributor base and GitHub community health_percentage from banked profiles","reuse":""},{"label":"Benchmark harness discipline (gates, neg-controls, determinism)","nishi":1,"organ":"runtime/_hdl_build/nx_evalsota_census.nx","note":"SWE-bench is the harness-rigor reference; Nishi matches on its own domains -- every rung gated with neg-controls and determinism floors","reuse":""},{"label":"Meta-evaluation -- the evaluator audits itself","nishi":2,"organ":"runtime/_hdl_build/nx_census_uiaudit.nx","note":"Nishi scans its own 162 censuses and FLAGS the 83 giving false SOTA signals; the field's tools rarely self-audit","reuse":"reuse: nx_census_uiaudit flags any product-surface census missing U-axes; run ecosystem-wide"},{"label":"Drift and claim/proof gating (a claim needs its evidence)","nishi":2,"organ":"runtime/_hdl_build/nx_ws_coach.nx","note":"the coach REFUSES any parity/exceed claim without a paired gate/live-URL token; no eval tool enforces this on its own verdicts","reuse":"reuse: nx_ws_coach gates every workstream rung (claim/proof + freshness); run before claiming SOTA"},{"label":"Automated product audits (perf, a11y, UX -- Lighthouse-class)","nishi":0,"organ":"runtime/_hdl_build/nx_evalsota_census.nx","note":"Lighthouse leads automated product audits; Nishi has ui-judge/contrast/exceed graders but they are not yet wired census-wide","reuse":""},{"label":"Model-as-judge for subjective axes (G-Eval class)","nishi":0,"organ":"runtime/_hdl_build/nx_evalsota_census.nx","note":"domain judges are gated; a general sovereign LLM-judge is buildable now that the no-float model generates faithfully -- named next rung","reuse":""},{"label":"Living refresh and scheduled cadence","nishi":0,"organ":"runtime/_hdl_build/nx_ws_coach.nx","note":"the freshness gate exists (current-quarter check); automated re-run cadence on the bg census infra is the next rung","reuse":""},{"label":"Composition and SBOM analysis (SPDX class)","nishi":0,"organ":"runtime/_hdl_build/nx_evalsota_census.nx","note":"our own stack is a trivially-complete SBOM (zero third-party deps); SPDX-interop composition analysis of evaluated suites is absent","reuse":""}]}
