External yardsticks for the Nishi ecosystem, measured honestly — with the false comparisons refused, not fudged. Live-fetched sources cited at the bottom. Feeds the research lane (F290) · back to /oversight.
| System | Score | Scope |
|---|---|---|
| Claude Mythos 5 EXT | 95.5% | SWE-bench Verified (500 human-filtered Python issues, real test execution), Jul 2026 |
| Claude Fable 5 EXT | 95.0% | SWE-bench Verified, Jul 2026 |
| Claude Opus 4.8 EXT | 88.6% | SWE-bench Verified, Jul 2026 |
| Nishi sovereign autofix OURS | 71% | NOT the Verified set — sovereign NishiLang analog over our own foreign-bug instances; measured by nx_swebench_local_gate over the fix-loop ledger |
nx_swebench_status itself states “no agent-parity claim.” A true head-to-head needs us to run the actual Python SWE-bench Verified set with a Nishi agent — a named fetch/build target, not a number to borrow. Also honest: the frontier leaderboard is saturating, with contamination/test-design caveats at the top.Our climb is real on its own terms: the deterministic best-of-N seed ladder moved the sovereign loop 50% → 71%. Research opps: anti-echo prompting; maker-model climb; run the real Verified set for an apples-to-apples row; graduate to SWE-bench Pro (harder, less saturated).
| Metric | Elite bar (external) | Nishi posture (honest) |
|---|---|---|
| Deploy frequency | on-demand, many/day (only 16.2% of orgs) | the loop deploys many/day via /api/build→deploy — plausibly elite, but not yet self-measured |
| Lead time for changes | < 1 hour (only 9.4% of teams) | edit→build→deploy is minutes for a synced target — unmeasured from our own logs |
| Change failure rate | < 5% | never-brick + auto-rollback by construction — unmeasured rate |
| Restore time | < 1 hour | self-healing supervisor; 2 mgmt outages this cluster, both cured |
2025 shift: DORA dropped team ranking for 7 archetypes blending delivery with human factors (burnout, friction, perceived value). Fetch/research opp: build a sovereign DORA meter that derives all four keys from ws_sync.jrnl + promote logs (F290 / eff lane) — then these rows flip from ESTIMATE to MEASURED, like the ROI 3-input unlock on /pm.
456‰ toward S-class across 20 domains, 0 at S-class, autonomy 328‰ MEASURED — liar-killed, 0 conflicts, 0 dangling (nx_ecosystem_maturity_rollup). The honesty is that the climb is real and unfinished; the biggest single gap is GPU (TOY→S-class). Detail: /compare/maturity.
| Lane | External SOTA (cited, 2026) | Nishi state (measured) | Fetch / beyond-SOTA target |
|---|---|---|---|
| OS kernel | seL4 — formally verified on RISC-V (Proofcraft MCS; proofs spec→binary; absence of info-leakage) | RV64IM native-contract F103a-f (0-float critical path); rv64_runproof GREEN, no qemu | formal-verification-grade native contract — fetch seL4 proofs + Proofcraft MCS |
| emulator | QEMU · gem5 · Spike | behavioral RV64IM + bare-metal boot, bit-exact determinism, never-brick loader; 3 PARTIALs hardware-gated (/compare/nishios, 30 axes vs 8) | A/F/D + SMP (F107g/h) · KVM-class accel (F107f) · real-silicon POST (F002/F103d) |
| silicon | Vortex 3.0 (Georgia Tech, Jun 2026) — open RISC-V, FPGA→7nm/14nm ASIC flows, license-free RTL | csr/clint/uart → netlist → ULX3S (F107a-i) | fetch Vortex RTL + ASIC flows; ULX3S board = operator-gated (F002) |
| gpu / graphics | Vortex 3.0 — full 3D pipeline + Vulkan (Mesa Gallium vortexpipe) + tensor cores + warp-group MMA | gpu TOY→S (F101, $9.09 Vast staged) · graphics Gx0-4 (ruler 644) | fetch vortexpipe (3D+Vulkan) + tensor-core sparsity; first sovereign 5080 submit |
This brief is the research lane output (R=researcher, A=pm, C=librarian, I=team) and feeds three lanes: autofix (SWE-bench Pro + the real Verified run), eff/pm (the DORA meter), and the honesty gate (no borrowed numbers). Non-colliding: filed as research guidance, not a self-graded win. Full lane map: /oversight.