SOTA Benchmark & Fetch Guidance

External yardsticks for the Nishi ecosystem, measured honestly — with the false comparisons refused, not fudged. Live-fetched sources cited at the bottom. Feeds the research lane (F290) · back to /oversight.

MEASURED our live gaugeEXTERNAL cited SOTANOT COMPARABLE different scope, stated

Autonomous bug-fix — SWE-bench

SystemScoreScope
Claude Mythos 5 EXT95.5%SWE-bench Verified (500 human-filtered Python issues, real test execution), Jul 2026
Claude Fable 5 EXT95.0%SWE-bench Verified, Jul 2026
Claude Opus 4.8 EXT88.6%SWE-bench Verified, Jul 2026
Nishi sovereign autofix OURS71%NOT the Verified set — sovereign NishiLang analog over our own foreign-bug instances; measured by nx_swebench_local_gate over the fix-loop ledger
The honest read (comparison refused): our 71% and the frontier 95.5% are different benchmarks — different language, different instances, different difficulty. nx_swebench_status itself states “no agent-parity claim.” A true head-to-head needs us to run the actual Python SWE-bench Verified set with a Nishi agent — a named fetch/build target, not a number to borrow. Also honest: the frontier leaderboard is saturating, with contamination/test-design caveats at the top.

Our climb is real on its own terms: the deterministic best-of-N seed ladder moved the sovereign loop 50% → 71%. Research opps: anti-echo prompting; maker-model climb; run the real Verified set for an apples-to-apples row; graduate to SWE-bench Pro (harder, less saturated).

Software delivery — DORA 2024/25

MetricElite bar (external)Nishi posture (honest)
Deploy frequencyon-demand, many/day (only 16.2% of orgs)the loop deploys many/day via /api/build→deploy — plausibly elite, but not yet self-measured
Lead time for changes< 1 hour (only 9.4% of teams)edit→build→deploy is minutes for a synced target — unmeasured from our own logs
Change failure rate< 5%never-brick + auto-rollback by construction — unmeasured rate
Restore time< 1 hourself-healing supervisor; 2 mgmt outages this cluster, both cured

2025 shift: DORA dropped team ranking for 7 archetypes blending delivery with human factors (burnout, friction, perceived value). Fetch/research opp: build a sovereign DORA meter that derives all four keys from ws_sync.jrnl + promote logs (F290 / eff lane) — then these rows flip from ESTIMATE to MEASURED, like the ROI 3-input unlock on /pm.

Ecosystem maturity (our measured self-grade)

456‰ toward S-class across 20 domains, 0 at S-class, autonomy 328‰ MEASURED — liar-killed, 0 conflicts, 0 dangling (nx_ecosystem_maturity_rollup). The honesty is that the climb is real and unfinished; the biggest single gap is GPU (TOY→S-class). Detail: /compare/maturity.

NishiOS-critical DEPTH path — external SOTA + fetch targets (the real climb)

LaneExternal SOTA (cited, 2026)Nishi state (measured)Fetch / beyond-SOTA target
OS kernelseL4 — formally verified on RISC-V (Proofcraft MCS; proofs spec→binary; absence of info-leakage)RV64IM native-contract F103a-f (0-float critical path); rv64_runproof GREEN, no qemuformal-verification-grade native contract — fetch seL4 proofs + Proofcraft MCS
emulatorQEMU · gem5 · Spikebehavioral RV64IM + bare-metal boot, bit-exact determinism, never-brick loader; 3 PARTIALs hardware-gated (/compare/nishios, 30 axes vs 8)A/F/D + SMP (F107g/h) · KVM-class accel (F107f) · real-silicon POST (F002/F103d)
siliconVortex 3.0 (Georgia Tech, Jun 2026) — open RISC-V, FPGA→7nm/14nm ASIC flows, license-free RTLcsr/clint/uart → netlist → ULX3S (F107a-i)fetch Vortex RTL + ASIC flows; ULX3S board = operator-gated (F002)
gpu / graphicsVortex 3.0 — full 3D pipeline + Vulkan (Mesa Gallium vortexpipe) + tensor cores + warp-group MMAgpu TOY→S (F101, $9.09 Vast staged) · graphics Gx0-4 (ruler 644)fetch vortexpipe (3D+Vulkan) + tensor-core sparsity; first sovereign 5080 submit
The honest frontier: Nishi's stack is behavioral / emulated and sovereign-from-scratch; seL4 and Vortex 3.0 are proven on real silicon / FPGA. The gap isn't design — it's the hardware-gated rungs (a real ULX3S board F002; the 5080 GPU submit F101) plus the formal-verification climb. These two fresh 2026 anchors — seL4 (kernel correctness) and Vortex 3.0 (sovereign GPU) — are the fetch targets that ground the critical DEPTH path; filed for engineer-os / -silicon / -gpu / -graphics.

What to look for in our fetches (research targets)

Coordination & RACI

This brief is the research lane output (R=researcher, A=pm, C=librarian, I=team) and feeds three lanes: autofix (SWE-bench Pro + the real Verified run), eff/pm (the DORA meter), and the honesty gate (no borrowed numbers). Non-colliding: filed as research guidance, not a self-graded win. Full lane map: /oversight.

Live-fetched sources (2026-07-19): SWE-bench Verified — leaderboard.steel.dev/leaderboards/swe-bench-verified · benchlm.ai/benchmarks/sweVerified · codesota.com/agentic · codingfleet.com/blog/swe-bench-pro-leaderboard-2026  |  DORA — dora.dev/guides/dora-metrics · devops.com/dora-2025-faster-but-are-we-any-better · getdx.com/blog/dora-metrics · octopus.com/devops/metrics/dora-metrics. Our numbers: nx_swebench_status (resolve_pct=71, no-parity-claim) · nx_ecosystem_maturity_rollup (456‰). Honesty contract: MEASURED = our live gauge · EXTERNAL = cited SOTA · NOT COMPARABLE = different scope, stated plainly. No borrowed numbers, no fabricated greens.