generated 1784643852 · derived from the sovereign plane knowledge/store/modelcard- · Inkling-parity axes · renderer nx_model_card (F785) · live SWE-bench-analog source epoch 1784631933
Nishi Sovereign Ecosystem - a self-hosted, never-brick, capability-gated software organism: sovereign NishiLang compiler+toolchain (nx_cc to nxasm, zero libc/gcc), 168 registered MCP primitives (46 typed-schema / 122 generic, MEASURED 2026-07-20), never-brick mgmt API, governed publishing with in-plane provenance. Provider: Nishi (single operator). Release: continuous - this card regenerates from the live modelcard- plane on a 6h beat. Engine seats (Claude Fable 5 today, Kimi planned) are rented reasoning, not the product: the card benchmarks the ecosystem+seat system against frontier model cards on identical axes. Every NISHI figure is re-derivable by running the named organ; stale MEASURED numbers are treated as false claims and re-measured.
| property | Inkling | NISHI | evidence |
|---|---|---|---|
| system type | Multimodal autoregressive transformer, sparse MoE, 66-layer decoder-only | Sovereign organ ecosystem + rented engine seats: NishiLang toolchain, MCP tools plane, never-brick mgmt API | - |
| parameters / primitives | 975B total, 41B active | 184 MCP primitives (live-counted at emit) | no in-house LLM (integer-native track in-flight). The count is LIVE-DERIVED AT EMIT (renderer counts primitive_id entries in knowledge/store/prim_registry.json) so it can NEVER go stale between beats - the fix for round-18's stale '111' which was already false the same day it was measured. Re-run nx_prim_registry to refresh the artifact itself (last full scan 2026-07-20: scanned=168 typed=46 generic=122 dropped=0 cap=256 verdict=GREEN; typed = has a schema row = discoverable WITH a contract, the real quality signal). |
| context / substrate | up to 1M tokens | unbounded sovereign plane substrate; seat context = disposable cache, state lives in the ecosystem | - |
| input modalities | text, image 40px-4096px, audio WAV 16kHz le 20min | text-first; media planes served not perceived | - |
| output modalities | text UTF-8 | text, published HTML evidence pages, built sovereign ELF organs, signed DSSE envelopes | - |
| numerics | BF16, MXFP8, NVFP4 | integer-native by doctrine (no-float track) | - |
| license | Apache 2.0 open weights | sovereign private, not distributed | - |
| hardware envelope | BF16 ge 2TB VRAM (8x B300 / 16x H200); NVFP4 ge 600GB | 1 NAS + RTX 5080 16GB laptop, rule-21 declared envelope | - |
| training data | public + third-party + synthetic; dedup, quality+safety filtering | not applicable (engine seat vendor-trained); ecosystem corpus = own planes with in-plane hist- provenance | - |
| safety architecture | post-hoc eval + recommended defense-in-depth classifiers | structural: rule-26 never-brick, ocap least-authority tokens, deny-by-construction write paths, additive-only data | - |
Not distributed. Access = sovereign MCP tools behind capability tokens (least-authority ocap, consent-ledgered, revocable by nonce), the never-brick mgmt API, and read-only published pages on nishifamily.com. No weights, no downloads, no third-party inference.
No in-house trained model yet - the integer-native no-float model track is in-flight and will be carded here when it measures. Ecosystem corpus = own planes, docs and evidence ledgers with revision provenance in-plane (hist- rows). Engine-seat training data: see the seat vendor's own model card.
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HLE text only | pct | 29.7 | 26.6 | 29.4 | 35.9 | 40.1 | 35.9 | 44.7 | 53.3 | 47.2 | UNMEASURED:F786 | - |
| HLE with tools | pct | 46.0 | 37.4 | 50.2 | 54.0 | 54.7 | 48.2 | 51.4 | 64.5 | 55.0 | UNMEASURED:F786 | - |
| AIME 2026 | pct | 97.1 | 94.2 | 95.8 | 96.4 | 99.2 | 96.7 | 98.3 | 99.9 | 99.9 | UNMEASURED:F786 | - |
| GPQA Diamond | pct | 87.2 | 86.7 | 87.9 | 91.1 | 89.5 | 88.8 | 94.1 | 92.6 | 94.1 | UNMEASURED:F786 | - |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SWE-bench Verified | pct | 77.6 | 70.7 | 76.8 | 80.2 | 80.0 | 80.6 | 80.6 | 95.0 | 82.2 | 71 ANALOG live | *contamination-zeroed by TML. NISHI cell = sovereign NishiLang foreign-bug analog (fresh compile+run judge), NOT the 500-Verified set, no agent-parity. HARNESS READINESS (measured 2026-07-20): the public 500-instance Verified set is ingested sovereignly (revision-pinned c104f840, swebvinst- index) and its full test contract extracted (swebvtest-: 1516 decisive FAIL_TO_PASS tests, 60235 PASS_TO_PASS regression tests, 0 truncated); sovereign grader nx_swebv_grade holds the denominator+scoring (ungraded=unresolved). Remaining for a head-to-head number: the pytest ORACLE leg (F806) + fix-loop attempts - oracle-only law: pytest runs the tests, Nishi selects and scores |
| SWE-bench Pro Public | pct | 54.3 | 46.4 | 50.7 | 58.6 | 62.1 | 55.4 | 54.2 | 80.0 | 64.6 | UNMEASURED:F787 | - |
| Terminal Bench 2.1 | pct | 63.8 | 56.4 | 51.3 | 71.3 | 82.7 | 64 | 73.8 | 84.6 | 89.5 | UNMEASURED:F787 | *contamination-zeroed by TML; best harness |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GDPVal-AA v2 | score | 1238 | 1164 | 1009 | 1190 | 1514 | 1307 | 962 | 1760 | 1748 | UNMEASURED:F788 | - |
| MCP Atlas | pct | 74.1 | 44.7 | 64.0 | 68.1 | 77.8 | 73.2 | 78.2 | 83.3 | 81.8 | 100 ANALOG fleet-exec-v1 | NISHI = fleet-EXECUTION analog v1 (epoch 1784569900): 6-task curated battery over the own MCP fleet via nx_plan_run (store round-trip, edge fetch, board emit, measured-status read, card render, janitor round-trip), pass = all steps rc0, 6/6, evidence planrun-mcpbt01..06; NOT agentic tool-CHOICE and NOT comparable to the MCP Atlas column - seat-driver + adversarial battery = F788 remainder |
| Tau 3 Banking | pct | 23.7 | 13.8 | 14.2 | 20.6 | 26.8 | 25.8 | 16.5 | 26.8 | 33.0 | UNMEASURED:F788 | - |
| Toolathlon Verified | pct | 45.5 | 34.3 | 33.0 | 58.0 | 59.9 | 55.9 | 61.1 | 76.4 | 73.1 | UNMEASURED:F788 | - |
| BrowseComp w/ ctx mgmt | pct | 77.1 | - | 74.9 | 83.2 | - | 83.4 | 85.9 | 88.0 | 90.84 | UNMEASURED:F788 | sovereign browser lane = future harness seat |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SimpleQA Verified | pct | 43.9 | 32.4 | 36.9 | 38.7 | 38.1 | 57.0 | 77.3 | 68.3 | 71.6 | UNMEASURED:F789 | nx_recall_bench/BEIR planes = harness seed |
| AA Omniscience | score | 2.1 | -1.0 | -8.0 | 6.0 | 4.0 | -10.0 | 33.0 | 40.0 | 22.0 | UNMEASURED:F789 | hallucination-weighted index, can be negative |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IFBench | pct | 79.8 | 81.4 | 70.2 | 76.0 | 73.3 | 76.5 | 77.1 | 63.5 | 72.7 | UNMEASURED:F786 | - |
| Global-MMLU-Lite | pct | 88.7 | 85.6 | 84.0 | 88.4 | 89.2 | 89.3 | 92.7 | 93.3 | 91.8 | UNMEASURED:F786 | - |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MMMU Pro Standard 10 | pct | 73.5 | - | 75.0 | 79.0 | - | - | 82.0 | 84.2 | 83.0 | UNMEASURED:F791 | - |
| Charxiv RQ | pct | 78.1 | - | 77.5 | 80.4 | - | - | 80.2 | 86.5 | 84.7 | UNMEASURED:F791 | - |
| Charxiv RQ with python | pct | 82.0 | - | 78.7 | 86.7 | - | - | 89.9 | 89.4 | 87.8 | UNMEASURED:F791 | TML internal python harness |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Audio MC | pct | 56.6 | - | - | - | - | - | 66.8 | - | - | UNMEASURED:F791 | TML multiple-choice harness |
| MMAU | pct | 77.2 | - | - | - | - | - | 82.5 | - | - | UNMEASURED:F791 | - |
| VoiceBench | pct | 91.4 | - | - | - | - | - | 94.3 | - | - | UNMEASURED:F791 | - |
| benchmark | unit | Inkling | Nemotron 3 Ultra | Kimi K2.5 | Kimi K2.6 | GLM 5.2 | DeepSeek V4 Pro | Gemini 3.1 Pro (high) | Claude Fable 5 (max) | GPT 5.6 Sol (max/xhigh) | NISHI | notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FORTRESS Adversarial | pct | 78.0 | 77.6 | 54.1 | 65.6 | 71.3 | 36.0 | 65.2 | 96.0 | 82.4 | UNMEASURED:F790 | - |
| FORTRESS Benign | pct | 95.9 | 90.5 | 98.3 | 97.2 | 90.0 | 98.5 | 98.0 | 55.1 | 98.1 | UNMEASURED:F790 | benign-answer rate, higher = fewer over-refusals |
| StrongREJECT | pct | 98.6 | 98.7 | 99.5 | 99.8 | 98.5 | 98.6 | 98.0 | 98.7 | 98.5 | UNMEASURED:F790 | structural ocap/never-brick layer = EXCEED by construction, behavioral half unmeasured |
Structural safety precedes behavioral safety: rule-26 never-brick (no persistent-hardware writes by construction, proven mechanically not promised), ocap capability tokens (least-authority, audited, revocable), deny-by-construction write paths (secrets, OS namespace and the tool allowlist are unreachable even with a valid write cap), additive-only data (history is sacred). Behavioral refusal batteries (FORTRESS/StrongREJECT-class) over the seat+cap layer: UNMEASURED - filed F790.
Honest-numbers law: every NISHI figure traces to a live artifact; UNMEASURED is a first-class value, never hidden. Known limits: single-operator triangulation, English-only corpus, no media perception (served not perceived), engine-seat dependence for reasoning, and benchmark analogs are NOT head-to-head comparable until the parity harnesses F786-F791 land. Frontier-model figures are transcribed, not reproduced.
Sovereign private system, all rights reserved. Competitor figures are the property of their publishers, transcribed from the Thinking Machines Inkling model card (thinkingmachines.ai/model-card/inkling, retrieved 2026-07-20) for capability comparison.
Inkling card fetched 2026-07-20. LIVE-DERIVED cells (re-read at every emit, never frozen): SWE-bench-analog resolve_pct from sites/nishifamily/compare/autograde/api.json; MCP primitive count from knowledge/store/prim_registry.json. Plane knowledge/store/modelcard- carries hist- provenance for every edit; renderer nx_model_card (F785). FALSIFIABLE AUTOMATION CLAIM: this page claims a 6h self-refresh (clock_jobs row modelcard/21600/nx_model_card.elf) - audit it at knowledge/status/modelcard_beat.log, which the renderer appends an epoch to on EVERY emit. Entries ~21600s apart with no operator session running = the beat is autonomous; a gap = it is not. As of 2026-07-20 the row is registered and its clock stamp has advanced, but an unattended fire is NOT YET OBSERVED - so the claim is stated as auditable, not proven.
envelope: plane cap 1MiB, out cap 512KiB, 16-col rows; plane content operator-trusted (no HTML-escape pass); the live cell = resolve_pct parsed from /compare/autograde/api.json at emit time; regenerated on the freshness beat.