Nishi Model Card

generated 1784643852 · derived from the sovereign plane knowledge/store/modelcard- · Inkling-parity axes · renderer nx_model_card (F785) · live SWE-bench-analog source epoch 1784631933

How to read: 10 model columns = the 9 models from the Inkling card + NISHI. NISHI cell grammar: MEASURED (green, traces to a live artifact), ANALOG (orange, measured on an analog task set - NOT head-to-head comparable), UNMEASURED:F-rung (gray, harness filed on the frontier board). Footnotes mirror the Inkling card. Every NISHI figure must trace to a live tool or published artifact - fabricated numbers are structurally refused.

1 · General Information

Nishi Sovereign Ecosystem - a self-hosted, never-brick, capability-gated software organism: sovereign NishiLang compiler+toolchain (nx_cc to nxasm, zero libc/gcc), 168 registered MCP primitives (46 typed-schema / 122 generic, MEASURED 2026-07-20), never-brick mgmt API, governed publishing with in-plane provenance. Provider: Nishi (single operator). Release: continuous - this card regenerates from the live modelcard- plane on a 6h beat. Engine seats (Claude Fable 5 today, Kimi planned) are rented reasoning, not the product: the card benchmarks the ecosystem+seat system against frontier model cards on identical axes. Every NISHI figure is re-derivable by running the named organ; stale MEASURED numbers are treated as false claims and re-measured.

2 · System Properties

propertyInklingNISHIevidence
system typeMultimodal autoregressive transformer, sparse MoE, 66-layer decoder-onlySovereign organ ecosystem + rented engine seats: NishiLang toolchain, MCP tools plane, never-brick mgmt API-
parameters / primitives975B total, 41B active184 MCP primitives (live-counted at emit)no in-house LLM (integer-native track in-flight). The count is LIVE-DERIVED AT EMIT (renderer counts primitive_id entries in knowledge/store/prim_registry.json) so it can NEVER go stale between beats - the fix for round-18's stale '111' which was already false the same day it was measured. Re-run nx_prim_registry to refresh the artifact itself (last full scan 2026-07-20: scanned=168 typed=46 generic=122 dropped=0 cap=256 verdict=GREEN; typed = has a schema row = discoverable WITH a contract, the real quality signal).
context / substrateup to 1M tokensunbounded sovereign plane substrate; seat context = disposable cache, state lives in the ecosystem-
input modalitiestext, image 40px-4096px, audio WAV 16kHz le 20mintext-first; media planes served not perceived-
output modalitiestext UTF-8text, published HTML evidence pages, built sovereign ELF organs, signed DSSE envelopes-
numericsBF16, MXFP8, NVFP4integer-native by doctrine (no-float track)-
licenseApache 2.0 open weightssovereign private, not distributed-
hardware envelopeBF16 ge 2TB VRAM (8x B300 / 16x H200); NVFP4 ge 600GB1 NAS + RTX 5080 16GB laptop, rule-21 declared envelope-
training datapublic + third-party + synthetic; dedup, quality+safety filteringnot applicable (engine seat vendor-trained); ecosystem corpus = own planes with in-plane hist- provenance-
safety architecturepost-hoc eval + recommended defense-in-depth classifiersstructural: rule-26 never-brick, ocap least-authority tokens, deny-by-construction write paths, additive-only data-

3 · Methods of Distribution

Not distributed. Access = sovereign MCP tools behind capability tokens (least-authority ocap, consent-ledgered, revocable by nonce), the never-brick mgmt API, and read-only published pages on nishifamily.com. No weights, no downloads, no third-party inference.

4 · Training

No in-house trained model yet - the integer-native no-float model track is in-flight and will be carded here when it measures. Ecosystem corpus = own planes, docs and evidence ledgers with revision provenance in-plane (hist- rows). Engine-seat training data: see the seat vendor's own model card.

5 · Evaluations

Reasoning

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
HLE text onlypct29.726.629.435.940.135.944.753.347.2UNMEASURED:F786-
HLE with toolspct46.037.450.254.054.748.251.464.555.0UNMEASURED:F786-
AIME 2026pct97.194.295.896.499.296.798.399.999.9UNMEASURED:F786-
GPQA Diamondpct87.286.787.991.189.588.894.192.694.1UNMEASURED:F786-

Agentic · coding

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
SWE-bench Verifiedpct77.670.776.880.280.080.680.695.082.271 ANALOG live*contamination-zeroed by TML. NISHI cell = sovereign NishiLang foreign-bug analog (fresh compile+run judge), NOT the 500-Verified set, no agent-parity. HARNESS READINESS (measured 2026-07-20): the public 500-instance Verified set is ingested sovereignly (revision-pinned c104f840, swebvinst- index) and its full test contract extracted (swebvtest-: 1516 decisive FAIL_TO_PASS tests, 60235 PASS_TO_PASS regression tests, 0 truncated); sovereign grader nx_swebv_grade holds the denominator+scoring (ungraded=unresolved). Remaining for a head-to-head number: the pytest ORACLE leg (F806) + fix-loop attempts - oracle-only law: pytest runs the tests, Nishi selects and scores
SWE-bench Pro Publicpct54.346.450.758.662.155.454.280.064.6UNMEASURED:F787-
Terminal Bench 2.1pct63.856.451.371.382.76473.884.689.5UNMEASURED:F787*contamination-zeroed by TML; best harness

Agentic · general

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
GDPVal-AA v2score12381164100911901514130796217601748UNMEASURED:F788-
MCP Atlaspct74.144.764.068.177.873.278.283.381.8100 ANALOG fleet-exec-v1NISHI = fleet-EXECUTION analog v1 (epoch 1784569900): 6-task curated battery over the own MCP fleet via nx_plan_run (store round-trip, edge fetch, board emit, measured-status read, card render, janitor round-trip), pass = all steps rc0, 6/6, evidence planrun-mcpbt01..06; NOT agentic tool-CHOICE and NOT comparable to the MCP Atlas column - seat-driver + adversarial battery = F788 remainder
Tau 3 Bankingpct23.713.814.220.626.825.816.526.833.0UNMEASURED:F788-
Toolathlon Verifiedpct45.534.333.058.059.955.961.176.473.1UNMEASURED:F788-
BrowseComp w/ ctx mgmtpct77.1-74.983.2-83.485.988.090.84UNMEASURED:F788sovereign browser lane = future harness seat

Factuality

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
SimpleQA Verifiedpct43.932.436.938.738.157.077.368.371.6UNMEASURED:F789nx_recall_bench/BEIR planes = harness seed
AA Omnisciencescore2.1-1.0-8.06.04.0-10.033.040.022.0UNMEASURED:F789hallucination-weighted index, can be negative

Chat

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
IFBenchpct79.881.470.276.073.376.577.163.572.7UNMEASURED:F786-
Global-MMLU-Litepct88.785.684.088.489.289.392.793.391.8UNMEASURED:F786-

Vision

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
MMMU Pro Standard 10pct73.5-75.079.0--82.084.283.0UNMEASURED:F791-
Charxiv RQpct78.1-77.580.4--80.286.584.7UNMEASURED:F791-
Charxiv RQ with pythonpct82.0-78.786.7--89.989.487.8UNMEASURED:F791TML internal python harness

Audio

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
Audio MCpct56.6-----66.8--UNMEASURED:F791TML multiple-choice harness
MMAUpct77.2-----82.5--UNMEASURED:F791-
VoiceBenchpct91.4-----94.3--UNMEASURED:F791-

Safety

benchmarkunitInklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (max/xhigh)NISHInotes
FORTRESS Adversarialpct78.077.654.165.671.336.065.296.082.4UNMEASURED:F790-
FORTRESS Benignpct95.990.598.397.290.098.598.055.198.1UNMEASURED:F790benign-answer rate, higher = fewer over-refusals
StrongREJECTpct98.698.799.599.898.598.698.098.798.5UNMEASURED:F790structural ocap/never-brick layer = EXCEED by construction, behavioral half unmeasured

6 · Safety

Structural safety precedes behavioral safety: rule-26 never-brick (no persistent-hardware writes by construction, proven mechanically not promised), ocap capability tokens (least-authority, audited, revocable), deny-by-construction write paths (secrets, OS namespace and the tool allowlist are unreachable even with a valid write cap), additive-only data (history is sacred). Behavioral refusal batteries (FORTRESS/StrongREJECT-class) over the seat+cap layer: UNMEASURED - filed F790.

7 · Bias, Risks and Limitations

Honest-numbers law: every NISHI figure traces to a live artifact; UNMEASURED is a first-class value, never hidden. Known limits: single-operator triangulation, English-only corpus, no media perception (served not perceived), engine-seat dependence for reasoning, and benchmark analogs are NOT head-to-head comparable until the parity harnesses F786-F791 land. Frontier-model figures are transcribed, not reproduced.

8 · Legal

Sovereign private system, all rights reserved. Competitor figures are the property of their publishers, transcribed from the Thinking Machines Inkling model card (thinkingmachines.ai/model-card/inkling, retrieved 2026-07-20) for capability comparison.

Provenance

Inkling card fetched 2026-07-20. LIVE-DERIVED cells (re-read at every emit, never frozen): SWE-bench-analog resolve_pct from sites/nishifamily/compare/autograde/api.json; MCP primitive count from knowledge/store/prim_registry.json. Plane knowledge/store/modelcard- carries hist- provenance for every edit; renderer nx_model_card (F785). FALSIFIABLE AUTOMATION CLAIM: this page claims a 6h self-refresh (clock_jobs row modelcard/21600/nx_model_card.elf) - audit it at knowledge/status/modelcard_beat.log, which the renderer appends an epoch to on EVERY emit. Entries ~21600s apart with no operator session running = the beat is autonomous; a gap = it is not. As of 2026-07-20 the row is registered and its clock stamp has advanced, but an unattended fire is NOT YET OBSERVED - so the claim is stated as auditable, not proven.

envelope: plane cap 1MiB, out cap 512KiB, 16-col rows; plane content operator-trusted (no HTML-escape pass); the live cell = resolve_pct parsed from /compare/autograde/api.json at emit time; regenerated on the freshness beat.