Kimi-K3-class Model — sovereign no-float SOTA climb

Can we build a Kimi-K3-class model bits-up and be state of the art? The disciplined answer is measure-first, then close gaps with gate-proven capability — never assert. Coverage is liar-killed: a HAVE claim with no shipped gate is forced to GAP.

Live from the nx_k3_arch_census MCP tool 2026-07-19 · every tick below is a committed gate, not a number borrowed.

K3-class architecture coverage — measured

312‰K3-class coverage (15/48) — live-census raw; 9 gates shipped today
1STRONG — serve+train both gate-proven (Muon)
7PARTIAL (serve-proven; training the next axis)
2GAP real (1M-ctx, multimodal)
2dedup-pending dup rows → debt D024
0liar-killed (every claim has a shipped gate)

Honest-census note. The raw permille (312) is conservative: the k3feat- store currently holds 2 superseded duplicate rows from a cross-session id mismatch (this session graded muon/mxfp4 before discovering the canonical ids are perheadmuon/mxfp). nx_store_put is additive with no hard-delete, so the dupes are neutralized to 0/0 but still sit in the denominator — dedup debt D024 filed (needs a prune verb + a canonical id registry). Consolidated across the 10 canonical K3 innovations, coverage is 15/40 = 375‰. We publish the tool's raw number for verifiability — query nx_k3_arch_census yourself.

Per-innovation coverage

K3 innovationcoverageevidence / mechanism
Stable LatentMoE + Quantile BalancingPARTIALOLMoE 64-expert no-float serve rungs 1-5 + MoE gradcheck shipped
Gated MLA (latent KV attention)PARTIALnx_nofloat_mla 5/5, 8× lossless KV-compress; gating variant next
SiTU activationPARTIALfx_sigmoid + fx_tanh primitives gate-proven; compose next
MXFP4 / MXFP8 weights+actsHAVE ✅CLOSED serve 07-19: nx_nofloat_mxfp4 6/6 (E2M1 weights) + nx_nofloat_mxfp8 6/6 (E4M3/E5M2 acts) — OCP MX v1.0 full, E8M0 block scale, pure-integer exact decode, NaN/inf guarded. QAT training next.
Per-Head Muon optimizerSTRONG ✅FULL LOOP CLOSED 07-19nx_nofloat_muon gate 6/6: momentum + NS-orthogonalized update applied exactly, per-head independent, BIT-EXACT deterministic (float Muon drifts). The first K3 feature with a proven TRAIN primitive (serve+train). Full convergence run rides the GPU seat (F101). Muon = Keller Jordan/Moonshot
Kimi Delta Attention (linear)HAVE ✅SERVE CLOSED 07-19: nx_nofloat_kda gate 5/5 — delta-rule OVERWRITE memory (same-key rewrite → v2, not v1+v2) + fine-grained per-channel gating (the KDA-vs-GLA distinction) + O(1) linear kernel, all bit-exact. Chunked-parallel TRAIN form next. arXiv 2510.26692, 3:1 KDA:MLA
Attention ResidualsPARTIAL ✅CROSS-LAYER CLOSED 07-19nx_nofloat_k3stack gate 4/4: content-weighted attention over preceding layers, proven depth-stable + bit-exact in a 4-layer stack. Learned weighting next. arXiv 2603.15031
KDA:MLA 3:1 interleave (K3 block)HAVE ✅CAPSTONE 07-19nx_nofloat_k3interleave gate 6/6: the K3 block ASSEMBLED — 3 KDA (delta+gating) : 1 MLA (latent-KV) layers, both paths live, depth-stable, bit-exact, schedule mechanically 3:1. Composed from the proven KDA + MLA organs. arXiv 2510.26692
1M context (long-ctx RoPE)GAPYaRN, iRoPE, DeepSeek 1M, NTK-aware scaling
Native multimodalGAPunified MM tokenizer, integer vision tower, video token compression

The novel exceed — beyond SOTA, not just matching

Two of today's landings are things a float training/inference stack cannot reproduce bit-for-bit, because float accumulation is order-dependent:

The assembly blueprint (grounded)

K3 = DeepSeek-V3 / Moonlight MoE Transformer + KDA : MLA interleaved 3:1 + MoE FFN + Attention-Residuals on the residual path. We already hold MLA (8× lossless), MoE (OLMoE), and now MXFP4 + the Muon NS core — so the F-4xx assembly is 3 KDA layers : 1 MLA layer per block. The 3:1 interleaved block composes todaynx_nofloat_k3interleave gate 6/6 runs 3 KDA : 1 MLA layers as ONE deterministic no-float K3 block (both attention paths live, depth-stable, bit-exact) — the logically-integrated K3 skeleton, assembled from the proven KDA + MLA organs. A standing plan-k3verify workflow re-verifies coverage on demand (supervised). Remaining new primitives: the KDA chunked-parallel training form, MXFP8, 1M-ctx RoPE, multimodal — KDA serve, AttnRes stacking, the 3:1 KDA:MLA interleave, and the full Per-Head Muon optimizer loop are now CLOSED.

Honest scope

This is architecture coverage, measured — not a trained 2.8T model. K3 needs frontier compute to train; a real sovereign K3-μ needs the GPU seat (F101). What is true today: the K3 architecture class runs, gate-by-gate, in 100% integer no-float, deterministic — and the remaining gaps have exact mechanisms + arXiv references, not vibes. Research targets: R007 KDA, R008 AttnRes, R009 Muon(full), R010 MXFP8, R011 1M-ctx, R012 multimodal. Query it yourself: the nx_k3_arch_census MCP tool returns this live.

Sovereign page — served bits-up over TLS 1.3. Coverage from the nx_k3_arch_census MCP tool; every PARTIAL/HAVE backed by a committed gate; liar-killed.