Nishi Family › Compare › Autonomous Coding & Code-LLM Harness
Nishi Compare · measured, not asserted
Autonomous Coding & Code-LLM Harness
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi forge vs Cursor and SWE-agent and Qwen2.5-Coder and Aider
Overview
The purpose, declared position and evidence coverage of this domain. Source presence and completed acceptance are different measures.
Where we are. Nishi forge is a real sovereign coding harness: organ-emitted verified context pack, a DETERMINISTIC build plus byte-exact judge (exceed: the verify signal is noise-free), a bounded repair loop on verbatim compiler stderr, a self-retiring vetted-rules registry (exceed), local no-float Coder 0.5B and 1.5B tiers loaded bits-up (exceed), a liar-killed requirements dossier, ChatML prompting, sovereign semantic code retrieval with an externally-validated integer embedding path. It is NOT at frontier capability: retrieval is base-weight quality, prefill is per-token, no best-of-N, no grammar-constrained decoding, single-file scope, hand-authored oracles, and zero external-leaderboard results.
Where we need to go. Stand on a public leaderboard with a fully sovereign stack, then climb it by the levers the record names: batched prefill for throughput, best-of-N against the free verifier, grammar masking into NishiLang, a contrastive retrieval tune on our own verified ledger, multi-file repair, and generated oracles -- each proven by the byte-exact judge, never by an LLM grading an LLM.
Research bar. Aider-Polyglot is measured on 225-task multi-language edit benchmark, scaffold-sensitive (little-coder 2026: 19 to 45 percent by scaffold alone). Theirs: public pass-rate leaderboard. Ours: CD6 publishes our number on this page.
Research bar. SWE-bench is measured on repository-level issue resolution. Theirs: the repo-level frontier. Ours: CD5 is the parity rung; SWE-bench-Lite is the stretch.
Research bar. Large Language Monkeys (arXiv 2407.21787) is measured on coverage scales with samples given a verifier (15.9 to 56 percent at 250 samples). Theirs: best-of-N with a verifier. Ours: CD2 -- our verifier is perfect and free.
Latest recorded release
No valid dated release entry is recorded for this domain.
Release entries describe recorded changes; they do not establish that every capability passed evaluation.
11 of 18 capabilities measured|5 of them measured exceeds|7 open|coverage 611/1000|adoption 10 full / 1 partial
Evidence profile — what the gaps on this board actually are
Measured by nx_swcompare_evidence, read back by nx_evprofile_lib. Every figure is a count with its denominator — there is deliberately no score, no grade and no percentage anywhere in this band, because a stored scalar is a field a seat can edit and a counted partition is not.
evidence|grounded 11/11|unsupported 0|gates green 1/1|proven able to fail 1/1|never bitten 0|green at 0/0 0|open gaps 7|of them unnamed 0|of them proof withheld 0|flips ready 0
proven able to fail counts the gates that have a RECORDED RED — nx_gate_bite mutated the gate subject, rebuilt it, watched the gate go red, and that record is inside the shared TTL. never bitten is its complement over the same denominator: those gates ran and were green, and nothing has ever shown them able to detect anything, so their green is a statement about this run and not about the gate. green at 0/0 is a separate and much weaker observation — the gate printed GREEN on a zero denominator, so its own tooth counter says it examined nothing. A gate can be green, non-zero, and still never bitten; that is the common case and it is now visible instead of implied.
partition: grounded + unsupported = 11 vs present 11 · named + unnamed + withheld = 7 vs open 7 · both reconcile
liar-kill conj=GPQN · all four conjuncts held
graded document: BUILDROOT tree, 8645 bytes · gates map: PRIMARY · stamped 2d 7h ago · source ../knowledge/status/evstamp_coding.verdict
| Gap class | What it is, and the work it names |
|---|---|
| NONE | No gap class fires on this board: every published claim is grounded, the declared gates ran and were green, no watch contract is unnamed or already landed, and the stamp is inside its TTL. That is a statement about evidence, not about depth, scale or polish. |
nx_swcompare_evidence coding itself and are not carried on the stamp, so this page names the classes and the producer names the rows. That split is stated rather than hidden: a count without a worklist is not actionable, and this band is honest about which half of that it is.Production map
Follow the dependencies, declared acceptance criteria and recorded priorities. Inspect source binding before treating a rank as executable work.
Ranking source binding: PLAN_MATRIX_BOUND_ONLY. Recorded priorities require current acceptance evidence and resource checks before execution.
Ranking matches the captured plan and matrix only. Latest execution outcome, research freshness, accepted delivery and investment return are unverified.
Recorded priority estimates
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: listed before new work by this heuristic. Priority is not measured delivery cost or execution readiness. Stamp: # asof=1789541549 domain=coding target_version=0.1 rungs=11 done=4 open=7 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | 0.1 | Aider-Polyglot runner (CD6) fg_polyglot_bench | 9600 | v=16 m=3 c=5 |
| #2 | later | Batched prefill (CD1) nsv_prefill_batch | 3714 | v=13 m=2 c=7 |
| #3 | later | Multi-file editing (CD5) fg_multi_file | 3428 | v=12 m=2 c=7 |
| #4 | later | Best-of-N against the verifier (CD2) fg_best_of_n | 3000 | v=3 m=2 c=2 |
| #5 | later | Oracle generation (CD7) fg_oracle_gen | 2400 | v=6 m=2 c=5 |
| #6 | later | Contrastive retrieval tune (CD4) ci_train_contrastive | 900 | v=9 m=2 c=20 |
| #7 | later | Grammar-constrained decoding (CD3) nsv_grammar_mask | 857 | v=3 m=2 c=7 |
Declared roadmap — contract, acceptance, executor, effort
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Byte-exact judge (CB0) | fg_run | Deterministic build plus byte-exact judge -- LANDED, the verifier every open rung leans on | Organ | 0 u |
| Bounded repair loop (CB1) after CB0 | fg_bank | Verbatim compiler stderr fed back, gate 5/5 -- LANDED | Organ | 0 u |
| Local model tier (CB2) | load_named_q16 | Coder 0.5B and 1.5B q8_0 through the no-float stack -- LANDED | Organ | 0 u |
| Semantic code retrieval (CB3) | ci_query_top1 | nx_code_index, jina recipe, 5-tooth gate -- LANDED at base-weight quality | Organ | 0 u |
| Batched prefill (CD1) after CB2 | nsv_prefill_batch | Prefill runs the prompt as ONE batched GEMM (m greater than 1) instead of m looped m=1 matmuls; pass-fail tooth: per-token prefill cost must FALL with prompt length on nx_embed_bench at m=8,32,64,125 -- flat means we did not enter the batched regime and must not claim it | Organ | 1.5 u |
| Best-of-N against the verifier (CD2) after CB0,CD1 | fg_best_of_n | Sample N candidates per task, run each through the byte-exact judge, keep the first GREEN; gate proves pass@N is monotone in N on the L4 mini-pack and that a task with no GREEN in N reports NONE rather than the best-scoring wrong answer | Organ | 1 u |
| Grammar-constrained decoding (CD3) after CB2 | nsv_grammar_mask | The serve masks logits to the NishiLang grammar's legal next tokens; gate proves every sampled program at least PARSES and a control with the mask off reproduces the current parse-failure rate | Organ | 1.5 u |
| Contrastive retrieval tune (CD4) after CB3 | ci_train_contrastive | Train the embedding head on the forge's own compile-verified query-answer pairs through the no-float autograd; done-rule: top-1 on the held-out ruler EXCEEDS the 3-of-6 base-weight ratchet, declared before the run, or it does not ship | Organ | 2 u |
| Multi-file editing (CD5) after CB1 | fg_multi_file | A task may name several organs; the repair loop applies a coordinated patch set and the judge builds the whole closure; gate proves a two-file rename task lands and a patch that breaks a third importer is refused by the build | Organ | 3 u |
| Aider-Polyglot runner (CD6) after CB0,CB2 | fg_polyglot_bench | A sovereign runner over the public task set using the byte-exact judge, publishing pass rate per language with the task count; the first number is whatever it is -- an honest standing beats a claimed one | Organ | 2 u |
| Oracle generation (CD7) after CB0 | fg_oracle_gen | Generate candidate test oracles for a task and ACCEPT one only when it discriminates a known-good from a known-bad solution (anti-vacuity by construction); gate proves a generated oracle fails the planted mutant and passes the reference | Organ | 2 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M1 · Batched and sampled | CD1,CD2 | 4.5 u |
| M2 · On a public leaderboard | CD6 | 2 u |
| M3 · Frontier-shaped harness | CD3,CD4,CD5,CD7 | 13.5 u |
Inspect a rung and its prerequisites
Declared nodes 11. Rank input binding: PLAN_MATRIX_BOUND_ONLY. Dependency order is authored. Implementation, acceptance evidence, authority and resource readiness are unverified. No action is recommended or dispatched here.
Plan SHA-256 03f6c43c9254afc5dfeb99d984bf5032049fbd99cd45aa19d0ec85961eef93f4. Target rows 0; role rows 0. Existing risks and release worklog retain their own scope; no node completion is inferred.
Use Enter or Space on a rung to inspect its contract. Prerequisite links locate another rung in this list; open its summary to inspect it. Estimates are authored effort, not forecasts.
CB0 — Byte-exact judge
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
fg_runAcceptance: Deterministic build plus byte-exact judge -- LANDED, the verifier every open rung leans on
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CB1 — Bounded repair loop
Prerequisites: CB0 (acceptance unverified)
Contract:
fg_bankAcceptance: Verbatim compiler stderr fed back, gate 5/5 -- LANDED
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CB2 — Local model tier
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
load_named_q16Acceptance: Coder 0.5B and 1.5B q8_0 through the no-float stack -- LANDED
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CB3 — Semantic code retrieval
Prerequisites: None declared; this does not establish execution eligibility.
Contract:
ci_query_top1Acceptance: nx_code_index, jina recipe, 5-tooth gate -- LANDED at base-weight quality
Authored effort: 0. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD1 — Batched prefill
Prerequisites: CB2 (acceptance unverified)
Contract:
nsv_prefill_batchAcceptance: Prefill runs the prompt as ONE batched GEMM (m greater than 1) instead of m looped m=1 matmuls; pass-fail tooth: per-token prefill cost must FALL with prompt length on nx_embed_bench at m=8,32,64,125 -- flat means we did not enter the batched regime and must not claim it
Authored effort: 1.5. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD2 — Best-of-N against the verifier
Prerequisites: CB0 (acceptance unverified), CD1 (acceptance unverified)
Contract:
fg_best_of_nAcceptance: Sample N candidates per task, run each through the byte-exact judge, keep the first GREEN; gate proves pass@N is monotone in N on the L4 mini-pack and that a task with no GREEN in N reports NONE rather than the best-scoring wrong answer
Authored effort: 1. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD3 — Grammar-constrained decoding
Prerequisites: CB2 (acceptance unverified)
Contract:
nsv_grammar_maskAcceptance: The serve masks logits to the NishiLang grammar's legal next tokens; gate proves every sampled program at least PARSES and a control with the mask off reproduces the current parse-failure rate
Authored effort: 1.5. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD4 — Contrastive retrieval tune
Prerequisites: CB3 (acceptance unverified)
Contract:
ci_train_contrastiveAcceptance: Train the embedding head on the forge's own compile-verified query-answer pairs through the no-float autograd; done-rule: top-1 on the held-out ruler EXCEEDS the 3-of-6 base-weight ratchet, declared before the run, or it does not ship
Authored effort: 2. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD5 — Multi-file editing
Prerequisites: CB1 (acceptance unverified)
Contract:
fg_multi_fileAcceptance: A task may name several organs; the repair loop applies a coordinated patch set and the judge builds the whole closure; gate proves a two-file rename task lands and a patch that breaks a third importer is refused by the build
Authored effort: 3. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD6 — Aider-Polyglot runner
Prerequisites: CB0 (acceptance unverified), CB2 (acceptance unverified)
Contract:
fg_polyglot_benchAcceptance: A sovereign runner over the public task set using the byte-exact judge, publishing pass rate per language with the task count; the first number is whatever it is -- an honest standing beats a claimed one
Authored effort: 2. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
CD7 — Oracle generation
Prerequisites: CB0 (acceptance unverified)
Contract:
fg_oracle_genAcceptance: Generate candidate test oracles for a task and ACCEPT one only when it discriminates a known-good from a known-bad solution (anti-vacuity by construction); gate proves a generated oracle fails the planted mutant and passes the reference
Authored effort: 2. Executor kind: Organ. Responsible, accountable and verifier not established. Inspect retained worklog.
Learning and practice paths
No structured learning path is declared for this plan. Existing research, roadmap and worklog remain available above.
Capability comparisons
Compare the field, search individual capabilities and open their source and adoption evidence. Documented presence does not establish comparative quality.
Position map — centrality and distinctiveness
The four-quadrant map the field uses for brand strategy (Dawar and Bagga, HBR June 2015), re-derived from this matrix on every publish. Centrality is the share of the category's feature mass a player covers, each feature weighted by how many hold it; distinctiveness is the average lead over each rival on the rows the player holds; breadth is the depth-weighted share of the whole matrix (the bubble); depth is how deeply the rows held are held; momentum is the day-over-day move off the spine (green rising, red falling, grey until day two); the dashed path runs first day → previous day → today. Dividers are the category means. Axes are fitted to the field of play, so read the tick numerals, not the frame. Rival marks are documented presence, so a rival's position reads the record, never its quality. The picture grades its own readability below; the table beside it is the same data for a screen reader or a second method.
knowledge/compare/coding.cdmap (cut|x|y, cube|x|y|z); absent = the HBR defaultsbubble area = breadth · ring = momentum (green rising, red falling, grey until day two) · dashed = category means · axes fitted to the field of play: centrality 500–1000, distinctiveness 100–600 of 0–1000 permil (the full range put every player in one corner)
readability of the 2D cut, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 0 (a fact of the data: two players that close are that close) · data spread 620 permil of the plot · quadrant words unseated 0
text contrast, measured with wcag2-ratio (floors from contrast.conf), light theme: labels 16.24 (floor 4.50) · notes 16.24 (floor 4.50) · quadrant words 5.89 (floor 4.50) · tick numerals 5.89 (floor 4.50) · axis titles 5.59 (floor 4.50) · dark theme: labels 13.78 (floor 4.50) · quadrant words 5.55 (floor 4.50) · axis titles 5.97 (floor 4.50) · dark classes under their floor 0 (one figure serves both themes: a dark shortfall is a token to fix, never a class to hide) · classes refused under their floor 0 (a refused class is not drawn; the scale classes are measured, never hidden) · export: SVG PNG (receipt, rendered by the estate's own rasteriser from this page)
readability of the 3D cube, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 0 (a fact of the data: two players that close are that close) · data spread 409 permil of the plot
10 panels over 5 registered axes, every pair once (the lower axis on x, the higher on y) · each panel fitted to its own field of play, first and last tick numerals shown · no labels in a small cut, the legend names the colours; the grade below reports the mark terms only (MM summed over the panels, spread averaged)
readability of the small multiples, self-graded by the layout ruler: label overlaps 0 · labels over marks 0 · off-canvas 0 · unresolved labels 0 · mark overlaps 1 (a fact of the data: two players that close are that close) · data spread 486 permil of the plot
| Player | Quadrant | centrality | distinctiveness | breadth | depth | momentum | Rows held | Days on spine | First seen |
|---|---|---|---|---|---|---|---|---|---|
| Nishi | Unconventional | 573 | 477 | 500 | 818 | 0 since 2026-09-15 | 11 | 2 | 2026-09-15 |
| Cursor | Mainstream | 868 | 243 | 574 | 794 | 0 since 2026-09-15 | 13 | 2 | 2026-09-15 |
| SWE-agent | Mainstream | 868 | 205 | 537 | 743 | 0 since 2026-09-15 | 13 | 2 | 2026-09-15 |
| Qwen2.5-Coder | Unconventional | 721 | 340 | 555 | 833 | 0 since 2026-09-15 | 12 | 2 | 2026-09-15 |
| Aider | Mainstream | 819 | 152 | 462 | 694 | 0 since 2026-09-15 | 12 | 2 | 2026-09-15 |
| Day | Matrix rows reviewed | Players recorded |
|---|---|---|
| 2026-09-15 | 18 | 5 |
| 2026-09-16 | 18 | 5 |
players 5|matrix rows 18|feature mass 61|centrality mean 769|distinctiveness mean 283|axes 5|spine days 2 (shown 2)|rows written today 0|readability defects 2D 0 cube 0|spine knowledge/status/cdmap/coding.spine
comparewatch- plane row flips with it. A dark tag means the organ file EXISTS but does not declare the contracted symbol: something shipped there under another name, and until the contract is repointed to the real entry point (the plan rung and this row) or the function is renamed, that capability is invisible to this board — a build lost to darkness, named so it is not. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1789497476, gate census asof 1789498715 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | Cursor | SWE-agent | Qwen2.5-Coder | Aider |
|---|---|---|---|---|---|
Verified context pack (organ-emitted, regenerable)Measured:fc_emit exists in runtime/nx_forge_ctx.nx, verified at emit. Cursor codebase-index [cursor-indexing] + Aider repo-map [aider-repomap] are the bar; ours is vetted+auto-regenerated from sources (gate 6/6) Adoption: LIB-WIRED importers=7 nonval=4 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; ahead of Qwen2.5-Coder; fully adopted on the estate ladderCons behind Cursor (leads) | ● | ◉ | ● | ○ | ● |
Deterministic build + byte-exact judgeMeasured exceed:fg_run in runtime/nx_forge.nx, verified at emit. EXCEED: integer toolchain -> GREEN is a PROVABLE equality, noise-free repair signal; all named tools rely on flaky test suites [sweagent24] Adoption: LIB-WIRED importers=8 nonval=5 — fully adopted (top of its ladder).Pros Nishi leads, a measured exceed; ahead of Cursor, SWE-agent, Qwen2.5-Coder, Aider; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ◉ | ◐ | ◐ | ○ | ◐ |
Bounded repair loop (verbatim compiler stderr)Measured:fg_bank exists in runtime/nx_forge.nx, verified at emit. ACT+Debugger/Reflexion/Self-Debug [selfdebug23] are the field (57->65 HumanEval); ours gate-proven 5/5 with a real parse-trap repair Adoption: LIB-WIRED importers=8 nonval=5 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; ahead of Qwen2.5-Coder; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ● | ● | ● | ○ | ● |
Self-retiring vetted-rules registryMeasured exceed:pv_run_witness in runtime/nx_pack_rules_vet_gate.nx, verified at emit. EXCEED: each language restriction carries a witness probe + measured WHY, auto-RETIRES when the compiler heals it; no prompt tool measures its own rules Adoption: GATE:BUILT-UNPROMOTED trial=- — PARTIAL: compiled, never promoted to the serving root: /api/promote it, then roster it.Pros Nishi leads, a measured exceed; ahead of Cursor, SWE-agent, Qwen2.5-Coder, AiderCons not adopted yet: compiled, never promoted to the serving root: /api/promote it, then roster it | ◉ | ○ | ○ | ○ | ○ |
Model backend (serve /gen + fenced extraction)Measured:fm_post_gen exists in runtime/nx_forge_model.nx, verified at emit. Table stakes; ours proven over a live loopback round-trip + model->candidate->judge GREEN rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind them Adoption: LIB-WIRED importers=3 nonval=2 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ● | ● | ● | ● | ● |
Local sovereign model tier loaded bits-upMeasured exceed:load_named_q16 in runtime/nx_nofloat_llm.nx, verified at emit. EXCEED (sovereignty): Coder-0.5B AND Coder-1.5B q8_0 (sha=ETag verified) [qwen25coder24] both run through our own no-float integer stack via the arch-config core, deterministic; 1.5B measured 8-10/10 on the inversion ruler (0.5B 1/10) -- scale is the capability lever Adoption: LIB-WIRED importers=33 nonval=8 — fully adopted (top of its ladder).Pros Nishi leads, a measured exceed; ahead of Cursor, SWE-agent, Qwen2.5-Coder, Aider; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ◉ | ○ | ○ | ◐ | ○ |
Requirements dossier evidence-forked + liar-killedMeasured exceed:rq_row in runtime/nx_forge_req_gate.nx, verified at emit. EXCEED: 14 requirement rows each fork their evidence on disk + a negctl; the census grades itself, cheat-proof by construction Adoption: GATE:LIVE trial=- — fully adopted (top of its ladder).Pros Nishi leads, a measured exceed; ahead of Cursor, SWE-agent, Qwen2.5-Coder, Aider; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ◉ | ○ | ○ | ○ | ○ |
Best-of-N sampling against the free verifierDARK —runtime/nx_forge.nx EXISTS but does not declare fg_best_of_n: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. Large-Language-Monkeys [llmonkeys24]: 15.9->56% SWE-bench-Lite [swebench23] at 250 samples WITH a verifier; ours is perfect+free but the loop is not built yetPros none measured yetCons behind Cursor, SWE-agent, Aider; open contract, nothing on disk yet | ○ | ● | ● | ○ | ● |
Chat-template / instruct promptingMeasured:nsv_chatml_ids exists in runtime/nx_nofloat_serve_core.nx, verified at emit. Qwen chat template is native to the model; ours PROVEN 07-15 (ChatML special-token splice into the no-float serve, byte-identical, L4 ran with it; SSE slot-collision fixed) rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind them Adoption: LIB-WIRED importers=16 nonval=9 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; fully adopted on the estate ladderCons behind Qwen2.5-Coder (leads) | ● | ● | ● | ◉ | ● |
Batched prefill (throughput)DARK —runtime/nx_nofloat_serve_core.nx EXISTS but does not declare nsv_prefill_batch: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. Per-token prefill; every serving stack batches -- our N6 keystone. NOTE 07-15: a SEPARATE wall (tokenizer O(n^2) linear-scan of 151k merges/pair, ~40min on a 4KB prompt) was ROOT-FIXED to hash+incremental-rank BYTE-EXACT; batched prefill itself still absent rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind themPros none measured yetCons behind Cursor (leads), SWE-agent (leads), Qwen2.5-Coder (leads), Aider; open contract, nothing on disk yet | ○ | ◉ | ◉ | ◉ | ● |
Grammar-constrained decoding (logit mask)DARK —runtime/nx_nofloat_serve_core.nx EXISTS but does not declare nsv_grammar_mask: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. Outlines/XGrammar [xgrammar24] (100x, near-zero overhead); sovereign-exclusive for us since only our OWN serve can mask logits to the NishiLang grammarPros none measured yetCons behind Cursor, SWE-agent, Qwen2.5-Coder; open contract, nothing on disk yet | ○ | ● | ● | ● | ○ |
Own-model training (compiler-filtered data)Measured exceed:nfa_backward in runtime/nx_nofloat_autograd.nx, verified at emit. EXCEED (determinism): 100% integer no-float train+infer is bit-exact reproducible; HONEST scale = toy (modelwright 476); the forge ledger is the compile-verified training set rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind them Adoption: LIB-WIRED importers=47 nonval=7 — fully adopted (top of its ladder).Pros Nishi leads, a measured exceed; ahead of Cursor, SWE-agent, Aider; fully adopted on the estate ladderCons none on the measured axes (rival marks are documented presence, not depth) | ◉ | ○ | ○ | ◉ | ○ |
Code retrieval + reranking (CoREB)Measured:ci_query_top1 exists in runtime/nx_code_index.nx, verified at emit. NOW HAVE 07-15: nx_code_index semantic organ retrieval (jina-recipe arXiv:2508.21290 [jinacode25], NXCI1 sovereign store, gated 5-tooth incl cross-request determinism + corrupt-store-refused); base-weight quality MEASURED 3/6 top-1 (ratchet, behind trained rerankers); e5 contrastive tune (llm2vec recipe) = the quality closer Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; fully adopted on the estate ladderCons behind Cursor (leads), Qwen2.5-Coder (leads) | ● | ◉ | ● | ◉ | ● |
Sovereign code embeddings, external-oracle validatedMeasured:nsv_embed exists in runtime/nx_nofloat_serve_core.nx, verified at emit. NEW 07-15: our 100%-integer last-token-pooled embed path RAN jina-code-embeddings-0.5b [jinacode25] (Qwen2.5-Coder backbone [qwen25coder24], cc-by-nc ORACLE) and REPRODUCED its bf16 2x2 signature 822/176/122/617 vs oracle 817/124/120/553 pm (matched cells within <=5pm; tuned margins 646/495 vs base 72/49) -- a sovereign-stack result no API embedder exposes; Qwen col=2 (jina fine-tuned ON that backbone = the bar) Adoption: LIB-WIRED importers=16 nonval=9 — fully adopted (top of its ladder).Pros Nishi has it, measured on disk; ahead of Cursor, SWE-agent, Aider; fully adopted on the estate ladderCons behind Qwen2.5-Coder (leads) | ● | ○ | ○ | ◉ | ○ |
Repository-level / multi-file editingDARK —runtime/nx_forge.nx EXISTS but does not declare fg_multi_file: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. SWE-bench issue resolution [swebench23] is the frontier; forge is single-file organ scope todayPros none measured yetCons behind Cursor (leads), SWE-agent (leads), Qwen2.5-Coder, Aider (leads); open contract, nothing on disk yet | ○ | ◉ | ◉ | ● | ◉ |
Test / oracle generationDARK —runtime/nx_forge.nx EXISTS but does not declare fg_oracle_gen: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. The field generates tests/oracles; ours are hand-authored gates -- the generator is a future rung rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind themPros none measured yetCons behind Cursor, SWE-agent (leads), Qwen2.5-Coder, Aider; open contract, nothing on disk yet | ○ | ● | ◉ | ● | ● |
Retrieval quality closer (contrastive tune on the compile-verified ledger)DARK —runtime/nx_code_index.nx EXISTS but does not declare ci_train_contrastive: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. The retrieval row above is PRESENT at base-weight quality (3 of 6 top-1, a ratchet); trained rerankers lead -- the named closer is the e5/llm2vec-style contrastive tune over the forge's own verified pairs, through the no-float autograd, so the ruler moves by measurement not by swapping in a foreign model rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind themPros none measured yetCons behind Cursor (leads), SWE-agent, Qwen2.5-Coder (leads), Aider; open contract, nothing on disk yet | ○ | ◉ | ● | ◉ | ● |
External leaderboard standing (Aider-Polyglot run)DARK —runtime/nx_forge.nx EXISTS but does not declare fg_polyglot_bench: something shipped at this path under another name, and this mark would read open forever. Repoint the contract to the real entry point (the plan rung AND this row) or rename the function; the flip follows on the next beat, and the comparewatch- plane row reads DARK until then. HONEST STANDING: zero external-suite results today; little-coder (2026) proved sub-10B is benchmarkable on Aider-Polyglot (19 to 45 percent by scaffold alone, N=225) -- the contract is a sovereign runner over the public task set with the byte-exact judge, publishing pass rates the field can compare rival marks uncited — the Yes, Best or Part codes on this row are an observation read with no reference mark behind themPros none measured yetCons behind Cursor (leads), SWE-agent (leads), Qwen2.5-Coder (leads), Aider (leads); open contract, nothing on disk yet | ○ | ◉ | ◉ | ◉ | ◉ |
Delivery and evidence
Inspect risks, technical debt, rendered observations, experiments and references. Read scope and limitations alongside every result.
Risk register
| Risk | Likelihood x impact | Mitigation |
|---|---|---|
| Best-of-N on a model that pattern-copies (0.5B) buys nothing -- scale is the lever the record already measured | likely x medium | CD2 reports per-tier; the 1.5B tier (8-10 of 10 on the inversion ruler) is the default subject. |
| A retrieval tune that improves the training ruler and not the held-out one is overfitting wearing a win | possible x high | CD4 done-rule is pre-declared on the held-out ruler; a miss stays UNWIRED as the BEIR arm did. |
| A leaderboard number taken once under a favourable scaffold is a sample, not a standing | possible x medium | CD6 publishes N and the scaffold hash with every number; re-run on the beat. |
Person · product · place — not yet measured for this domain
knowledge/compare/coding.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain coding, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).The field — discovered, not chosen
Rows written by nx_field_discover from coding.seeds: the industry's own lists (Wikipedia wikitext, GitHub topics, awesome lists) read mechanically, every candidate counted across seeds. The matrix columns above are a SEAT'S pick; this band is the population they were picked from, and the stats line measures one against the other. A rival here is a lead, never a verdict — it earns a column when its capabilities are read and pinned.
field|seeds=4|scanned=4|fetched=4|reused=0|failed=0|named=4|candidates=4|mentions=4|capped=0
rival|Cursor|1|1|col0-cursor|https://cursor.com/blog/secure-codebase-indexing|named
rival|SWE-agent|1|1|col1-swe|https://arxiv.org/abs/2405.15793|named
rival|Qwen2.5-Coder|1|1|col2-qwen2|https://arxiv.org/abs/2409.12186|named
rival|Aider|1|1|col3-aider|https://aider.chat/docs/repomap.html|named
| Rank | Rival | Seeds | Mentions | First seed | Kind | Link |
|---|---|---|---|---|---|---|
| 1 | Cursor | 1 | 1 | col0-cursor | named | https://cursor.com/blog/secure-codebase-indexing |
| 2 | SWE-agent | 1 | 1 | col1-swe | named | https://arxiv.org/abs/2405.15793 |
| 3 | Qwen2.5-Coder | 1 | 1 | col2-qwen2 | named | https://arxiv.org/abs/2409.12186 |
| 4 | Aider | 1 | 1 | col3-aider | named | https://aider.chat/docs/repomap.html |
field candidates 4|shown 4 of 4|matrix columns in the field 4 of 4|discovered rivals with no column 0|malformed rows 0 (counted, never rendered)|read-capped 0
Gaps from the record — what the estate does that no board carries
The record census (nx_goalmap record) reads the invoked-tool population and every plan queue row and files each organ or directive that NO matrix, plan or gates row names. A row here is a callout the boards missed: adjudicate it onto a board or declare it infrastructure. Census state BLIND (age 74396 s), sources read 4 of 7 declared — a BLIND census is a FLOOR: unread sources can only add rows.
| kind | name | board | source | evidence |
|---|
rows shown 0|this board's directives 0|estate-wide un-boarded organs 492 (listed in full on /compare/ecosystem)|census rows 611|malformed 0 (counted, never rendered)
References
- [swebench23] Jimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024 (arXiv:2310.06770). publisher · read in our library
knowledge/fetched/cmp_coding_swebench23.html· pinhcd170cab3a56f891ee94070d5748d923685f2a84a91201d6336a221436fd0f9f· accessed 2026-08-18 · published-paperGrounds: The "Repository-level / multi-file editing" row: SWE-bench issue resolution over real GitHub repositories is the frontier task that row names, and the reason forge's single-file organ scope grades 0 against 2/2/1/2; it is also the benchmark behind the SWE-bench-Lite figure in the "Best-of-N sampling against the free verifier" row. - [sweagent24] Yang, Jimenez, Wettig, Lieret, Yao, Narasimhan, Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024 (arXiv:2405.15793). publisher · read in our library
knowledge/fetched/cmp_coding_sweagent24.html· pinh1589f706b0380d637196a65970b23cb89037465fe6d883af1007e34a41d8bdbf· accessed 2026-08-18 · published-paperGrounds: The SWE-agent column (@cols): the agent-computer-interface harness those column codes describe, and the flaky-test-suite verify signal the "Deterministic build + byte-exact judge" EXCEED row contrasts against (a byte-exact GREEN is a provable equality, a test-suite pass is not). - [qwen25coder24] Hui, Yang, Cui, Yang, Liu, Zhang et al. (Qwen Team). Qwen2.5-Coder Technical Report. arXiv:2409.12186, 2024. publisher · read in our library
knowledge/fetched/cmp_coding_qwen25coder24.html· pinh449fe8f906c0ebde1f5c7f7957692a28dff03a1a935e12882a0767c271555390· accessed 2026-08-18 · published-paperGrounds: The Qwen2.5-Coder column and the "Local sovereign model tier loaded bits-up" row: Coder-0.5B and Coder-1.5B are this model family, run q8_0 through the no-float integer stack; it is also the backbone named in the "Sovereign code embeddings, external-oracle validated" row. - [llmonkeys24] Brown, Juravsky, Ehrlich, Clark, Le, Re, Mirhoseini. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv:2407.21787, 2024. publisher · read in our library
knowledge/fetched/cmp_coding_llmonkeys24.html· pinh922be22acaa0999ba5690c50a680d29228db6c0e73aa6c89ffb769acb7cab68c· accessed 2026-08-18 · published-paperGrounds: The "Best-of-N sampling against the free verifier" row: the 15.9 to 56 percent SWE-bench Lite result at 250 samples with a verifier is this paper's headline number; the row grades our loop absent although our verifier is exact and free. - [selfdebug23] Chen, Lin, Schaerli, Zhou. Teaching Large Language Models to Self-Debug. arXiv:2304.05128, 2023. publisher · read in our library
knowledge/fetched/cmp_coding_selfdebug23.html· pinh967b70ab50990ac9fd716a27160f07d6d494283bdefdfa432a05d55e9b99af7c· accessed 2026-08-18 · published-paperGrounds: The "Bounded repair loop (verbatim compiler stderr)" row: Self-Debug is one of the three named field methods (ACT+Debugger, Reflexion, Self-Debug) that feed execution feedback back to the model; ours feeds verbatim compiler stderr under a bounded loop, gate-proven 5/5. - [jinacode25] Kryvosheieva, Sturua, Guenther, Xiao. Efficient Code Embeddings from Code Generation Models (jina-code-embeddings). arXiv:2508.21290, 2025. publisher · read in our library
knowledge/fetched/cmp_coding_jinacode25.html· pinhe933c046f06596bf19f75f0acfe3e3d28c5b50a860227d76cf492f824637de49· accessed 2026-08-18 · published-paperGrounds: The "Code retrieval + reranking (CoREB)" and "Sovereign code embeddings, external-oracle validated" rows: the jina-code-embeddings recipe (Qwen2.5-Coder backbone, last-token pooling) is the external ORACLE our 100 percent integer embed path reproduced within 5 permil on matched cells, and the arXiv id the matrix cites. - [xgrammar24] Dong, Ruan, Cai, Lai, Xu, Zhao, Chen. XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models. arXiv:2411.15100, 2024. publisher · read in our library
knowledge/fetched/cmp_coding_xgrammar24.html· pinh16539c51e47d62e3672b6fa34da90194c587a5708df9c346ef0ff5723c44c56f· accessed 2026-08-18 · published-paperGrounds: The "Grammar-constrained decoding (logit mask)" row: XGrammar (with Outlines) is the named engine whose near-zero-overhead logit masking sets the bar for the sovereign-exclusive NishiLang-grammar mask this row declares absent. - [aider-repomap] Aider documentation: Repository map -- a concise tree-sitter derived map of the classes and functions in the whole git repository with their signatures, ranked to fit the context window. publisher · read in our library
knowledge/fetched/cmp_coding_aider-repomap.html· pinh77708d7004b0755018668f7db831cb09b7cb5883da4e46d90bd65fe5683b07f0· accessed 2026-08-18 · vendor-docGrounds: The Aider column and the "Verified context pack (organ-emitted, regenerable)" row: the Aider repo-map is one of the two named bars for that row; ours is vetted and auto-regenerated from sources rather than ranked from a symbol graph. - [cursor-indexing] Cursor. Securely indexing large codebases (engineering blog, 2026-01-27): Merkle-tree change detection, syntactic chunking, embeddings computed asynchronously, per-codebase vector namespace with obfuscated paths. publisher · read in our library
knowledge/fetched/cmp_coding_cursor-indexing.html· pinh150ddab0184554edda2844c71304792378f8034a3e88da5556912c1024d3640c· accessed 2026-08-18 · vendor-docGrounds: The Cursor column and the "Verified context pack (organ-emitted, regenerable)" row where the Cursor codebase-index is graded 2 (Best): this is the vendor's own description of that index; ours grades 0 as a regenerable vetted pack, not a semantic index of the whole tree.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/coding.matrix · source checks show implementation presence; runtime and user-outcome evidence are reported separately · JavaScript supports page controls