Nishi Research — the beyond-SOTA frontier & fetch spec

What to build to EXCEED the state of the art, and exactly what to look for when we fetch. Accountable: researcher. Every row = a filed frontier rung, momentum-ranked from real corpora (OpenAlex 4MB · Crossref 3.8MB) via nx_swcompare_gapmap. Map: /org · graph: /atlas · benches+radars: /compare.

The frontier — per lane: the top unproven opportunity + what to fetch

LaneBeyond-SOTA opportunityWhat to fetch (look for)Rung
videoend-to-end NEURAL codec · implicit neural representation · learned entropy/RDO · perceptual/generativeDCVC-FM, hyperprior/learned-entropy, VVC/AV1 tool-set, GAN-perceptual metrics — look for RD curves vs x264/x265 at equal SSIMF611
recall/searchneural RERANK · RRF hybrid BM25+dense · web-scale indexBEIR suite (13 tasks), cross-encoder rerankers, SPLADE/ColBERT — look for nDCG@10 + latency + which signals fuseF231/236
atlas/recombinelink-prediction on the dep graph · workflow mining · evolutionary recombinationnode2vec/GraphSAGE, van der Aalst process mining, AlphaEvolve/OpenEvolve, CodeScene change-coupling, Structure101/Lattix DSM, ArchUnit fitness-functionsF225a
living-docsGraphRAG over the doc-graph · executable/literate docs · staleness-detectOpenAlex momentum: literate-docs (12) · GraphRAG (6) · doc-staleness (3) · LLM-authoring (3) — look for retrieval grounding + freshness treatmentF250
pm/ROIautonomous-agent productivity metrics (a new category) · real-time EVMJellyfish, LinearB, Swarmia, DX getdx, Cortex — the DORA report, the SPACE paper (Forsgren), DX Core 4; look for source-of-metric + uncertainty/provenance treatment (none tag it = our exceed)F740-743
chain-of-evidencePQ signatures · transparency log · ZK proofs · witness quorumSigstore/Rekor, in-toto/SLSA, C2PA 2.x, AWS QLDB — interop-EXPORT mappings only (3rd-party substrate refused by sovereignty doctrine); look for the trust-root + revocation modelF707
model/LLMtrain-our-own AT SCALE (today: honest TOY) · no-float trainingMegatron/DeepSpeed parallelism, K-quant/GGUF, the scaling-law papers — look for tokens/param + the integer-determinism boundaryF235
gpufirst sovereign submit on the 5080 · C0 GEMM · tensor-core pathsCUDA/ROCm kernel patterns, CUTLASS tiling, the 5080 ISA — look for occupancy + the bit-exact-vs-fast tradeoffF101
civiclegislative tracking · grounded aggregation · corruption signalCrossref 3.8MB corpus, court-docket + SEC-EDGAR + USPTO feeds — look for citations verifiable vs fabricated (the zombie-loop firewall, atlas-hygiene F263)F318

The fetch doctrine — what EVERY fetch looks for

  1. Grounding over fluency — a claim is only usable if it traces to a real source (DOI/patent/filing); fabricated citations are blocked at ingest (atlas-hygiene zombie-loop firewall F263).
  2. Freshness + provenance-tag — record when the source was published and mark it PRIMARY vs AI-SYNTHESIZED; never let a synthesized summary become primary validation.
  3. Exceed vs parity — a fetch answers "where is SOTA, and which axis can we EXCEED sovereignly?" not "how do we copy it"; sub-SOTA is a filed rung, never a design excuse.
  4. Measured, not asserted — the comparison must run on real corpora (gapmap momentum) or a real bench (BEIR/DORA), liar-killed, with the envelope declared.

The infrastructure

Corpora fetched sovereignly over our own TLS (nx_https_get): OpenAlex (4MB, 2024-26 works) + Crossref (3.8MB). Per-domain .q = the fetch-query spec; .axes = the coverage axes; nx_swcompare_gapmap ranks opportunities by 2025/26 momentum (liar-killed, envelope-declared, silent-truncation banned). Live radars: video · livingdocs · civic · search · and every /compare/<domain>/frontier.

Accountable = researcher; Responsible = the per-lane researcher + the lane engineer who ships the proof. Consulted = the domain census (referee/librarian); Informed = pm (triages accepted rungs into the board via /plan). Every opportunity here is UNPROVEN by design — it becomes real only when a gate or a run proves it, then it enters the atlas and the next fetch looks further out.