Nishi FamilyCompare › Deep Research and Answer Engines

Nishi Compare · measured, not asserted

Deep Research and Answer Engines

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

Nishi vs Perplexity and DeepResearcher and Tongyi DeepResearch and deep-searcher

Layer 1 · Executive

Where we are. The sovereign research engine is real and deterministic: fetch-and-bank over own TLS, BM25 and dense and PPMI retrieval, a pretrained sovereign Qwen reader measured on-protocol (SQuAD 510, Hotpot 505 vs mechanical 250), F1 scoring, a compounding findings ledger, corpus semantic retrieval, and adversarial refutation that is AUTO-INVOKED in-run: every numeric claim graded against its cited source and cross-validated across independent origins before publish. Two exceeds: fully sovereign at zero dollars per query, and integer-deterministic crash-resumable with no link-rot. Behind the field on: RL-trained agency, fluent synthesis, a trained cross-encoder reranker, multimodal output, and a served product API.

Where we need to go. A research engine someone else can call: a served API over the existing pipeline, fluent synthesis written by the sovereign model but GATED by the same refutation that already guards numbers, a trained reranker that beats the lexical arm on a pre-declared ruler, charts and figures in the report, and an agent whose policy is learned from its own verified runs -- never a hosted LLM in the trust path.

The unit. 1 u = one measured session-leg (estate calibration: graphics R21 in one leg 2026-08-15). Local evidence: auto-invoked refutation and multi-source cross-validation each landed gated (4/4, 5/5, 7/7, 10/10) in one leg apiece on 2026-07-14; estimates are relative to those.
Where we are: 5 open rungs. Retrieval, reading and verification present and honest; product surface and learned components absent. Counts measured at emit below this line.
Cost to a callable product: 1.5 u. DR1 the served research API over nx_research_unified -- the pipeline already runs end to end, the door is missing.
Cost to fluent, gated synthesis: 2 u. DR2 the sovereign model writes the answer; nxr_assess grades every claim in it before it ships, exactly as it grades extractive output today.
Cost to learned quality: 8 u. DR3 the trained cross-encoder reranker (accept rule pre-declared against BM25 321 permil), DR4 charts in the report, DR5 the learned agent policy over the verified-run ledger.

Research bar. Perplexity is measured on served API, polished synthesis, multimodal answers. Theirs: the product bar. Ours: DR1 DR2 DR4 measured on this page.

Research bar. DeepResearcher (arXiv 2504.03160) is measured on GRPO-trained research agent, multi-source cross-validation. Theirs: the learned-agent bar. Ours: DR5; cross-validation already matched.

Research bar. arXiv query API is measured on whether a machine can ask the scholarly frontier a question without a key or a browser. Theirs: a public Atom API with a Boolean field grammar. Ours: DN0 already composes it as a data-declared source row [@arxiv-api].

Research bar. OpenAlex works API is measured on a second key-free index, which is what makes a wide sweep a pace problem rather than a quota problem. Theirs: free REST over the open catalogue, no key. Ours: DN1 must admit it before the sweep can run fleet-wide [@openalex-api].

14 of 23 capabilities measured|4 of them measured exceeds|9 open|coverage 608/1000|adoption 7 full / 7 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787904987 domain=deepresearch target_version=0.1 rungs=12 done=4 open=8 finish=1 ranker=nx_dr_ocm

#StageRungPriorityDerivation
FFINISHUnified refutation in-run (DB0) nxr_assessLIB-GATE-ONLY importers=3imported only by validation organs (gates, tests, benches): wire it into a shipping program
#10.1Served research API (DR1) rapi_serve2000v=15 m=2 c=15
#2laterGated synthesis (DR2) ls_synthesize1050v=7 m=3 c=20
#3laterTrained cross-encoder reranker (DR3) nr_cross_score533v=8 m=2 c=30
#4laterLearned agent policy (DR5) ra_policy_step400v=7 m=2 c=35
#5laterCharts in the report (DR4) mr_chart_emit133v=2 m=1 c=15
#6laterFleet sweep on a beat (DN1) fs_sweep_all133v=1 m=2 c=15
#7laterProposal adjudication in the hive (DN2) fs_adjudicate0v=0 m=2 c=10
#8laterMirror bound to its own URL (DN3) crg_mirror_url_bound0v=0 m=1 c=10

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Unified refutation in-run (DB0)nxr_assessEvery numeric claim graded vs its source and cross-validated before publish -- LANDED gatedOrgan0 u
Sovereign reader (DB1)qwx_extractOn-protocol SQuAD 510, Hotpot 505 -- LANDEDOrgan0 u
Lexical retrieval (DB2)idf_q10BM25 baseline, integer-deterministic -- LANDEDOrgan0 u
Served research API (DR1)
after DB0
rapi_serveAn authenticated HTTPS door that runs the unified pipeline for a question and returns the graded report with per-claim verdicts; gate proves a request over the edge returns a report whose claims carry their sources and a fabricated-number fixture is refused by the same pathOrgan1.5 u
Gated synthesis (DR2)
after DB0,DB1
ls_synthesizeThe sovereign model composes prose from the extractive findings; every sentence with a number or a citation is graded by nxr_assess and a discredited claim forces rewrite or STOP-INCOMPLETE, never ships; gate proves a planted false number in the synthesis is caught and the report abstainsOrgan2 u
Trained cross-encoder reranker (DR3)
after DB2
nr_cross_scoreA no-float cross-encoder scoring (query, passage) pairs, trained on the ledger's verified pairs; ACCEPT RULE declared before the run: must EXCEED the BM25 arm (321 permil on the BEIR ruler) on the same queries or it stays UNWIRED, as the semppmi and RRF arms didOrgan3 u
Charts in the report (DR4)
after DR2
mr_chart_emitNumeric findings rendered as sovereign SVG charts inside the report, each chart citing the claims it plots; gate proves a chart's plotted values equal the graded claims byte-for-byte and a chart of an ungraded number is refusedOrgan1.5 u
Learned agent policy (DR5)
after DB0,DR3
ra_policy_stepA policy that chooses the next action (fetch, read, refute, stop) trained on the verified-run ledger through the no-float autograd; accept rule: higher on-protocol F1 per fetch than the hand-written loop on a held-out question set, or UNWIREDOrgan3.5 u

Milestones

MilestoneRungsCumulative
M1 · CallableDR11.5 u
M2 · Fluent and gatedDR2,DR45 u
M3 · LearnedDR3,DR511.5 u
RungCloses withDefinition of done (pre-declared)ExecutorEst.
The noticer itself (DN0)fs_deficitLANDED 2026-08-20. Seeds derived from the domain's own board, outside pulled through the sovereign fetcher, findings diffed against our rows and refs, GAP PROPOSALS filed to a worklist and the frontierprop- plane. Done-rule: gate GREEN with a negative control proving a no-gap board issues ZERO fetches and ZERO proposals, and a tooth proving a run never writes the board it readsOrgan0 u
Fleet sweep on a beat (DN1)
after DN0
fs_sweep_allEvery /compare domain scanned on a cadence instead of one. Done-rule: a per-source pace budget that the run PRINTS and obeys, a second key-free index admitted as a source row, and a full-fleet run whose outbound request count matches rows times sources times domains exactly -- a sweep that cannot state its own request count is a load incident waiting to happenOrgan1.5 u
Proposal adjudication in the hive (DN2)
after DN0
fs_adjudicateA seat accepts or rejects a proposal by ONE call that closes the plane row and, on accept, emits the ready-to-paste ref row with the mirror already fetched and pinned. Done-rule: an accepted proposal produces a refs row that nx_compare_refs_gate passes on first run, and a rejected one leaves the board byte-identicalOrgan1 u
Mirror bound to its own URL (DN3)crg_mirror_url_boundThe referee proves a mirror is a capture of the url in ITS OWN ROW, not merely that some file exists and its pin matches. Done-rule: the four 2026-08-18 charsim rows that pass today must FAIL this tooth, and every honestly-fetched row must still pass -- the bite is already sitting in the corpusOrgan1 u
M4 · The estate notices its own gapsDN0,DN1,DN2,DN33.5 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiPerplexityDeepResearcherTongyideep-searcher
Web fetch and bank pipelineMeasured: rf_fetch_bank exists in runtime/nx_research_engine.nx, verified at emit. RAG fetch loop; all fetch the live web Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
adoption BUILT-UNPROMOTED
Lexical retrieval (BM25)Measured: idf_q10 exists in runtime/nx_intlog.nx, verified at emit. Sparse retrieval baseline Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).
Dense vector retrievalMeasured: nx_vec_index exists in runtime/nx_vec_index.nx, verified at emit. deep-searcher leads (Milvus / Zilliz vector DB) [deepsearcher] Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Distributional embeddingsMeasured: nx_distrib_embed exists in runtime/nx_distrib_embed.nx, verified at emit. Co-occurrence semantic vectors Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
PPMI semantic matrixMeasured: nx_semppmi_build exists in runtime/_hdl_build/nx_semppmi_build.nx, verified at emit. Ours explicit sparse; theirs neural Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-DARK
Trained embeddingsMeasured: nx_embed_train exists in runtime/_hdl_build/nx_embed_train.nx, verified at emit. Ours PPMI-factorization toy; theirs neural leads Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-UNAUTHORISED
Multi-hop QA readerMeasured: qwx_extract exists in runtime/nx_qwen_extract.nx, verified at emit. Pretrained sovereign Qwen 1-shot + s0-window opt-in: on-protocol SQuAD 510 and Hotpot 505 (above mech-oracle 368) [yang2018hotpot] vs mech 250; MuSiQue 122 ties the composition wall [trivedi2022musique] (NAS-measured, multi-alias). A 2-shot surface-form variant won +104 on the single-gold standalone bench but did NOT transfer to the multi-alias on-protocol metric (SQuAD 427 vs 510) and was reverted (confirmed by nx_reader_ruler_adversary_gate 9/9). Evidence-grounded: SQuAD human F1 86.8 per arxiv:1606.05250 [rajpurkar2016]; published SQuAD v1.1 SOTA = F1 932 (BERT, arxiv:1810.04805) [devlin2018] VERIFIED via nx_research_verify -- our reader 510 = 55 percent of SOTA, which still leads Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder).
QA F1 / EM scoringMeasured: qs_f1 exists in runtime/nx_qa_score_lib.nx, verified at emit. SQuAD-style measured evaluation Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).
Findings ledger (compounding memory)Open — no implementing organ is measured for this axis yet. Persistent provenance ledger
Corpus semantic retrievalMeasured: nx_lib_semantic exists in runtime/nx_lib_semantic.nx, verified at emit. deep-searcher leads Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
RL-trained research agentOpen — watching runtime/nx_research_agent.nx : ra_policy_step, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. DeepResearcher / Tongyi GRPO-trained lead; Nishi none [shao2024grpo]
watching ra_policy_step
LLM synthesis / reasoningOpen — watching runtime/nx_llm_synth.nx : ls_synthesize, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (polished synthesis); Nishi extractive only
watching ls_synthesize
Adversarial refutationMeasured: nxr_assess exists in runtime/nx_research_unified.nx, verified at emit. DeepResearcher cross-validates across sources (leads). Nishi now has nx_claim_verify (single) + nx_research_verify (batch STAGE) -- autonomous: fetch a source over sovereign TLS then confirm/discredit/flag-unverifiable a claim vs primary evidence, never fabricating (LIVE-PROVEN: claim 3/3, batch 4/4, and grounded the SQuAD SOTA 932 above) + metric-ruler adversary (nx_reader_ruler_adversary_gate 9/9). Single-source numeric-grounding (catches fabrication/alteration) vs multi-source cross-validation. NOW AUTO-INVOKED in-run (2026-07-14): nx_research_unified (nxr_assess) grades every numeric claim vs its cited source and GATES PUBLISH -- a discredited/fabricated number forces CONTINUE or STOP-INCOMPLETE, never ships; gate-proven nx_research_verify_synth_gate 4/4 + nx_research_unified_gate 5/5 + nx_research_refute_gate 7/7. So PRESENT (auto-invoked, gated). AND multi-source cross-validation is NOW ALSO auto-invoked (2026-07-14b): cvx_answer_verify (nx_research_crossval) inside nxr_assess corroborates each claim across ALL sources with the independence discipline -- >=2 distinct origins required, same-origin echo never counts, conflict surfaced not averaged; crossval gate 10/10 incl a LIVE 3-host proof (GPT-3 175B CORROBORATED across arxiv.org + api.datacite.org while a rate-limited semanticscholar stub was absorbed as unverifiable, not faked). DeepResearcher's agentic multi-source loop still leads on breadth, so NOT exceed [deepresearcher2025] Adoption: LIB-GATE-ONLY importers=3 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program.
adoption LIB-GATE-ONLY importers=3
Neural reranker (cross-encoder)Open — watching runtime/nx_neural_reranker.nx : nr_cross_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Trained rerank; Nishi lexical + sparse only
watching nr_cross_score
Multimodal report outputOpen — watching runtime/nx_multimodal_report.nx : mr_chart_emit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (charts / images); Nishi text only
watching mr_chart_emit
Served research API / productOpen — watching runtime/nx_research_api.nx : rapi_serve, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (product API); Nishi organ-only [perplexity-api]
watching rapi_serve
Sovereign -- own substrate, TLS, index; zero OpenAI / cloud dep; 0 dollar per queryMeasured exceed: ss_fold_cp in runtime/nx_seg_store.nx, verified at emit. All use hosted LLMs; Tongyi open-weights partial [tongyi2025]; Nishi fully sovereign Adoption: LIB-WIRED importers=441 nonval=321 — fully adopted (top of its ladder).
Integer-deterministic, crash-resumable, no link-rotMeasured exceed: idf_q10 in runtime/nx_intlog.nx, verified at emit. Reproducible; peers are non-deterministic LLM pipelines Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).
Capability-gap proposals diffed against our own boardMeasured exceed: fs_deficit in runtime/nx_frontier_scan.nx, verified at emit. LANDED 2026-08-20, nx_frontier_scan_gate GREEN 14 of 14 with two teeth BITE-PROVEN. The organ derives its query seeds from a /compare domain's OWN data (the .axes frontier-keyword field, the .matrix row labels), pulls the outside through the sovereign fetcher over sources that are themselves data [arxiv-api], and emits GAP PROPOSALS naming the row each candidate challenges plus its fetched mirror. THE FOUR RIVAL COLUMNS ARE CODED 0 AS A CATEGORY STATEMENT AND NOT AS A FEATURE AUDIT: all four are answer engines that respond to a question a human asked, none maintains a capability board, so there is no comparable surface to measure them on and none is fabricated here. It emits proposals and NEVER admits one -- the measured law is that no margin threshold makes auto-declaration safe, and a completion signal that keys on a name rewards writing the name. Adoption: RUN-BY:cron — fully adopted (top of its ladder).
Frontier noticing across every /compare domain on a beatOpen — watching runtime/nx_frontier_scan.nx : fs_sweep_all, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN. One domain is on a daily beat today (cron.reg, 07:35) at 12 outbound requests per run, because outbound cost is rows times sources. A fleet sweep over 58 domains at the same settings is roughly 700 requests a day, which is a rate-limit and host-load decision before it is a scheduling one, so the rung must admit a second key-free scholarly index [openalex-api] and a per-source pace budget before it can run wide. Rival columns unmeasured for the category reason given on the row above.
watching fs_sweep_all
Citation mirror provably bound to the URL it claimsMeasured exceed: crg_prov_state in runtime/nx_compare_refs_gate.nx, verified at emit. LANDED 2026-08-20, nx_compare_refs_gate GREEN 18 of 18. THE CONTRACT WAS DECLARED AS crg_mirror_url_bound AND THE SHIPPED SYMBOL IS crg_prov_state -- the rename is recorded here rather than left to read as a moved goalpost: the contract named a property, the function names the three-state answer. The binding is made at FETCH time, the only moment both facts are in hand -- nx_research_fetch appends a row carrying epoch, requested-url, mirror, sha256, bytes and redirect-hops to knowledge/status/fetch_provenance.jrnl, and this gate joins every declared mirror against it. THREE STATES because two would lie: PROVEN, MISMATCH (this mirror under a DIFFERENT url), and UNPROVEN (no journal row at all, which is every mirror captured before the journal existed and is NOT a defect). Live 2026-08-20: jrnl_rows 29, proven 4, mismatch 0, unproven 617, declared_mirrors 621, and the partition SUMS. MISMATCH is ARMED at zero; UNPROVEN is RATCHETED and self-baselined at 617, because arming it fleet-wide would have turned 68 domains RED in one step and taught everyone to ignore the gate. Every unproven mirror is NAMED in knowledge/status/refs_provenance_unproven.txt, because a count without a worklist is not actionable. Bite-proven in BOTH directions on fixtures assembled at runtime, with a tooth asserting the fixtures actually parsed before any verdict is read off them. RESIDUAL, STATED IN THE GATE VERDICT ITSELF AND CARRIED AS THE ROW BELOW: provenance proves the mirror is OF the url, never that it SUPPORTS the claim. THE HOLE IT CLOSED, kept here so the fix keeps its reason: it is the exact hole the 2026-08-18 metahuman episode fell through. Four references were appended whose URLs were never opened; three carry mirror and pin both absent, and the fourth points at a page fetched three days earlier from a DIFFERENT url -- and the referee passed the file 9 of 9, because it proves the mirror EXISTS and the pin EQUALS the filehash of that mirror, never that the mirror is a capture of the row's own url. A PIN PROVES THE BYTES DID NOT CHANGE, NEVER THAT THEY ARE THE RIGHT DOCUMENT. The same blindness has a second cause now closed: until 2026-08-20 the TLS-1.2 fetch leg truncated bodies near 256 KiB while printing a clean SAVED, so a partial page could be pinned as whole. Rival columns unmeasured for the category reason given two rows above. Adoption: GATE:LIVE trial=- — fully adopted (top of its ladder).
Citation mirror provably SUPPORTS the claim it backsOpen — watching runtime/nx_compare_refs_gate.nx : crg_claim_supported, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN, and it is the half that fetch-time provenance CANNOT close -- named as its own contract rather than left implied, because the row above could easily be read as having ended the class. Provenance proves the mirror is a capture of the url in its row. It says NOTHING about whether that document contains the claim the row cites it for, and on 2026-08-20 that second question had to be answered BY HAND over four mirrors and 22 claims: PRESENT 6, PARTIAL 2, ABSENT 13. Every one of the 13 sat behind a real file with a valid pin, and after this rung ships they would still sit behind a real file with a valid pin AND a correct provenance row. The done-rule is a checker that takes a row grounds field and its mirror and returns SUPPORTED, UNSUPPORTED or UNDECIDABLE with the witnessing byte offsets, ratcheted the same way and armed on nothing until it has a calibrated false-positive rate on the existing corpus. Rival columns unmeasured for the category reason given three rows above.
watching crg_claim_supported
Frontier sources that are not scholarly indexes -- vendor release notes and conference programmesOpen — watching runtime/nx_frontier_scan.nx : fs_src_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN, and the reason it is open is a COVERAGE BOUNDARY, not a scoring failure. THE LOOP CANNOT SEE THESE, IT DID NOT MERELY MISS THEM. The two wired sources are arXiv and HN Algolia. Measured against the four references a human added to charsim on 2026-08-18: metahuman57 is an Epic Developer Community release note and rtxhair26 is a GDC schedule session page -- NEITHER INDEX CARRIES EITHER KIND OF PAGE AT ALL; cc5-reallusion is a trade-press article, reachable through HN only on the condition that some human happened to submit it, which is not a capability of the source; sig26-avatars is a SIGGRAPH changelog index page, itself unreachable, though the papers it lists are on arXiv under different citations. SO THE HONEST DENOMINATOR IS ZERO: of the four cited urls, ZERO were in-principle reachable from the sources wired, and the replays 0-of-4 therefore means the loop missed 0 of the 0 it could see. THE MECHANISM ITSELF WORKS WHERE ITS SOURCES CAN SEE -- the 2026-08-20 v2 replay surfaced Show HN SwitchLight, a browser relighting product, against the SKIN full-PBR-and-normal-map-lighting row, unaided and on-topic. PRE-DECLARED ACCEPT RULE, WRITTEN BEFORE THE SOURCE EXISTS SO IT CANNOT BE RATIONALIZED AFTERWARDS: wiring a vendor and release-note source kind is accepted only if replaying the 2026-08-18 PRE-STATE board surfaces the MetaHuman-class surface-layer gap unaided, as at least one proposal whose url is a vendor release note challenging a SKIN, HAIR, MESH or ANIM row, and only if that proposal mirror carries a row in knowledge/status/fetch_provenance.jrnl binding it to that url. THE SECOND CONDITION IS NOT DECORATION: the same audit proved Epic SILENTLY REPOINTED the metahuman-5-7 slug to the 5.8 notes, so a vendor page can change identity underneath a citation, and a vendor source without fetch-time provenance would rebuild the exact hole this lane just spent a day retracting. Rival columns unmeasured for the category reason given four rows above.
watching fs_src_page

Risk register

RiskLikelihood x impactMitigation
Synthesis that reads well hides a wrong number better than extractive text didlikely x highDR2 is graded sentence by sentence by the existing referee; fluency never bypasses the gate.
A learned reranker or policy that wins on the training ruler and loses held-out is an illusion of progresspossible x highDR3 and DR5 carry pre-declared accept rules on held-out sets; a miss stays UNWIRED with the number published.
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The coverage above is capability presence measured against source — not depth, scale, or polish, where mature rivals may lead. Exceeds are claimed only where a mechanism backs them. Every open gap is a watch contract: it names the organ and symbol that closes it, and this page flips the cell itself when that workstream ships.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/deepresearch.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain deepresearch, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [perplexity-api] Perplexity. API Platform overview -- Router, Agent, Search and Embeddings APIs (real-time web-grounded answers). Vendor documentation. publisher · read in our library knowledge/fetched/cmp_deepresearch_perplexity.html · pin hae63a992df31746dde53a7ba823cfe0479d11e091814e8fcaaf8dd86c7bc0e46 · accessed 2026-08-18 · vendor-docGrounds: The Perplexity column: Best on "Served research API / product", "LLM synthesis / reasoning" and "Multimodal report output" -- the documented product API is the served-answer-engine capability those codes describe, and it is a hosted LLM service, which is the foil for the "Sovereign -- own substrate, TLS, index; zero OpenAI / cloud dep" exceed row.
  2. [deepresearcher2025] Zheng, Y., Fu, D., Hu, X., Cai, X., Ye, L., Lu, P., Liu, P. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments. arXiv:2504.03160, April 2025. publisher · read in our library knowledge/fetched/cmp_deepresearch_deepresearcher2025.html · pin hfc6c302637b8af3e2c88af6cc438c81002200150b0c2641ad1fb7b62fc4f0549 · accessed 2026-08-18 · published-paperGrounds: The DeepResearcher column: Best on "RL-trained research agent" and "Adversarial refutation" -- the abstract describes end-to-end RL training of deep-research agents in real web environments and reports emergent cross-validation of information across multiple sources, which is exactly the multi-source breadth the row note concedes still leads ours.
  3. [tongyi2025] Tongyi DeepResearch Team, Li, B., Zhang, B. et al. Tongyi DeepResearch Technical Report. arXiv:2510.24701, October 2025 (30.5 billion total parameters, 3.3 billion activated per token; model, framework and solutions open-sourced). publisher · read in our library knowledge/fetched/cmp_deepresearch_tongyi2025.html · pin h24ac13b5b8fafe85408e0b6c254151046e31913659f1c4fc1a2afa52bf3b35e8 · accessed 2026-08-18 · published-paperGrounds: The Tongyi column: Best on "RL-trained research agent" (agentic mid-training plus post-training) and Part on the "Sovereign" row ("Tongyi open-weights partial") -- the report states the model is open-sourced, which is why its column is the only rival with a non-zero code on that exceed row.
  4. [deepsearcher] Zilliz. deep-searcher -- open source deep research alternative to reason and search on private data, written in Python (GitHub repository, zilliztech). publisher · read in our library knowledge/fetched/cmp_deepresearch_deepsearcher.html · pin hf906e0ee916636ecf7db998d4a955ff020af6a15cb02bb76220a32c1f768d46b · accessed 2026-08-18 · source-readGrounds: The deep-searcher column: Best on "Dense vector retrieval" and "Corpus semantic retrieval" (the note: Milvus / Zilliz vector DB) -- the repository is the vector-database-backed private-data research agent those codes describe.
  5. [rajpurkar2016] Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P. SQuAD: 100,000+ Questions for Machine Comprehension of Text. EMNLP 2016; arXiv:1606.05250. Human performance F1 86.8 percent (abstract). publisher · read in our library knowledge/fetched/cmp_deepresearch_rajpurkar2016.html · pin hec5af640421c7c422bd13b49c0d2deac62a706d71047fad2304930cc756f48ce · accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" and "QA F1 / EM scoring" rows: the SQuAD protocol (span extraction, F1 and EM) is the estate's reader ruler, and the row note's "SQuAD human F1 86.8 per arxiv:1606.05250" is quoted from this abstract.
  6. [devlin2018] Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2019; arXiv:1810.04805. SQuAD v1.1 Test F1 93.2 (abstract). publisher · read in our library knowledge/fetched/cmp_deepresearch_devlin2018.html · pin hffa5d3591fa543ce7b833bea05f1455d40a6be8680417720f09da8807ada6abe · accessed 2026-08-18 · published-paperGrounds: The "Multi-hop QA reader" row note: "published SQuAD v1.1 SOTA = F1 932 (BERT, arxiv:1810.04805)" -- the 93.2 F1 is quoted from this abstract, and our reader's 510 is graded as 55 percent of it, honestly BEHIND.
  7. [yang2018hotpot] Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Manning, C.D. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. EMNLP 2018; arXiv:1809.09600. publisher · read in our library knowledge/fetched/cmp_deepresearch_yang2018hotpot.html · pin h1d2aba7b7aa74242566af15165d3861311bc478db918ebf7a28ebc94d8b20ffd · accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" row: the Hotpot 505 figure in the note is measured on this dataset's multi-hop protocol; it is the second of the three benchmarks the reader ruler runs.
  8. [trivedi2022musique] Trivedi, H., Balasubramanian, N., Khot, T., Sabharwal, A. MuSiQue: Multihop Questions via Single-hop Question Composition. TACL 2022; arXiv:2108.00573. publisher · read in our library knowledge/fetched/cmp_deepresearch_trivedi2022musique.html · pin hd8ca40a3303d9287245cad135d31a338a646f67e97764dd3ce4c59c321a578c7 · accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" row: "MuSiQue 122 ties the composition wall" -- MuSiQue is built by composing single-hop questions so that shortcut readers fail, which is the wall the note names; the third benchmark of the reader ruler.
  9. [shao2024grpo] Shao, Z., Wang, P., Zhu, Q. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024 -- introduces Group Relative Policy Optimization (GRPO), a PPO variant (abstract). publisher · read in our library knowledge/fetched/cmp_deepresearch_shao2024grpo.html · pin h32e37ca6d2af1522b1372c824d8ace3f400f44b3701a3cc69fd624254a39aeb5 · accessed 2026-08-18 · published-paperGrounds: The "RL-trained research agent" row note ("DeepResearcher / Tongyi GRPO-trained lead; Nishi none"): GRPO is defined in this paper, so the training method the row credits the leaders with has a single published definition, and the _ABSENT_ nx_research_agent watch names what a sovereign no-float RL loop must implement.
  10. [arxiv-api] arXiv. arXiv API User's Manual (info.arxiv.org): the search_query grammar with field prefixes such as all: ti: au:, the Boolean operators AND OR ANDNOT, sortBy submittedDate, and the Atom feed schema returned by the query endpoint. publisher · read in our library knowledge/fetched/cmp_deepresearch_arxivapi.html · pin hf6c3ce2936f629e03fd42faf911604a7b0677ae3475c9c8608a543e9a7579671 · accessed 2026-08-20 · vendor-docGrounds: The "Capability-gap proposals diffed against our own board" row and knowledge/frontier_sources.conf: this manual is why every source row carries an explicit term-join operator. A space-joined search_query does NOT AND its terms, so an unjoined seed returns the newest submissions of the whole archive with HTTP 200 -- measured on the first metahuman replay, which proposed a Bose-Einstein condensate paper against a skin-shader row.
  11. [openalex-api] OpenAlex. API Overview (docs.openalex.org): a free and open catalogue of scholarly works behind a public REST API that requires no key, with works filterable and sortable by publication date. publisher · read in our library knowledge/fetched/cmp_deepresearch_openalex.html · pin ha4405ec3a16a534736c39084c43a6c945b3850c27c080a0c4eed4798cd7ae0bf · accessed 2026-08-20 · vendor-docGrounds: The "Frontier noticing across every /compare domain on a beat" watch row: a fleet sweep is a rate-limit problem before it is a scheduling problem, and a second key-free scholarly index is what makes the wide sweep affordable without hammering a single source. It names the source row that rung must admit.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/deepresearch.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers