Nishi Family › Compare › Deep Research and Answer Engines
Nishi Compare · measured, not asserted
Deep Research and Answer Engines
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs Perplexity and DeepResearcher and Tongyi DeepResearch and deep-searcher
Where we are. The sovereign research engine is real and deterministic: fetch-and-bank over own TLS, BM25 and dense and PPMI retrieval, a pretrained sovereign Qwen reader measured on-protocol (SQuAD 510, Hotpot 505 vs mechanical 250), F1 scoring, a compounding findings ledger, corpus semantic retrieval, and adversarial refutation that is AUTO-INVOKED in-run: every numeric claim graded against its cited source and cross-validated across independent origins before publish. Two exceeds: fully sovereign at zero dollars per query, and integer-deterministic crash-resumable with no link-rot. Behind the field on: RL-trained agency, fluent synthesis, a trained cross-encoder reranker, multimodal output, and a served product API.
Where we need to go. A research engine someone else can call: a served API over the existing pipeline, fluent synthesis written by the sovereign model but GATED by the same refutation that already guards numbers, a trained reranker that beats the lexical arm on a pre-declared ruler, charts and figures in the report, and an agent whose policy is learned from its own verified runs -- never a hosted LLM in the trust path.
Research bar. Perplexity is measured on served API, polished synthesis, multimodal answers. Theirs: the product bar. Ours: DR1 DR2 DR4 measured on this page.
Research bar. DeepResearcher (arXiv 2504.03160) is measured on GRPO-trained research agent, multi-source cross-validation. Theirs: the learned-agent bar. Ours: DR5; cross-validation already matched.
Research bar. arXiv query API is measured on whether a machine can ask the scholarly frontier a question without a key or a browser. Theirs: a public Atom API with a Boolean field grammar. Ours: DN0 already composes it as a data-declared source row [@arxiv-api].
Research bar. OpenAlex works API is measured on a second key-free index, which is what makes a wide sweep a pace problem rather than a quota problem. Theirs: free REST over the open catalogue, no key. Ours: DN1 must admit it before the sweep can run fleet-wide [@openalex-api].
14 of 23 capabilities measured|4 of them measured exceeds|9 open|coverage 608/1000|adoption 7 full / 7 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883285 domain=deepresearch target_version=0.1 rungs=12 done=4 open=8 finish=1 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| F | FINISH | Unified refutation in-run (DB0) nxr_assess | LIB-GATE-ONLY importers=3 | imported only by validation organs (gates, tests, benches): wire it into a shipping program |
| #1 | 0.1 | Served research API (DR1) rapi_serve | 2000 | v=15 m=2 c=15 |
| #2 | later | Gated synthesis (DR2) ls_synthesize | 1050 | v=7 m=3 c=20 |
| #3 | later | Trained cross-encoder reranker (DR3) nr_cross_score | 533 | v=8 m=2 c=30 |
| #4 | later | Learned agent policy (DR5) ra_policy_step | 400 | v=7 m=2 c=35 |
| #5 | later | Charts in the report (DR4) mr_chart_emit | 133 | v=2 m=1 c=15 |
| #6 | later | Fleet sweep on a beat (DN1) fs_sweep_all | 133 | v=1 m=2 c=15 |
| #7 | later | Proposal adjudication in the hive (DN2) fs_adjudicate | 0 | v=0 m=2 c=10 |
| #8 | later | Mirror bound to its own URL (DN3) crg_mirror_url_bound | 0 | v=0 m=1 c=10 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Unified refutation in-run (DB0) | nxr_assess | Every numeric claim graded vs its source and cross-validated before publish -- LANDED gated | Organ | 0 u |
| Sovereign reader (DB1) | qwx_extract | On-protocol SQuAD 510, Hotpot 505 -- LANDED | Organ | 0 u |
| Lexical retrieval (DB2) | idf_q10 | BM25 baseline, integer-deterministic -- LANDED | Organ | 0 u |
| Served research API (DR1) after DB0 | rapi_serve | An authenticated HTTPS door that runs the unified pipeline for a question and returns the graded report with per-claim verdicts; gate proves a request over the edge returns a report whose claims carry their sources and a fabricated-number fixture is refused by the same path | Organ | 1.5 u |
| Gated synthesis (DR2) after DB0,DB1 | ls_synthesize | The sovereign model composes prose from the extractive findings; every sentence with a number or a citation is graded by nxr_assess and a discredited claim forces rewrite or STOP-INCOMPLETE, never ships; gate proves a planted false number in the synthesis is caught and the report abstains | Organ | 2 u |
| Trained cross-encoder reranker (DR3) after DB2 | nr_cross_score | A no-float cross-encoder scoring (query, passage) pairs, trained on the ledger's verified pairs; ACCEPT RULE declared before the run: must EXCEED the BM25 arm (321 permil on the BEIR ruler) on the same queries or it stays UNWIRED, as the semppmi and RRF arms did | Organ | 3 u |
| Charts in the report (DR4) after DR2 | mr_chart_emit | Numeric findings rendered as sovereign SVG charts inside the report, each chart citing the claims it plots; gate proves a chart's plotted values equal the graded claims byte-for-byte and a chart of an ungraded number is refused | Organ | 1.5 u |
| Learned agent policy (DR5) after DB0,DR3 | ra_policy_step | A policy that chooses the next action (fetch, read, refute, stop) trained on the verified-run ledger through the no-float autograd; accept rule: higher on-protocol F1 per fetch than the hand-written loop on a held-out question set, or UNWIRED | Organ | 3.5 u |
Milestones
| Milestone | Rungs | Cumulative | ||
|---|---|---|---|---|
| M1 · Callable | DR1 | 1.5 u | ||
| M2 · Fluent and gated | DR2,DR4 | 5 u | ||
| M3 · Learned | DR3,DR5 | 11.5 u |
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| The noticer itself (DN0) | fs_deficit | LANDED 2026-08-20. Seeds derived from the domain's own board, outside pulled through the sovereign fetcher, findings diffed against our rows and refs, GAP PROPOSALS filed to a worklist and the frontierprop- plane. Done-rule: gate GREEN with a negative control proving a no-gap board issues ZERO fetches and ZERO proposals, and a tooth proving a run never writes the board it reads | Organ | 0 u |
| Fleet sweep on a beat (DN1) after DN0 | fs_sweep_all | Every /compare domain scanned on a cadence instead of one. Done-rule: a per-source pace budget that the run PRINTS and obeys, a second key-free index admitted as a source row, and a full-fleet run whose outbound request count matches rows times sources times domains exactly -- a sweep that cannot state its own request count is a load incident waiting to happen | Organ | 1.5 u |
| Proposal adjudication in the hive (DN2) after DN0 | fs_adjudicate | A seat accepts or rejects a proposal by ONE call that closes the plane row and, on accept, emits the ready-to-paste ref row with the mirror already fetched and pinned. Done-rule: an accepted proposal produces a refs row that nx_compare_refs_gate passes on first run, and a rejected one leaves the board byte-identical | Organ | 1 u |
| Mirror bound to its own URL (DN3) | crg_mirror_url_bound | The referee proves a mirror is a capture of the url in ITS OWN ROW, not merely that some file exists and its pin matches. Done-rule: the four 2026-08-18 charsim rows that pass today must FAIL this tooth, and every honestly-fetched row must still pass -- the bite is already sitting in the corpus | Organ | 1 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | Perplexity | DeepResearcher | Tongyi | deep-searcher |
|---|---|---|---|---|---|
Web fetch and bank pipelineMeasured:rf_fetch_bank exists in runtime/nx_research_engine.nx, verified at emit. RAG fetch loop; all fetch the live web Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ● | ● | ● | ● |
Lexical retrieval (BM25)Measured:idf_q10 exists in runtime/nx_intlog.nx, verified at emit. Sparse retrieval baseline Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
Dense vector retrievalMeasured:nx_vec_index exists in runtime/nx_vec_index.nx, verified at emit. deep-searcher leads (Milvus / Zilliz vector DB) [deepsearcher] Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ● | ◉ |
Distributional embeddingsMeasured:nx_distrib_embed exists in runtime/nx_distrib_embed.nx, verified at emit. Co-occurrence semantic vectors Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ● | ● |
PPMI semantic matrixMeasured:nx_semppmi_build exists in runtime/_hdl_build/nx_semppmi_build.nx, verified at emit. Ours explicit sparse; theirs neural Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ◐ | ◐ | ◐ | ◐ |
Trained embeddingsMeasured:nx_embed_train exists in runtime/_hdl_build/nx_embed_train.nx, verified at emit. Ours PPMI-factorization toy; theirs neural leads Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ◉ | ◉ | ◉ | ◉ |
Multi-hop QA readerMeasured:qwx_extract exists in runtime/nx_qwen_extract.nx, verified at emit. Pretrained sovereign Qwen 1-shot + s0-window opt-in: on-protocol SQuAD 510 and Hotpot 505 (above mech-oracle 368) [yang2018hotpot] vs mech 250; MuSiQue 122 ties the composition wall [trivedi2022musique] (NAS-measured, multi-alias). A 2-shot surface-form variant won +104 on the single-gold standalone bench but did NOT transfer to the multi-alias on-protocol metric (SQuAD 427 vs 510) and was reverted (confirmed by nx_reader_ruler_adversary_gate 9/9). Evidence-grounded: SQuAD human F1 86.8 per arxiv:1606.05250 [rajpurkar2016]; published SQuAD v1.1 SOTA = F1 932 (BERT, arxiv:1810.04805) [devlin2018] VERIFIED via nx_research_verify -- our reader 510 = 55 percent of SOTA, which still leads Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder). | ● | ◉ | ◉ | ◉ | ● |
QA F1 / EM scoringMeasured:qs_f1 exists in runtime/nx_qa_score_lib.nx, verified at emit. SQuAD-style measured evaluation Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
Findings ledger (compounding memory)Open — no implementing organ is measured for this axis yet. Persistent provenance ledger | ○ | ● | ● | ● | ◐ |
Corpus semantic retrievalMeasured:nx_lib_semantic exists in runtime/nx_lib_semantic.nx, verified at emit. deep-searcher leads Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ● | ◉ |
RL-trained research agentOpen — watchingruntime/nx_research_agent.nx : ra_policy_step, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. DeepResearcher / Tongyi GRPO-trained lead; Nishi none [shao2024grpo] | ○ | ● | ◉ | ◉ | ○ |
LLM synthesis / reasoningOpen — watchingruntime/nx_llm_synth.nx : ls_synthesize, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (polished synthesis); Nishi extractive only | ○ | ◉ | ● | ● | ● |
Adversarial refutationMeasured:nxr_assess exists in runtime/nx_research_unified.nx, verified at emit. DeepResearcher cross-validates across sources (leads). Nishi now has nx_claim_verify (single) + nx_research_verify (batch STAGE) -- autonomous: fetch a source over sovereign TLS then confirm/discredit/flag-unverifiable a claim vs primary evidence, never fabricating (LIVE-PROVEN: claim 3/3, batch 4/4, and grounded the SQuAD SOTA 932 above) + metric-ruler adversary (nx_reader_ruler_adversary_gate 9/9). Single-source numeric-grounding (catches fabrication/alteration) vs multi-source cross-validation. NOW AUTO-INVOKED in-run (2026-07-14): nx_research_unified (nxr_assess) grades every numeric claim vs its cited source and GATES PUBLISH -- a discredited/fabricated number forces CONTINUE or STOP-INCOMPLETE, never ships; gate-proven nx_research_verify_synth_gate 4/4 + nx_research_unified_gate 5/5 + nx_research_refute_gate 7/7. So PRESENT (auto-invoked, gated). AND multi-source cross-validation is NOW ALSO auto-invoked (2026-07-14b): cvx_answer_verify (nx_research_crossval) inside nxr_assess corroborates each claim across ALL sources with the independence discipline -- >=2 distinct origins required, same-origin echo never counts, conflict surfaced not averaged; crossval gate 10/10 incl a LIVE 3-host proof (GPT-3 175B CORROBORATED across arxiv.org + api.datacite.org while a rate-limited semanticscholar stub was absorbed as unverifiable, not faked). DeepResearcher's agentic multi-source loop still leads on breadth, so NOT exceed [deepresearcher2025] Adoption: LIB-GATE-ONLY importers=3 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program. | ● | ● | ◉ | ● | ○ |
Neural reranker (cross-encoder)Open — watchingruntime/nx_neural_reranker.nx : nr_cross_score, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Trained rerank; Nishi lexical + sparse only | ○ | ● | ● | ● | ◐ |
Multimodal report outputOpen — watchingruntime/nx_multimodal_report.nx : mr_chart_emit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (charts / images); Nishi text only | ○ | ◉ | ○ | ○ | ○ |
Served research API / productOpen — watchingruntime/nx_research_api.nx : rapi_serve, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Perplexity leads (product API); Nishi organ-only [perplexity-api] | ○ | ◉ | ◐ | ◐ | ● |
Sovereign -- own substrate, TLS, index; zero OpenAI / cloud dep; 0 dollar per queryMeasured exceed:ss_fold_cp in runtime/nx_seg_store.nx, verified at emit. All use hosted LLMs; Tongyi open-weights partial [tongyi2025]; Nishi fully sovereign Adoption: LIB-WIRED importers=441 nonval=321 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ◐ | ○ |
Integer-deterministic, crash-resumable, no link-rotMeasured exceed:idf_q10 in runtime/nx_intlog.nx, verified at emit. Reproducible; peers are non-deterministic LLM pipelines Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Capability-gap proposals diffed against our own boardMeasured exceed:fs_deficit in runtime/nx_frontier_scan.nx, verified at emit. LANDED 2026-08-20, nx_frontier_scan_gate GREEN 14 of 14 with two teeth BITE-PROVEN. The organ derives its query seeds from a /compare domain's OWN data (the .axes frontier-keyword field, the .matrix row labels), pulls the outside through the sovereign fetcher over sources that are themselves data [arxiv-api], and emits GAP PROPOSALS naming the row each candidate challenges plus its fetched mirror. THE FOUR RIVAL COLUMNS ARE CODED 0 AS A CATEGORY STATEMENT AND NOT AS A FEATURE AUDIT: all four are answer engines that respond to a question a human asked, none maintains a capability board, so there is no comparable surface to measure them on and none is fabricated here. It emits proposals and NEVER admits one -- the measured law is that no margin threshold makes auto-declaration safe, and a completion signal that keys on a name rewards writing the name. Adoption: RUN-BY:cron — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Frontier noticing across every /compare domain on a beatOpen — watchingruntime/nx_frontier_scan.nx : fs_sweep_all, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN. One domain is on a daily beat today (cron.reg, 07:35) at 12 outbound requests per run, because outbound cost is rows times sources. A fleet sweep over 58 domains at the same settings is roughly 700 requests a day, which is a rate-limit and host-load decision before it is a scheduling one, so the rung must admit a second key-free scholarly index [openalex-api] and a per-source pace budget before it can run wide. Rival columns unmeasured for the category reason given on the row above. | ○ | ○ | ○ | ○ | ○ |
Citation mirror provably bound to the URL it claimsMeasured exceed:crg_prov_state in runtime/nx_compare_refs_gate.nx, verified at emit. LANDED 2026-08-20, nx_compare_refs_gate GREEN 18 of 18. THE CONTRACT WAS DECLARED AS crg_mirror_url_bound AND THE SHIPPED SYMBOL IS crg_prov_state -- the rename is recorded here rather than left to read as a moved goalpost: the contract named a property, the function names the three-state answer. The binding is made at FETCH time, the only moment both facts are in hand -- nx_research_fetch appends a row carrying epoch, requested-url, mirror, sha256, bytes and redirect-hops to knowledge/status/fetch_provenance.jrnl, and this gate joins every declared mirror against it. THREE STATES because two would lie: PROVEN, MISMATCH (this mirror under a DIFFERENT url), and UNPROVEN (no journal row at all, which is every mirror captured before the journal existed and is NOT a defect). Live 2026-08-20: jrnl_rows 29, proven 4, mismatch 0, unproven 617, declared_mirrors 621, and the partition SUMS. MISMATCH is ARMED at zero; UNPROVEN is RATCHETED and self-baselined at 617, because arming it fleet-wide would have turned 68 domains RED in one step and taught everyone to ignore the gate. Every unproven mirror is NAMED in knowledge/status/refs_provenance_unproven.txt, because a count without a worklist is not actionable. Bite-proven in BOTH directions on fixtures assembled at runtime, with a tooth asserting the fixtures actually parsed before any verdict is read off them. RESIDUAL, STATED IN THE GATE VERDICT ITSELF AND CARRIED AS THE ROW BELOW: provenance proves the mirror is OF the url, never that it SUPPORTS the claim. THE HOLE IT CLOSED, kept here so the fix keeps its reason: it is the exact hole the 2026-08-18 metahuman episode fell through. Four references were appended whose URLs were never opened; three carry mirror and pin both absent, and the fourth points at a page fetched three days earlier from a DIFFERENT url -- and the referee passed the file 9 of 9, because it proves the mirror EXISTS and the pin EQUALS the filehash of that mirror, never that the mirror is a capture of the row's own url. A PIN PROVES THE BYTES DID NOT CHANGE, NEVER THAT THEY ARE THE RIGHT DOCUMENT. The same blindness has a second cause now closed: until 2026-08-20 the TLS-1.2 fetch leg truncated bodies near 256 KiB while printing a clean SAVED, so a partial page could be pinned as whole. Rival columns unmeasured for the category reason given two rows above. Adoption: GATE:LIVE trial=- — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Citation mirror provably SUPPORTS the claim it backsOpen — watchingruntime/nx_compare_refs_gate.nx : crg_claim_supported, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN, and it is the half that fetch-time provenance CANNOT close -- named as its own contract rather than left implied, because the row above could easily be read as having ended the class. Provenance proves the mirror is a capture of the url in its row. It says NOTHING about whether that document contains the claim the row cites it for, and on 2026-08-20 that second question had to be answered BY HAND over four mirrors and 22 claims: PRESENT 6, PARTIAL 2, ABSENT 13. Every one of the 13 sat behind a real file with a valid pin, and after this rung ships they would still sit behind a real file with a valid pin AND a correct provenance row. The done-rule is a checker that takes a row grounds field and its mirror and returns SUPPORTED, UNSUPPORTED or UNDECIDABLE with the witnessing byte offsets, ratcheted the same way and armed on nothing until it has a calibrated false-positive rate on the existing corpus. Rival columns unmeasured for the category reason given three rows above. | ○ | ○ | ○ | ○ | ○ |
Frontier sources that are not scholarly indexes -- vendor release notes and conference programmesOpen — watchingruntime/nx_frontier_scan.nx : fs_src_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OPEN, and the reason it is open is a COVERAGE BOUNDARY, not a scoring failure. THE LOOP CANNOT SEE THESE, IT DID NOT MERELY MISS THEM. The two wired sources are arXiv and HN Algolia. Measured against the four references a human added to charsim on 2026-08-18: metahuman57 is an Epic Developer Community release note and rtxhair26 is a GDC schedule session page -- NEITHER INDEX CARRIES EITHER KIND OF PAGE AT ALL; cc5-reallusion is a trade-press article, reachable through HN only on the condition that some human happened to submit it, which is not a capability of the source; sig26-avatars is a SIGGRAPH changelog index page, itself unreachable, though the papers it lists are on arXiv under different citations. SO THE HONEST DENOMINATOR IS ZERO: of the four cited urls, ZERO were in-principle reachable from the sources wired, and the replays 0-of-4 therefore means the loop missed 0 of the 0 it could see. THE MECHANISM ITSELF WORKS WHERE ITS SOURCES CAN SEE -- the 2026-08-20 v2 replay surfaced Show HN SwitchLight, a browser relighting product, against the SKIN full-PBR-and-normal-map-lighting row, unaided and on-topic. PRE-DECLARED ACCEPT RULE, WRITTEN BEFORE THE SOURCE EXISTS SO IT CANNOT BE RATIONALIZED AFTERWARDS: wiring a vendor and release-note source kind is accepted only if replaying the 2026-08-18 PRE-STATE board surfaces the MetaHuman-class surface-layer gap unaided, as at least one proposal whose url is a vendor release note challenging a SKIN, HAIR, MESH or ANIM row, and only if that proposal mirror carries a row in knowledge/status/fetch_provenance.jrnl binding it to that url. THE SECOND CONDITION IS NOT DECORATION: the same audit proved Epic SILENTLY REPOINTED the metahuman-5-7 slug to the 5.8 notes, so a vendor page can change identity underneath a citation, and a vendor source without fetch-time provenance would rebuild the exact hole this lane just spent a day retracting. Rival columns unmeasured for the category reason given four rows above. | ○ | ○ | ○ | ○ | ○ |
Risk register
| Risk | Likelihood x impact | Mitigation |
|---|---|---|
| Synthesis that reads well hides a wrong number better than extractive text did | likely x high | DR2 is graded sentence by sentence by the existing referee; fluency never bypasses the gate. |
| A learned reranker or policy that wins on the training ruler and loses held-out is an illusion of progress | possible x high | DR3 and DR5 carry pre-declared accept rules on held-out sets; a miss stays UNWIRED with the number published. |
Person · product · place — not yet measured for this domain
knowledge/compare/deepresearch.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain deepresearch, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [perplexity-api] Perplexity. API Platform overview -- Router, Agent, Search and Embeddings APIs (real-time web-grounded answers). Vendor documentation. publisher · read in our library
knowledge/fetched/cmp_deepresearch_perplexity.html· pinhae63a992df31746dde53a7ba823cfe0479d11e091814e8fcaaf8dd86c7bc0e46· accessed 2026-08-18 · vendor-docGrounds: The Perplexity column: Best on "Served research API / product", "LLM synthesis / reasoning" and "Multimodal report output" -- the documented product API is the served-answer-engine capability those codes describe, and it is a hosted LLM service, which is the foil for the "Sovereign -- own substrate, TLS, index; zero OpenAI / cloud dep" exceed row. - [deepresearcher2025] Zheng, Y., Fu, D., Hu, X., Cai, X., Ye, L., Lu, P., Liu, P. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments. arXiv:2504.03160, April 2025. publisher · read in our library
knowledge/fetched/cmp_deepresearch_deepresearcher2025.html· pinhfc6c302637b8af3e2c88af6cc438c81002200150b0c2641ad1fb7b62fc4f0549· accessed 2026-08-18 · published-paperGrounds: The DeepResearcher column: Best on "RL-trained research agent" and "Adversarial refutation" -- the abstract describes end-to-end RL training of deep-research agents in real web environments and reports emergent cross-validation of information across multiple sources, which is exactly the multi-source breadth the row note concedes still leads ours. - [tongyi2025] Tongyi DeepResearch Team, Li, B., Zhang, B. et al. Tongyi DeepResearch Technical Report. arXiv:2510.24701, October 2025 (30.5 billion total parameters, 3.3 billion activated per token; model, framework and solutions open-sourced). publisher · read in our library
knowledge/fetched/cmp_deepresearch_tongyi2025.html· pinh24ac13b5b8fafe85408e0b6c254151046e31913659f1c4fc1a2afa52bf3b35e8· accessed 2026-08-18 · published-paperGrounds: The Tongyi column: Best on "RL-trained research agent" (agentic mid-training plus post-training) and Part on the "Sovereign" row ("Tongyi open-weights partial") -- the report states the model is open-sourced, which is why its column is the only rival with a non-zero code on that exceed row. - [deepsearcher] Zilliz. deep-searcher -- open source deep research alternative to reason and search on private data, written in Python (GitHub repository, zilliztech). publisher · read in our library
knowledge/fetched/cmp_deepresearch_deepsearcher.html· pinhf906e0ee916636ecf7db998d4a955ff020af6a15cb02bb76220a32c1f768d46b· accessed 2026-08-18 · source-readGrounds: The deep-searcher column: Best on "Dense vector retrieval" and "Corpus semantic retrieval" (the note: Milvus / Zilliz vector DB) -- the repository is the vector-database-backed private-data research agent those codes describe. - [rajpurkar2016] Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P. SQuAD: 100,000+ Questions for Machine Comprehension of Text. EMNLP 2016; arXiv:1606.05250. Human performance F1 86.8 percent (abstract). publisher · read in our library
knowledge/fetched/cmp_deepresearch_rajpurkar2016.html· pinhec5af640421c7c422bd13b49c0d2deac62a706d71047fad2304930cc756f48ce· accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" and "QA F1 / EM scoring" rows: the SQuAD protocol (span extraction, F1 and EM) is the estate's reader ruler, and the row note's "SQuAD human F1 86.8 per arxiv:1606.05250" is quoted from this abstract. - [devlin2018] Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2019; arXiv:1810.04805. SQuAD v1.1 Test F1 93.2 (abstract). publisher · read in our library
knowledge/fetched/cmp_deepresearch_devlin2018.html· pinhffa5d3591fa543ce7b833bea05f1455d40a6be8680417720f09da8807ada6abe· accessed 2026-08-18 · published-paperGrounds: The "Multi-hop QA reader" row note: "published SQuAD v1.1 SOTA = F1 932 (BERT, arxiv:1810.04805)" -- the 93.2 F1 is quoted from this abstract, and our reader's 510 is graded as 55 percent of it, honestly BEHIND. - [yang2018hotpot] Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Manning, C.D. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. EMNLP 2018; arXiv:1809.09600. publisher · read in our library
knowledge/fetched/cmp_deepresearch_yang2018hotpot.html· pinh1d2aba7b7aa74242566af15165d3861311bc478db918ebf7a28ebc94d8b20ffd· accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" row: the Hotpot 505 figure in the note is measured on this dataset's multi-hop protocol; it is the second of the three benchmarks the reader ruler runs. - [trivedi2022musique] Trivedi, H., Balasubramanian, N., Khot, T., Sabharwal, A. MuSiQue: Multihop Questions via Single-hop Question Composition. TACL 2022; arXiv:2108.00573. publisher · read in our library
knowledge/fetched/cmp_deepresearch_trivedi2022musique.html· pinhd8ca40a3303d9287245cad135d31a338a646f67e97764dd3ce4c59c321a578c7· accessed 2026-08-18 · datasetGrounds: The "Multi-hop QA reader" row: "MuSiQue 122 ties the composition wall" -- MuSiQue is built by composing single-hop questions so that shortcut readers fail, which is the wall the note names; the third benchmark of the reader ruler. - [shao2024grpo] Shao, Z., Wang, P., Zhu, Q. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024 -- introduces Group Relative Policy Optimization (GRPO), a PPO variant (abstract). publisher · read in our library
knowledge/fetched/cmp_deepresearch_shao2024grpo.html· pinh32e37ca6d2af1522b1372c824d8ace3f400f44b3701a3cc69fd624254a39aeb5· accessed 2026-08-18 · published-paperGrounds: The "RL-trained research agent" row note ("DeepResearcher / Tongyi GRPO-trained lead; Nishi none"): GRPO is defined in this paper, so the training method the row credits the leaders with has a single published definition, and the _ABSENT_ nx_research_agent watch names what a sovereign no-float RL loop must implement. - [arxiv-api] arXiv. arXiv API User's Manual (info.arxiv.org): the search_query grammar with field prefixes such as all: ti: au:, the Boolean operators AND OR ANDNOT, sortBy submittedDate, and the Atom feed schema returned by the query endpoint. publisher · read in our library
knowledge/fetched/cmp_deepresearch_arxivapi.html· pinhf6c3ce2936f629e03fd42faf911604a7b0677ae3475c9c8608a543e9a7579671· accessed 2026-08-20 · vendor-docGrounds: The "Capability-gap proposals diffed against our own board" row and knowledge/frontier_sources.conf: this manual is why every source row carries an explicit term-join operator. A space-joined search_query does NOT AND its terms, so an unjoined seed returns the newest submissions of the whole archive with HTTP 200 -- measured on the first metahuman replay, which proposed a Bose-Einstein condensate paper against a skin-shader row. - [openalex-api] OpenAlex. API Overview (docs.openalex.org): a free and open catalogue of scholarly works behind a public REST API that requires no key, with works filterable and sortable by publication date. publisher · read in our library
knowledge/fetched/cmp_deepresearch_openalex.html· pinha4405ec3a16a534736c39084c43a6c945b3850c27c080a0c4eed4798cd7ae0bf· accessed 2026-08-20 · vendor-docGrounds: The "Frontier noticing across every /compare domain on a beat" watch row: a fleet sweep is a rate-limit problem before it is a scheduling problem, and a second key-free scholarly index is what makes the wide sweep affordable without hammering a single source. It names the source row that rung must admit.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/deepresearch.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers