Nishi Experiential Testing — a SOTA-grounded census

An honest, evidence-grounded grade of where Nishi's experiential UX/CX testing stands against the real state of the art. Produced by the Nishi verification stack — researcher (sovereign-fetched competitor evidence), census (disk facts x that evidence, never self-scored), adversary (discipline teeth that forbid a fake pass), critic (verdict + keystone). THIS PAGE IS EMITTED BY THE CENSUS RUN (nx_s21_census) — it cannot drift from disk truth.

GAP 11PARITY 8AHEAD 019 attested axes
Verdict: the experiential surface is climbing but NOT closed: the keystone that closes ~9 of the remaining GAP axes at once is one unbuilt organ, nx_browser_drive (input drive over our stack). Every PRESENT organ caps at PARITY until a measured head-to-head beats the tool — a self-gate is necessary, not sufficient. The full build-out ladder with per-axis wiring (API + MCP + workstream + agent + RACI) lives at /evidence/experiential-testing.

The census

Each axis: how many of the 8 fetched SOTA-tool READMEs attest the capability (≥2 = corroborated), vs whether our implementing organ opens on disk. our=PRESENT caps the verdict at PARITY — AHEAD requires a measured head-to-head gate we have not built.

capability axiscompetitor (of 8)our organverdict
drive browser (headless control)5present (nx_browser_drive + nishi_gui drive lane)PARITY
cross-browser engines4absentGAP
WebKit engine1absentGAP
mobile / device emulation2absentGAP
screenshot capture (organ)4present (nishi_gui nxshot recording lane)PARITY
input drive (click / type)5present (nx_browser_drive (click/scroll/open; type queued))PARITY
auto-wait / retry5present (nx_browser_drive wait verb)PARITY
network interception3absentGAP
console / error capture6absentGAP
trace / replay2absentGAP
visual diff / regression2present (nx_visual_diff)PARITY
accessibility audit (WCAG)3absentGAP
performance audit (web vitals)2absentGAP
agentic goal-directed play4absentGAP
live-run intelligence / findings3present (nx_ux_study)PARITY
CI / sweep orchestration5absentGAP
video / media playtest4absentGAP
automated page / asset verify7present (nx_page_verify)PARITY
UI / CX quality judging3present (nx_uxcx_grade + nx_ux_study)PARITY

The adversary (why this can't be gamed)

Three discipline teeth run with the census and must all pass, or it fails closed:

The evidence (researcher)

Competitor capability is attested from live, sovereign-fetched sources (our own TLS 1.3, no third-party HTTP client), not from memory: playwright · puppeteer · cypress · selenium · axe-core · lighthouse · backstop.js · browser-use — public READMEs fetched into knowledge/census/*.raw. The census greps these for each capability term; >2 distinct tools = corroborated.

The critic (the path forward)

The honest read: the seeing/judging half is landing (screenshots, recordings, visual diff, page verify, UX study findings — all disk-real today), while the DRIVING half is the gap. One organ, nx_browser_drive, closes the drive-rooted axes together. The full ladder — every axis, its status, the number it moves, and its API/MCP/workstream/agent/RACI wiring — is maintained as data at /evidence/experiential-testing, and the dated proof runs live under /evidence.