Nishi Experiential Testing — a SOTA-grounded census
An honest, evidence-grounded grade of where Nishi's experiential UX/CX testing stands against the real state of the art. Produced by the Nishi verification stack — researcher (sovereign-fetched competitor evidence), census (disk facts x that evidence, never self-scored), adversary (discipline teeth that forbid a fake pass), critic (verdict + keystone). THIS PAGE IS EMITTED BY THE CENSUS RUN (nx_s21_census) — it cannot drift from disk truth.
nx_browser_drive (input drive over our stack). Every PRESENT organ caps at PARITY until a measured head-to-head beats the tool — a self-gate is necessary, not sufficient. The full build-out ladder with per-axis wiring (API + MCP + workstream + agent + RACI) lives at /evidence/experiential-testing.The census
Each axis: how many of the 8 fetched SOTA-tool READMEs attest the capability (≥2 = corroborated), vs whether our implementing organ opens on disk. our=PRESENT caps the verdict at PARITY — AHEAD requires a measured head-to-head gate we have not built.
| capability axis | competitor (of 8) | our organ | verdict |
|---|---|---|---|
| drive browser (headless control) | 5 | present (nx_browser_drive + nishi_gui drive lane) | PARITY |
| cross-browser engines | 4 | absent | GAP |
| WebKit engine | 1 | absent | GAP |
| mobile / device emulation | 2 | absent | GAP |
| screenshot capture (organ) | 4 | present (nishi_gui nxshot recording lane) | PARITY |
| input drive (click / type) | 5 | present (nx_browser_drive (click/scroll/open; type queued)) | PARITY |
| auto-wait / retry | 5 | present (nx_browser_drive wait verb) | PARITY |
| network interception | 3 | absent | GAP |
| console / error capture | 6 | absent | GAP |
| trace / replay | 2 | absent | GAP |
| visual diff / regression | 2 | present (nx_visual_diff) | PARITY |
| accessibility audit (WCAG) | 3 | absent | GAP |
| performance audit (web vitals) | 2 | absent | GAP |
| agentic goal-directed play | 4 | absent | GAP |
| live-run intelligence / findings | 3 | present (nx_ux_study) | PARITY |
| CI / sweep orchestration | 5 | absent | GAP |
| video / media playtest | 4 | absent | GAP |
| automated page / asset verify | 7 | present (nx_page_verify) | PARITY |
| UI / CX quality judging | 3 | present (nx_uxcx_grade + nx_ux_study) | PARITY |
The adversary (why this can't be gamed)
Three discipline teeth run with the census and must all pass, or it fails closed:
- No false parity — an organ that isn't on disk, against a capability the competitors attest, must read GAP.
- No self-gate inflation — an organ that exists but has no measured head-to-head gate reads PARITY, never AHEAD. A gate proving our thing works is necessary, not sufficient; only beating the tool earns AHEAD.
- No phantom capability — a bogus organ path grades ABSENT.
The evidence (researcher)
playwright · puppeteer · cypress · selenium · axe-core · lighthouse · backstop.js · browser-use — public READMEs fetched into knowledge/census/*.raw. The census greps these for each capability term; >2 distinct tools = corroborated.The critic (the path forward)
The honest read: the seeing/judging half is landing (screenshots, recordings, visual diff, page verify, UX study findings — all disk-real today), while the DRIVING half is the gap. One organ, nx_browser_drive, closes the drive-rooted axes together. The full ladder — every axis, its status, the number it moves, and its API/MCP/workstream/agent/RACI wiring — is maintained as data at /evidence/experiential-testing, and the dated proof runs live under /evidence.