Nishi FamilyCompare › Automated Testing and Competitive Intel (Racing Crew)

Nishi Compare · measured, not asserted

Automated Testing and Competitive Intel (Racing Crew)

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

Nishi vs Playwright and Cypress and Selenium and Applitools

Layer 1 · Executive

Where we are. Measured 2026-08-19. The crew tests the SHIPPED BYTES (wasm on our own VM, vision forensics, seeded repro) and runs competitive intel as a primitive; what it lacks is the interactive half of the field: no real-browser driver, no locator API, no mocking, no flaky quarantine, no codegen, no CI surface, no a11y organ. The a11y rulers (nx_uiq_atree, nx_uiq_contrast, page_verify a11y count) exist separately and the driver precedent exists as a laptop PowerShell wire -- both are composition rungs, not inventions.

Where we need to go. Drive real pages from an organ, script journeys on it, and turn the gate journals into the CI surface -- then scale over the swarm and learn selectors and triage on the sovereign seat, keeping the shipped-bytes harness as the exceed the browser-automation field does not have.

The unit. 1 u = one measured session-leg. Calibration from landed rungs: the mangagen panel compositor went from existing substrate to shipped and live-verified in ONE leg (2026-08-13); the citations rung went from 3 to 55 domains in one leg across seven seats (2026-08-18); a greenfield engine with a bite-proven gate has measured 2 to 4 legs. Estimates recalibrate as rungs land and PR7 actuals write back.
Cost to drive real pages: 5 u. Through M0. The driver is the long pole and unblocks five later rungs.
Cost to production parity: 11.5 u. Through M2.
Cost to scale and learn: 18.5 u. Everything below.

10 of 20 capabilities measured|2 of them measured exceeds|10 open|coverage 500/1000|adoption 2 full / 8 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883514 domain=racing target_version=0.1 rungs=10 done=0 open=10 finish=0 ranker=nx_dr_ocm

#StageRungPriorityDerivation
#10.1Locator API with auto-wait (R1) e2e_locate850v=17 m=1 c=20
#20.1Sovereign browser driver (CDP and the Nishi browser) (R0) wd_session666v=20 m=1 c=30
#3laterCI reports from the gate journals (R4) tr_report2000v=20 m=1 c=10
#4laterFlaky detection and quarantine (R3) fl_quarantine900v=9 m=1 c=10
#5laterCross-browser grid on the swarm (R8) tg_run700v=14 m=1 c=20
#6laterLLM test authoring and triage (R9) lt_author700v=14 m=1 c=20
#7laterNetwork interception and mocking (R2) nm_route666v=10 m=1 c=15
#8laterRecord and codegen (R6) tc_record450v=9 m=1 c=20
#9laterAccessibility audit organ (R5) aa_audit200v=2 m=1 c=10
#10laterVisual self-healing locators (R7) va_heal200v=6 m=1 c=30

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Sovereign browser driver (CDP and the Nishi browser) (R0)wd_sessionAn organ drives a real page -- navigate, click, type, wait, screenshot -- over CDP to Chromium and natively to the sovereign browser; the laptop-side nx_cdp.ps1 wire that proved WebGPU is the break-glass precedent this rung retires into an organOrgan3 u
Locator API with auto-wait (R1)
after R0
e2e_locateRole, text and CSS locators with auto-wait semantics over the driver; a scripted journey on /compare and /survey runs green and a planted DOM change fails itOrgan2 u
Flaky detection and quarantine (R3)fl_quarantineRe-run statistics over the gate roster journal name flaky gates by measurement; quarantined gates still run and report, they stop blocking; the list is printed, never a count aloneOrgan1 u
CI reports from the gate journals (R4)tr_reportJUnit-class XML plus a per-run page derived from gateroster.jrnl and biteall.log; every row links to its gate and seedOrgan1 u
Accessibility audit organ (R5)aa_auditOne axe-class report per page composing nx_uiq_atree, nx_uiq_contrast and the page_verify a11y count; runs over every compare page on the beat; a planted contrast defect turns it REDOrgan1 u
Network interception and mocking (R2)
after R0
nm_routeRoute rules intercept and stub requests in the driven page; a journey runs against a mocked backend deterministicallyOrgan1.5 u
Record and codegen (R6)
after R1
tc_recordUser actions recorded through the driver emit a replayable journey script; replay is byte-equal to the recording on the same buildOrgan2 u
Visual self-healing locators (R7)
after R1
va_healA learned visual matcher re-finds a moved element and the healed locator is reported with its evidence; precision on a banked layout-drift set PRE-DECLAREDLocal model3 u
Cross-browser grid on the swarm (R8)
after R0
tg_runJourneys fan out over the mesh to Chromium, Firefox and the sovereign browser; a matrix of results with every cell a linkOrgan2 u
LLM test authoring and triage (R9)
after R1
lt_authorThe sovereign seat drafts journeys from a page and triages failures by seed and frame; every draft runs before it is admittedLocal model2 u

Milestones

MilestoneRungsCumulative
M0 · Drive real pagesR0,R15 u
M1 · Suite hygieneR3,R4,R58 u
M2 · Script everythingR2,R611.5 u
M3 · Scale and learnR7,R8,R918.5 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiPlaywrightCypressSeleniumApplitools
Vision forensics on the SHIPPED artifact (not source, not a proxy)Measured: abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. Executes the real wasm on nx_wasm_vm and reads the framebuffer: blackout/frozen/magenta-leak/sky-ground structure; Applitools does screenshot AI-diff, but on browser captures not on the shipped bytes in our own VM Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Seeded-repro fuzz determinism (any red = exact repro)Measured: abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. LCG seed printed in every finding; frames 1-6 byte-identical across runs; the loop auto-found + localized 2 miscalibrations on runs 1 and 2 (2026-07-03) Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
No-reference perceptual image-quality judgeMeasured: ns_mscn_rho exists in runtime/nx_natstat.nx, verified at emit. MSCN / NIQE-class 5-axis natural-scene statistics (reverse-AI photoreal judge) [niqe13]; Applitools visual AI leads on learned diff Adoption: LIB-WIRED importers=9 nonval=5 — fully adopted (top of its ladder).
State invariants per tickMeasured: abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. camera-y bounds + pitch-clamp checked every tick, not just at frame samples Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Competitive-intel delta ledger (survey incumbents, drive roadmap)Measured: nx_cint_record exists in runtime/nx_competitive_intel.nx, verified at emit. Per-incumbent delta ledger with a 90-day survey cadence and next-due tracking (nx_cint_record/set_next_due); no test tool ships strategic competitive-intel as a primitive Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Measured pitwall board (1:1 no-proxy)Measured: nx_pitwall_stage exists in runtime/nx_pitwall_render.nx, verified at emit. The racing-crew scoreboard: measured position vs named incumbents, no proxy (nxasm 4x/size/determinism triple-exceed found + gated 2026-05-29) Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program.
adoption LIB-GATE-ONLY importers=1
Liar-killed competitive census planeMeasured: LIAR-KILL exists in runtime/nx_swcompare_matrix.nx, verified at emit. 32 live head-to-head comparisons; every Nishi cell requires the symbol on disk, fabricated cells die at build -- the racing crew's find-gaps-vs-competitors engine, productized Adoption: LIVE — fully adopted (top of its ladder).
UX / UI composition gradingMeasured: pmil exists in runtime/nx_ui_judge.nx, verified at emit. nx_ui_judge scores contrast/tokens/mobile; the field pairs with axe/Lighthouse, Applitools does visual UI diff Adoption: PROMOTED-UNREGISTERED — PARTIAL: a real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption PROMOTED-UNREGISTERED
Real-browser WebDriver / CDP driverOpen — watching runtime/nx_webdriver.nx : wd_session, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright CDP [cdp] + Selenium WebDriver [w3c-webdriver] are the bar; our harness runs wasm directly, the JS/canvas/DOM layer is outside it by construction -- the named next rung
watching wd_session
Selector-based E2E scripting APIOpen — watching runtime/nx_e2e_api.nx : e2e_locate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright locators + auto-wait are the bar [playwright-autowait]; we fuzz inputs, we do not script user journeys
watching e2e_locate
Visual-AI self-healing selectorsOpen — watching runtime/nx_visual_ai.nx : va_heal, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Applitools self-healing locators + AI diff are the bar [applitools-eyes]; our vision is heuristic forensics, not a learned model
watching va_heal
Cross-browser cloud test gridOpen — watching runtime/nx_test_grid.nx : tg_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Selenium Grid + BrowserStack are the bar [selenium-grid]; we run one interpreted VM
watching tg_run
Network interception and mockingOpen — watching runtime/nx_net_mock.nx : nm_route, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright route-mocking is the bar; absent here
watching nm_route
Flaky-test detection and quarantineOpen — watching runtime/nx_flaky.nx : fl_quarantine, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Every mature suite quarantines flakies [luo14]; ours is deterministic-by-seed so flakiness is structurally rarer, but no detector organ exists
watching fl_quarantine
Auto-record test codegenOpen — watching runtime/nx_test_codegen.nx : tc_record, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright codegen + Cypress Studio record user actions into tests [cypress-studio]; absent here (ties to the team census macro-recorder gap)
watching tc_record
CI reporting and dashboardsOpen — watching runtime/nx_test_report.nx : tr_report, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. JUnit/Allure reporting is the field standard; our forensics tables are console + the /compare board, no per-run CI surface
watching tr_report
Accessibility auto-auditOpen — watching runtime/nx_a11y_audit.nx : aa_audit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. axe-core / Lighthouse a11y are the bar [axe-core]; ties to the UI-quality doctrine but no automated a11y organ
watching aa_audit
LLM-driven test authoring and triageOpen — watching runtime/nx_llm_test.nx : lt_author, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The cross-roster convergence gap (5th radar to point here); nobody in this field fully ships it either
watching lt_author
Tests the shipped bytes on our own VM (sovereign, no external browser)Measured exceed: abug in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. The harness executes the ACTUAL deployed wasm in nx_wasm_vm and inspects VM memory + framebuffer host-side -- zero external browser, zero proxy, bit-exact; the field drives a real browser it does not own Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Survey-incumbents-and-drive-the-roadmap as a standing primitiveMeasured exceed: nx_cint_record in runtime/nx_competitive_intel.nx, verified at emit. The racing doctrine is a compiled ledger + cadence + pitwall + the 32-comparison plane: honest delta-to-incumbent drives every roadmap; the test-automation field ships no strategic competitive-intel layer Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The Racing crew is the meta-team the operator named -- auto-find bugs, auto-resolve, and race every effort against named competitors by default -- and its most distinctive capabilities are ones the browser-automation field simply does not have. It tests the SHIPPED BYTES: seeded-fuzz + vision forensics execute the actual wasm on our own VM and LOOK at the rendered framebuffer (blackout, frozen-frame, magenta-leak, sky-ground structure), so any red is an exact seed+tick repro; the loop auto-found and localized two real miscalibrations on its first two runs. And competitive-intel is a STANDING PRIMITIVE: a delta-to-incumbent ledger on a 90-day survey cadence plus the pitwall board and the 32-comparison liar-killed /compare plane -- surveying incumbents to drive the roadmap, which no test tool ships. What it lacks is the field's entire interactive half: no real-browser WebDriver/CDP driver, no selector E2E scripting, no visual-AI self-healing locators (Applitools), no cross-browser cloud grid (the JS/canvas layer is outside our harness by construction), no flaky-test quarantine, no record-codegen, no accessibility audit. The climb: a sovereign WebDriver/CDP driver so the harness drives real pages, wire competitive-intel to auto-open a track per census, then flaky-detection and a11y.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/racing.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain racing, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [playwright-autowait] Microsoft. Playwright documentation, Auto-waiting: actionability checks (visible, stable, receives events, enabled, editable) run before every locator action, failing with TimeoutError instead of flaking. publisher · read in our library knowledge/fetched/cmp_racing_playwright-autowait.html · pin had67cb07cdb4baaf7b3a553bd4c6a64219114e937a78ad3444a2ec213a81f32a · accessed 2026-08-18 · vendor-docGrounds: The Playwright column: "Selector-based E2E scripting API" (the row's "Playwright locators + auto-wait are the bar") and "Real-browser WebDriver / CDP driver" (Playwright CDP) -- the documented behaviour our absent nx_e2e_api and nx_webdriver are graded against.
  2. [cypress-studio] Cypress.io. Cypress Studio (Cypress Documentation): record clicks, typing and form interactions in the running app into Cypress test code, with AI-suggested assertions reviewed before saving. publisher · read in our library knowledge/fetched/cmp_racing_cypress-studio.html · pin hc67cd97e9b99a894146f3355d6711dd8fbf55f6d8cbcff6195237f99bab05140 · accessed 2026-08-18 · vendor-docGrounds: The Cypress column: "Auto-record test codegen" (the row's "Playwright codegen + Cypress Studio record user actions into tests") -- the recorder our nx_test_codegen is _ABSENT_ against, and the tie to the team census macro-recorder gap.
  3. [selenium-grid] Selenium project. Selenium Grid documentation: WebDriver scripts executed on remote machines by routing client commands to remote browser instances -- parallel runs across many machines, browser versions and operating systems. publisher · read in our library knowledge/fetched/cmp_racing_selenium-grid.html · pin h51f2d209d9994c6ac5fb133b847782ebdf5bf888d5632a539186991dec505254 · accessed 2026-08-18 · vendor-docGrounds: The Selenium column: "Cross-browser cloud test grid" (the row's "Selenium Grid + BrowserStack are the bar") -- distributed multi-browser execution versus our single interpreted VM.
  4. [w3c-webdriver] W3C. WebDriver. W3C Recommendation, 05 June 2018 -- "a remote control interface that enables introspection and control of user agents"; the wire protocol Selenium implements (WebDriver Level 2 continues as a Working Draft, 02 July 2026). publisher · accessed 2026-08-18 · published-standardGrounds: "Real-browser WebDriver / CDP driver" -- the row's "Selenium WebDriver" bar is this standard; a sovereign driver speaking it is the named next rung. Mirror absent by declaration: the sovereign fetch of w3.org returned a 5.8KB truncated body.
  5. [cdp] Chrome DevTools team (Google). Chrome DevTools Protocol: the domains, methods and events that let tools instrument, inspect, debug and profile Chromium-based browsers -- the protocol Playwright and Puppeteer drive. publisher · read in our library knowledge/fetched/cmp_racing_cdp.html · pin hf4119c5a11315df0f83d4968bd77b0a5539408f527a2c1aa9d75ef481c40cd02 · accessed 2026-08-18 · vendor-docGrounds: "Real-browser WebDriver / CDP driver" -- the CDP half of the row's "Playwright CDP + Selenium WebDriver are the bar"; the JS/canvas/DOM layer our wasm harness cannot reach is exactly what this protocol exposes.
  6. [applitools-eyes] Applitools. Applitools Documentation: Eyes Visual AI added to an existing test framework, the Ultrafast Grid for cross-browser visual runs, and the Execution Cloud whose self-healing lets tests run even when elements have changed. publisher · read in our library knowledge/fetched/cmp_racing_applitools.html · pin h226005936c94ab2229ea8aac907eb46988ad57adf9c99129da51171668407507 · accessed 2026-08-18 · vendor-docGrounds: The Applitools column: "Visual-AI self-healing selectors" (the row's "Applitools self-healing locators + AI diff are the bar"), "Vision forensics on the SHIPPED artifact" (Applitools screenshot AI-diff on browser captures, code 3) and "UX / UI composition grading" -- learned visual diff versus our heuristic framebuffer forensics.
  7. [niqe13] Mittal, Soundararajan, Bovik. Making a Completely Blind Image Quality Analyzer (NIQE). IEEE Signal Processing Letters, pp. 209-212, March 2013; with Mittal, Moorthy, Bovik, No-Reference Image Quality Assessment in the Spatial Domain (BRISQUE), IEEE TIP 2012 -- both listed on the LIVE lab no-reference IQA page. publisher · read in our library knowledge/fetched/cmp_racing_niqe13.html · pin hff89949bea4a7961618db4a8fb32f153aeaac5566a3ee3f9243df4f42fc735b0 · accessed 2026-08-18 · published-paperGrounds: "No-reference perceptual image-quality judge" -- the row's "MSCN / NIQE-class 5-axis natural-scene statistics": MSCN coefficients and the NIQE natural-scene-statistics model are these papers, the published basis of nx_natstat ns_mscn_rho.
  8. [axe-core] Deque Systems. axe-core: the accessibility engine for automated Web UI testing (GitHub repository and rule documentation). publisher · read in our library knowledge/fetched/cmp_racing_axe-core.html · pin h2ae56ef394658f8192bb22cc56f617bd0ec0189d7acef4596514e961be6db6af · accessed 2026-08-18 · vendor-docGrounds: "Accessibility auto-audit" -- the row's "axe-core / Lighthouse a11y are the bar"; the rule engine our absent nx_a11y_audit is graded against, tied to the UI-quality doctrine.
  9. [luo14] Luo, Hariri, Eloussi, Marinov. An Empirical Analysis of Flaky Tests. FSE 2014 (22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering), pp. 643-653 -- 201 flaky-fix commits across 51 projects, root-cause taxonomy and fix strategies. publisher · read in our library knowledge/fetched/cmp_racing_luo14.pdf · pin h36abbab3e917ad38300e8956a0311004423e4dcfede766b1bf6081b02fe1823b · accessed 2026-08-18 · published-paperGrounds: "Flaky-test detection and quarantine" -- the published root-cause taxonomy (async waits, concurrency, test-order dependence) behind why every mature suite quarantines flakies; our seed-deterministic harness makes the class structurally rarer, but no detector organ exists.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/racing.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers