Nishi Family › Compare › Automated Testing and Competitive Intel (Racing Crew)
Nishi Compare · measured, not asserted
Automated Testing and Competitive Intel (Racing Crew)
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs Playwright and Cypress and Selenium and Applitools
Where we are. Measured 2026-08-19. The crew tests the SHIPPED BYTES (wasm on our own VM, vision forensics, seeded repro) and runs competitive intel as a primitive; what it lacks is the interactive half of the field: no real-browser driver, no locator API, no mocking, no flaky quarantine, no codegen, no CI surface, no a11y organ. The a11y rulers (nx_uiq_atree, nx_uiq_contrast, page_verify a11y count) exist separately and the driver precedent exists as a laptop PowerShell wire -- both are composition rungs, not inventions.
Where we need to go. Drive real pages from an organ, script journeys on it, and turn the gate journals into the CI surface -- then scale over the swarm and learn selectors and triage on the sovereign seat, keeping the shipped-bytes harness as the exceed the browser-automation field does not have.
10 of 20 capabilities measured|2 of them measured exceeds|10 open|coverage 500/1000|adoption 2 full / 8 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883514 domain=racing target_version=0.1 rungs=10 done=0 open=10 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | 0.1 | Locator API with auto-wait (R1) e2e_locate | 850 | v=17 m=1 c=20 |
| #2 | 0.1 | Sovereign browser driver (CDP and the Nishi browser) (R0) wd_session | 666 | v=20 m=1 c=30 |
| #3 | later | CI reports from the gate journals (R4) tr_report | 2000 | v=20 m=1 c=10 |
| #4 | later | Flaky detection and quarantine (R3) fl_quarantine | 900 | v=9 m=1 c=10 |
| #5 | later | Cross-browser grid on the swarm (R8) tg_run | 700 | v=14 m=1 c=20 |
| #6 | later | LLM test authoring and triage (R9) lt_author | 700 | v=14 m=1 c=20 |
| #7 | later | Network interception and mocking (R2) nm_route | 666 | v=10 m=1 c=15 |
| #8 | later | Record and codegen (R6) tc_record | 450 | v=9 m=1 c=20 |
| #9 | later | Accessibility audit organ (R5) aa_audit | 200 | v=2 m=1 c=10 |
| #10 | later | Visual self-healing locators (R7) va_heal | 200 | v=6 m=1 c=30 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Sovereign browser driver (CDP and the Nishi browser) (R0) | wd_session | An organ drives a real page -- navigate, click, type, wait, screenshot -- over CDP to Chromium and natively to the sovereign browser; the laptop-side nx_cdp.ps1 wire that proved WebGPU is the break-glass precedent this rung retires into an organ | Organ | 3 u |
| Locator API with auto-wait (R1) after R0 | e2e_locate | Role, text and CSS locators with auto-wait semantics over the driver; a scripted journey on /compare and /survey runs green and a planted DOM change fails it | Organ | 2 u |
| Flaky detection and quarantine (R3) | fl_quarantine | Re-run statistics over the gate roster journal name flaky gates by measurement; quarantined gates still run and report, they stop blocking; the list is printed, never a count alone | Organ | 1 u |
| CI reports from the gate journals (R4) | tr_report | JUnit-class XML plus a per-run page derived from gateroster.jrnl and biteall.log; every row links to its gate and seed | Organ | 1 u |
| Accessibility audit organ (R5) | aa_audit | One axe-class report per page composing nx_uiq_atree, nx_uiq_contrast and the page_verify a11y count; runs over every compare page on the beat; a planted contrast defect turns it RED | Organ | 1 u |
| Network interception and mocking (R2) after R0 | nm_route | Route rules intercept and stub requests in the driven page; a journey runs against a mocked backend deterministically | Organ | 1.5 u |
| Record and codegen (R6) after R1 | tc_record | User actions recorded through the driver emit a replayable journey script; replay is byte-equal to the recording on the same build | Organ | 2 u |
| Visual self-healing locators (R7) after R1 | va_heal | A learned visual matcher re-finds a moved element and the healed locator is reported with its evidence; precision on a banked layout-drift set PRE-DECLARED | Local model | 3 u |
| Cross-browser grid on the swarm (R8) after R0 | tg_run | Journeys fan out over the mesh to Chromium, Firefox and the sovereign browser; a matrix of results with every cell a link | Organ | 2 u |
| LLM test authoring and triage (R9) after R1 | lt_author | The sovereign seat drafts journeys from a page and triages failures by seed and frame; every draft runs before it is admitted | Local model | 2 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M0 · Drive real pages | R0,R1 | 5 u |
| M1 · Suite hygiene | R3,R4,R5 | 8 u |
| M2 · Script everything | R2,R6 | 11.5 u |
| M3 · Scale and learn | R7,R8,R9 | 18.5 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | Playwright | Cypress | Selenium | Applitools |
|---|---|---|---|---|---|
Vision forensics on the SHIPPED artifact (not source, not a proxy)Measured:abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. Executes the real wasm on nx_wasm_vm and reads the framebuffer: blackout/frozen/magenta-leak/sky-ground structure; Applitools does screenshot AI-diff, but on browser captures not on the shipped bytes in our own VM Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ○ | ◐ |
Seeded-repro fuzz determinism (any red = exact repro)Measured:abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. LCG seed printed in every finding; frames 1-6 byte-identical across runs; the loop auto-found + localized 2 miscalibrations on runs 1 and 2 (2026-07-03) Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ● | ○ |
No-reference perceptual image-quality judgeMeasured:ns_mscn_rho exists in runtime/nx_natstat.nx, verified at emit. MSCN / NIQE-class 5-axis natural-scene statistics (reverse-AI photoreal judge) [niqe13]; Applitools visual AI leads on learned diff Adoption: LIB-WIRED importers=9 nonval=5 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ◉ |
State invariants per tickMeasured:abug exists in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. camera-y bounds + pitch-clamp checked every tick, not just at frame samples Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ● | ● | ● | ○ |
Competitive-intel delta ledger (survey incumbents, drive roadmap)Measured:nx_cint_record exists in runtime/nx_competitive_intel.nx, verified at emit. Per-incumbent delta ledger with a 90-day survey cadence and next-due tracking (nx_cint_record/set_next_due); no test tool ships strategic competitive-intel as a primitive Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ○ | ○ | ○ | ○ |
Measured pitwall board (1:1 no-proxy)Measured:nx_pitwall_stage exists in runtime/nx_pitwall_render.nx, verified at emit. The racing-crew scoreboard: measured position vs named incumbents, no proxy (nxasm 4x/size/determinism triple-exceed found + gated 2026-05-29) Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program. | ● | ○ | ○ | ○ | ○ |
Liar-killed competitive census planeMeasured:LIAR-KILL exists in runtime/nx_swcompare_matrix.nx, verified at emit. 32 live head-to-head comparisons; every Nishi cell requires the symbol on disk, fabricated cells die at build -- the racing crew's find-gaps-vs-competitors engine, productized Adoption: LIVE — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
UX / UI composition gradingMeasured:pmil exists in runtime/nx_ui_judge.nx, verified at emit. nx_ui_judge scores contrast/tokens/mobile; the field pairs with axe/Lighthouse, Applitools does visual UI diff Adoption: PROMOTED-UNREGISTERED — PARTIAL: a real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ● | ● | ○ | ◐ |
Real-browser WebDriver / CDP driverOpen — watchingruntime/nx_webdriver.nx : wd_session, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright CDP [cdp] + Selenium WebDriver [w3c-webdriver] are the bar; our harness runs wasm directly, the JS/canvas/DOM layer is outside it by construction -- the named next rung | ○ | ◉ | ◉ | ◐ | ● |
Selector-based E2E scripting APIOpen — watchingruntime/nx_e2e_api.nx : e2e_locate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright locators + auto-wait are the bar [playwright-autowait]; we fuzz inputs, we do not script user journeys | ○ | ◐ | ◉ | ◉ | ● |
Visual-AI self-healing selectorsOpen — watchingruntime/nx_visual_ai.nx : va_heal, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Applitools self-healing locators + AI diff are the bar [applitools-eyes]; our vision is heuristic forensics, not a learned model | ○ | ● | ● | ○ | ◐ |
Cross-browser cloud test gridOpen — watchingruntime/nx_test_grid.nx : tg_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Selenium Grid + BrowserStack are the bar [selenium-grid]; we run one interpreted VM | ○ | ◉ | ● | ◐ | ◉ |
Network interception and mockingOpen — watchingruntime/nx_net_mock.nx : nm_route, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright route-mocking is the bar; absent here | ○ | ◐ | ◐ | ● | ○ |
Flaky-test detection and quarantineOpen — watchingruntime/nx_flaky.nx : fl_quarantine, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Every mature suite quarantines flakies [luo14]; ours is deterministic-by-seed so flakiness is structurally rarer, but no detector organ exists | ○ | ◐ | ◉ | ● | ○ |
Auto-record test codegenOpen — watchingruntime/nx_test_codegen.nx : tc_record, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Playwright codegen + Cypress Studio record user actions into tests [cypress-studio]; absent here (ties to the team census macro-recorder gap) | ○ | ◉ | ◐ | ● | ○ |
CI reporting and dashboardsOpen — watchingruntime/nx_test_report.nx : tr_report, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. JUnit/Allure reporting is the field standard; our forensics tables are console + the /compare board, no per-run CI surface | ○ | ◐ | ◐ | ◉ | ◉ |
Accessibility auto-auditOpen — watchingruntime/nx_a11y_audit.nx : aa_audit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. axe-core / Lighthouse a11y are the bar [axe-core]; ties to the UI-quality doctrine but no automated a11y organ | ○ | ● | ● | ○ | ○ |
LLM-driven test authoring and triageOpen — watchingruntime/nx_llm_test.nx : lt_author, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The cross-roster convergence gap (5th radar to point here); nobody in this field fully ships it either | ○ | ◐ | ◉ | ○ | ◐ |
Tests the shipped bytes on our own VM (sovereign, no external browser)Measured exceed:abug in runtime/_hdl_build/nx_mineworld_autotest.nx, verified at emit. The harness executes the ACTUAL deployed wasm in nx_wasm_vm and inspects VM memory + framebuffer host-side -- zero external browser, zero proxy, bit-exact; the field drives a real browser it does not own Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ◉ | ○ | ○ | ○ | ○ |
Survey-incumbents-and-drive-the-roadmap as a standing primitiveMeasured exceed:nx_cint_record in runtime/nx_competitive_intel.nx, verified at emit. The racing doctrine is a compiled ledger + cadence + pitwall + the 32-comparison plane: honest delta-to-incumbent drives every roadmap; the test-automation field ships no strategic competitive-intel layer Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ◉ | ○ | ○ | ○ | ○ |
Person · product · place — not yet measured for this domain
knowledge/compare/racing.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain racing, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [playwright-autowait] Microsoft. Playwright documentation, Auto-waiting: actionability checks (visible, stable, receives events, enabled, editable) run before every locator action, failing with TimeoutError instead of flaking. publisher · read in our library
knowledge/fetched/cmp_racing_playwright-autowait.html· pinhad67cb07cdb4baaf7b3a553bd4c6a64219114e937a78ad3444a2ec213a81f32a· accessed 2026-08-18 · vendor-docGrounds: The Playwright column: "Selector-based E2E scripting API" (the row's "Playwright locators + auto-wait are the bar") and "Real-browser WebDriver / CDP driver" (Playwright CDP) -- the documented behaviour our absent nx_e2e_api and nx_webdriver are graded against. - [cypress-studio] Cypress.io. Cypress Studio (Cypress Documentation): record clicks, typing and form interactions in the running app into Cypress test code, with AI-suggested assertions reviewed before saving. publisher · read in our library
knowledge/fetched/cmp_racing_cypress-studio.html· pinhc67cd97e9b99a894146f3355d6711dd8fbf55f6d8cbcff6195237f99bab05140· accessed 2026-08-18 · vendor-docGrounds: The Cypress column: "Auto-record test codegen" (the row's "Playwright codegen + Cypress Studio record user actions into tests") -- the recorder our nx_test_codegen is _ABSENT_ against, and the tie to the team census macro-recorder gap. - [selenium-grid] Selenium project. Selenium Grid documentation: WebDriver scripts executed on remote machines by routing client commands to remote browser instances -- parallel runs across many machines, browser versions and operating systems. publisher · read in our library
knowledge/fetched/cmp_racing_selenium-grid.html· pinh51f2d209d9994c6ac5fb133b847782ebdf5bf888d5632a539186991dec505254· accessed 2026-08-18 · vendor-docGrounds: The Selenium column: "Cross-browser cloud test grid" (the row's "Selenium Grid + BrowserStack are the bar") -- distributed multi-browser execution versus our single interpreted VM. - [w3c-webdriver] W3C. WebDriver. W3C Recommendation, 05 June 2018 -- "a remote control interface that enables introspection and control of user agents"; the wire protocol Selenium implements (WebDriver Level 2 continues as a Working Draft, 02 July 2026). publisher · accessed 2026-08-18 · published-standardGrounds: "Real-browser WebDriver / CDP driver" -- the row's "Selenium WebDriver" bar is this standard; a sovereign driver speaking it is the named next rung. Mirror absent by declaration: the sovereign fetch of w3.org returned a 5.8KB truncated body.
- [cdp] Chrome DevTools team (Google). Chrome DevTools Protocol: the domains, methods and events that let tools instrument, inspect, debug and profile Chromium-based browsers -- the protocol Playwright and Puppeteer drive. publisher · read in our library
knowledge/fetched/cmp_racing_cdp.html· pinhf4119c5a11315df0f83d4968bd77b0a5539408f527a2c1aa9d75ef481c40cd02· accessed 2026-08-18 · vendor-docGrounds: "Real-browser WebDriver / CDP driver" -- the CDP half of the row's "Playwright CDP + Selenium WebDriver are the bar"; the JS/canvas/DOM layer our wasm harness cannot reach is exactly what this protocol exposes. - [applitools-eyes] Applitools. Applitools Documentation: Eyes Visual AI added to an existing test framework, the Ultrafast Grid for cross-browser visual runs, and the Execution Cloud whose self-healing lets tests run even when elements have changed. publisher · read in our library
knowledge/fetched/cmp_racing_applitools.html· pinh226005936c94ab2229ea8aac907eb46988ad57adf9c99129da51171668407507· accessed 2026-08-18 · vendor-docGrounds: The Applitools column: "Visual-AI self-healing selectors" (the row's "Applitools self-healing locators + AI diff are the bar"), "Vision forensics on the SHIPPED artifact" (Applitools screenshot AI-diff on browser captures, code 3) and "UX / UI composition grading" -- learned visual diff versus our heuristic framebuffer forensics. - [niqe13] Mittal, Soundararajan, Bovik. Making a Completely Blind Image Quality Analyzer (NIQE). IEEE Signal Processing Letters, pp. 209-212, March 2013; with Mittal, Moorthy, Bovik, No-Reference Image Quality Assessment in the Spatial Domain (BRISQUE), IEEE TIP 2012 -- both listed on the LIVE lab no-reference IQA page. publisher · read in our library
knowledge/fetched/cmp_racing_niqe13.html· pinhff89949bea4a7961618db4a8fb32f153aeaac5566a3ee3f9243df4f42fc735b0· accessed 2026-08-18 · published-paperGrounds: "No-reference perceptual image-quality judge" -- the row's "MSCN / NIQE-class 5-axis natural-scene statistics": MSCN coefficients and the NIQE natural-scene-statistics model are these papers, the published basis of nx_natstat ns_mscn_rho. - [axe-core] Deque Systems. axe-core: the accessibility engine for automated Web UI testing (GitHub repository and rule documentation). publisher · read in our library
knowledge/fetched/cmp_racing_axe-core.html· pinh2ae56ef394658f8192bb22cc56f617bd0ec0189d7acef4596514e961be6db6af· accessed 2026-08-18 · vendor-docGrounds: "Accessibility auto-audit" -- the row's "axe-core / Lighthouse a11y are the bar"; the rule engine our absent nx_a11y_audit is graded against, tied to the UI-quality doctrine. - [luo14] Luo, Hariri, Eloussi, Marinov. An Empirical Analysis of Flaky Tests. FSE 2014 (22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering), pp. 643-653 -- 201 flaky-fix commits across 51 projects, root-cause taxonomy and fix strategies. publisher · read in our library
knowledge/fetched/cmp_racing_luo14.pdf· pinh36abbab3e917ad38300e8956a0311004423e4dcfede766b1bf6081b02fe1823b· accessed 2026-08-18 · published-paperGrounds: "Flaky-test detection and quarantine" -- the published root-cause taxonomy (async waits, concurrency, test-order dependence) behind why every mature suite quarantines flakies; our seed-deterministic harness makes the class structurally rarer, but no detector organ exists.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/racing.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers