← product evidence + the ladder to state of the art
Nishi browser vs Chrome — the TOP-20 real-site suite (first honest baseline)
workstream: S21 experiential / S18 graphics — browser chrome-parity ratchetrun: browser-top20-2026-07-162026-07-16
Verdict: The benchmark is now 20 real sites we do not control (operator: one self-authored page proves nothing), re-runnable as one command (bench/nx_top20_parity.ps1) — and its FIRST HONEST LESSON is about the metric itself: the eyeball pass caught Nishi rendering google.com as a BLANK page that still scored 97.3% coarse, because google is mostly white space and a blank white page matches white space. Region-mean parity is GAMEABLE BY EMPTINESS on sparse pages — so google/reddit/ebay/example's high scores are VACUOUS until an ink-weight check lands (the next instrument rung: compare content-pixel mass per side; a side with a fraction of the oracle's ink cannot score parity). What stands after that discipline: wikipedia 90.6% coarse and hackernews 88.5% are REAL (dense text pages, visually confirmed); stackoverflow 81.7 / netflix 80.2 / linkedin 79.6 render real content with real gaps. The suite still earned its keep twice over: run 1 caught the canvas background-propagation bug (a banner div painted google's whole canvas dark; fixed to spec root-chain html->body, gate + /experiential void re-verified), and run 2's eyeball caught the blank-parity metric hole. Named gaps, each a ladder rung: google renders blank despite fetching (layout has 3378px of content — paint/visibility diagnosis queued); 5 sites FETCH-EMPTY (youtube/facebook/instagram/duckduckgo/craigslist); bing + github 0.0% (background images + attribute-selector themes); x/reddit tiny JS shells; amazon blocks headless Chrome (no oracle); ebay's oracle was a bot page.













Measured
| what | result | meaning |
|---|---|---|
| Suite verdict (renderable 13 of 20, ink-disciplined) | REAL parity: wikipedia 90.6 · hackernews 88.5 · stackoverflow 81.7 · netflix 80.2 · linkedin 79.6 · bbc 71.3 | VACUOUS (blank/sparse-matching, pending ink-weight): google 97.3 · ebay 99.4 · example 98.9 · reddit 92.1 — high numbers on near-empty renders count for NOTHING |
| Metric hole found by the eyeball pass | blank Nishi google scored 97.3% coarse | region-mean parity is gameable by emptiness on sparse pages -> NEXT instrument rung: ink-weight (content-pixel mass per side; ink ratio << oracle => BLANK-NISHI class, no parity awarded) |
| google renders BLANK despite fetching | nishi_h=3378 laid out; zero visible paint | paint/visibility diagnosis = the top engine rung from this run (text color? visibility rules? fold?) |
| The suite caught a real engine bug on run 1 | google/reddit coarse 0.0% -> 97.3%/92.1% | canvas bg propagated from ANY first non-white block (a banner div); spec fix = root chain only (html->body); /experiential void re-verified dark (luma 17) after the fix |
| FETCH-EMPTY (5) | youtube, facebook, instagram, duckduckgo, craigslist | nishi_fetch returns nothing — the fetch-reliability rung (TLS fingerprinting / redirect chains / HTTP-2-only to diagnose) |
| Oracle failures (2) | amazon (Chrome headless blocked), ebay (Chrome got a 354px bot page while Nishi got real content) | the CHROME-BLOCKED class exists so two blank pages can never again score 100% parity (run 1 did exactly that) |
| Named paint gaps from the suite | background-image (bing full-bleed photo), attribute-selector themes (github data-color-mode), JS shells (x, reddit skeletons) | each a ladder rung with its measuring site |
| Height ratios (renderable) | wikipedia 115 · bbc 112 · netflix 120 · hackernews 134 · stackoverflow 88 · linkedin 77 · github 51 · google 42 | over-100s = margin-collapse rung; under-50s = JS-populated content we honestly don't have |
| THE ORGAN EYE (nx_parity_judge) — the suite now sees these failures WITHOUT a human | gate GREEN 6/6 against the human truth set from this very run | autonomous classes per pair: BLANK-NISHI (google, netflix, ebay) · NO-ORACLE-CONTENT (bing) · CANVAS-OFF (github) · SPARSE-NISHI (linkedin, x, stackoverflow) · CONTENT-OK (wikipedia, hackernews, yahoo, example, reddit) |
| The organ out-judged the human once | hackernews "dark canvas" human read -> canvas-delta 23 (fine) | the dark region is the post-content area, not the page canvas — the eyeball was wrong, the measurement was right; caption corrected |
| Common failures, machine-countable now | 1 blank-paint class (google) · 1-char column collapse (bing/netflix strips) · canvas themes (github) · decoration/column loss (linkedin/x/so) · table scatter (HN) · attr-leak (SO) · overlap (yahoo/wiki) | judge v2 rungs: attr-leak + overlap need a layout-tree dump from the shot lane (pixel-only can't see them) |
| READING RECALL — the number that ends "it looks fine" (OCR tier: the same neutral instrument reads BOTH renders; sovereign scorer = truth-token recall, permille) | suite mean ~384 of 1000: the Nishi browser preserves ~38 percent of what a Chrome reader can read | per site: x 848 · example 789 (even the trivial control loses 21 percent to font quality) · linkedin 772 · wikipedia 766 (our best real page: "bad but legible") · bbc 543 · hackernews 400 (the table scatter halves readability) · stackoverflow 116 (the CONTENT-OK class hid an 88 percent reading loss) · github 108 · yahoo 84 · google/reddit/netflix/ebay 0 (blank) — instrument bench/nx_ocr_read.ps1 (Windows.Media.Ocr, typography-arc pattern) + organ _offc/nx_ocr_recall.elf (identity self-test 1000) |
| Precision tells its own story | wikipedia precision 507 vs recall 766 | we paint MORE tokens than Chrome shows (duplicated/garbled runs + menu text Chrome hides) — badly-placed text is as measurable as missing text |
| Repeatability | ONE command: bench/nx_top20_parity.ps1 — self-judging + read-scored | Chrome oracle + Nishi record + sovereign instruments + nx_parity_judge class + OCR reading-recall per site; results committed with the run |
Machine detail: api.json · related: /evidence/nishi-browser