Nishi Family › Compare › Web Scraping and Extraction
Nishi Compare · measured, not asserted
Web Scraping and Extraction
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs Firecrawl and Apify and Bright Data and Zyte
Where we are. Measured 2026-08-19. Sovereign TLS fetch, range fetch, a persisted frontier crawler, extraction, media harvest, near-duplicate detection, Common Crawl, WARC and CDX and a full-text index are present; the exceeds are sovereignty and zero-cost determinism. Absent: headless rendering, LLM extraction, managed scale, egress rotation and bot-wall handling -- the last two bounded by ethics (owned egress only, no CAPTCHA solving).
Where we need to go. Render JavaScript pages, extract structure with every field grounded, schedule on the swarm, and handle egress and bot walls within ethics -- never trading the sovereign, zero-SaaS stack for reach.
18 of 23 capabilities measured|2 of them measured exceeds|5 open|coverage 782/1000|adoption 13 full / 5 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883283 domain=webscraping target_version=0.1 rungs=10 done=6 open=4 finish=2 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| F | FINISH | Domain map with soft-404 fingerprint (R8) dm_scan | REGISTERED-UNAUTHORISED | registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked) |
| F | FINISH | Pacing decay (R9) pace_decay | PROMOTED-UNREGISTERED | a real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked) |
| #1 | 0.1 | Headless rendering (R0) hs_render | 360 | v=9 m=1 c=25 |
| #2 | later | Egress rotation on owned hosts (R3) pp_rotate | 900 | v=9 m=1 c=10 |
| #3 | later | Scheduling at scale on the swarm (R2) ssch_plan | 800 | v=12 m=1 c=15 |
| #4 | later | LLM structured extraction (R1) lx_extract | 400 | v=6 m=1 c=15 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Headless rendering (R0) | hs_render | The sovereign browser's JS execution renders SPA pages into the crawl; shared contract with the search domain's JS-rendered row; the browser census FETCH-EMPTY records flip | Organ | 2.5 u |
| LLM structured extraction (R1) | lx_extract | Schema-guided JSON extraction on the sovereign seat with every field grounded to a source span; a field without a span is refused, not invented; precision on a banked page set PRE-DECLARED | Local model | 1.5 u |
| Scheduling at scale on the swarm (R2) | ssch_plan | Crawl and extraction jobs planned and leased through the swarm job journal with robots and pace rules as data; a stalled job is re-leased, never lost | Organ | 1.5 u |
| Egress rotation on owned hosts (R3) | pp_rotate | Rotation across sovereign-owned egress endpoints only (never residential botnets); the pace gate stays in the path and the policy is rows | Organ | 1 u |
| Bot-wall classification (R4) | abt_classify | The research-fetch bot-wall witness generalised: Cloudflare challenge, cookie walls, Anubis and go-away recognised from captured specimens and reported as a verdict, never solved by a CAPTCHA farm; robots and terms are respected by construction | Organ | 1.5 u |
| Crawl sufficiency (adaptive stop) (R5) | cs_confidence | Coverage, consistency and saturation in integer permil over the pages fetched so far, weights and stop bars as conf rows, simhash-clustered duplicates discounted; the crawler and the research fetch stop when the bar is met and print which of the three moved; gate bite-proven on a fixture that must NOT stop early and one that must [@crawl4ai-adaptive] | Organ | 1.5 u |
| Boilerplate block scoring (R6) | bd_fit_text | Per-block text and link density scored in permil with weights and threshold as conf rows; the crawler indexes the fit text; accept rule: on the banked fixture every article sentence survives and the nav and footer blocks are dropped, and a page with no boilerplate is byte-identical [@crawl4ai-fit] | Organ | 1 u |
| Frontier priority and aging (R7) | wc_priority | The pull window is ranked shallow-first with a query-term bonus and rotated by a persisted cursor so every pending row gets a turn; accept rule: on a fixture frontier no row waits more than one full rotation and the ranking is a pure gate-tested function [@crawl4ai-readme] | Organ | 1 u |
| Domain map with soft-404 fingerprint (R8) | dm_scan | Sitemap index recursion with gzip, feeds, robots path mining, crt.sh subdomains and Wayback CDX composed into one appended-as-decided census whose partition sums; a soft-404 fingerprint taken from a known-bad path classifies probes; robots and pace stay in the path [@crawl4ai-readme] | Organ | 2 u |
| Pacing decay (R9) | pace_decay | A throttled host's backoff halves on each success instead of resetting to base, so one 200 after a 429 cannot re-trigger the limit; the pace gate's teeth state the declared policy | Organ | 0.5 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M0 · Render | R0 | 2.5 u |
| M1 · Extract and scale | R1,R2 | 5.5 u |
| M2 · Reach within ethics | R3,R4 | 8 u |
| M3 · Know when to stop | R5 | 9.5 u |
| M4 · Extract and prioritise | R6,R7 | 11.5 u |
| M5 · Map and pace | R8,R9 | 14 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | Firecrawl | Apify | Bright Data | Zyte |
|---|---|---|---|---|---|
Sovereign HTTP fetch (TLS)Measured:nx_https_fetch_follow exists in runtime/nx_https_fetch_follow.nx, verified at emit. All fetch over HTTP; Nishi owns its TLS Adoption: LIB-WIRED importers=232 nonval=211 — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
Byte-range fetchMeasured:nx_https_fetch_range exists in runtime/nx_https_fetch_follow.nx, verified at emit. Partial / resumable fetch Adoption: LIB-WIRED importers=232 nonval=211 — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
Frontier-persisted crawlerMeasured:nx_web_crawl_step exists in runtime/_hdl_build/nx_web_crawl_step.nx, verified at emit. Resumable BFS crawl frontier [rfc9309] Adoption: RUN-BY:fork:nx_hostctl — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
HTML content extractionMeasured:nx_html_to_text exists in runtime/nx_html_to_text.nx, verified at emit. Firecrawl leads (clean markdown extract) [kohlschuetter2010] Adoption: LIB-WIRED importers=56 nonval=41 — fully adopted (top of its ladder). | ● | ◉ | ● | ● | ● |
Media / asset harvestMeasured:nx_media_harvest exists in runtime/nx_media_harvest.nx, verified at emit. Image / video / stream URL extract Adoption: LIB-WIRED importers=6 nonval=5 — fully adopted (top of its ladder). | ● | ● | ● | ● | ● |
Near-duplicate detectionOpen — no implementing organ is measured for this axis yet. Content dedup; Bright Data mature [manku2007] | ○ | ◐ | ◐ | ● | ◐ |
Common Crawl ingest (index-only)Measured:nx_cc_ingest exists in runtime/_hdl_build/nx_cc_ingest.nx, verified at emit. CDX + range-fetch, Memex-style [commoncrawl] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ○ | ○ | ○ | ◐ |
WARC archive parseMeasured:nx_warc_index exists in runtime/nx_warc_index.nx, verified at emit. Web-archive format; live scrapers do not [warc11-iso28500] Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program. | ● | ○ | ○ | ○ | ○ |
CDX capture-index queryMeasured:nx_cdx_parse exists in runtime/nx_cdx_parse.nx, verified at emit. Wayback / Common Crawl capture index Adoption: LIB-WIRED importers=8 nonval=6 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Full-text index of scraped dataMeasured:nx_search_inverted exists in runtime/nx_search_inverted.nx, verified at emit. Nishi indexes; platforms mostly deliver raw Adoption: LIB-WIRED importers=38 nonval=32 — fully adopted (top of its ladder). | ● | ○ | ◐ | ◐ | ◐ |
JS-rendered / headless scrapingOpen — watchingruntime/nx_headless_scrape.nx : hs_render, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Firecrawl / Apify lead; Nishi static-HTML only | ○ | ◉ | ◉ | ● | ● |
Proxy rotation / residential IPsOpen — watchingruntime/nx_proxy_pool.nx : pp_rotate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Bright Data leads (residential pool); Nishi none [brightdata-residential] | ○ | ● | ◉ | ◉ | ● |
Anti-bot / CAPTCHA handlingMeasured:abt_classify exists in runtime/nx_antibot.nx, verified at emit. Rivals bypass or solve; Nishi CLASSIFIES from specimen rows (tier, vendor, marker) and names the vendor, never solves [crawl4ai-readme] Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder). | ● | ◉ | ● | ◉ | ● |
LLM structured extractionOpen — watchingruntime/nx_llm_extract.nx : lx_extract, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Firecrawl leads (schema / JSON extract); Nishi none [firecrawl-docs] | ○ | ◉ | ● | ● | ● |
Managed scheduling / scaleOpen — watchingruntime/nx_scrape_scheduler.nx : ssch_plan, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Apify actors + cloud scale; Nishi operator-stepped [apify-actors] | ○ | ● | ◉ | ◉ | ◉ |
Adaptive stop -- crawl sufficiency (coverage, consistency, saturation)Measured:cs_confidence exists in runtime/nx_crawl_sufficiency.nx, verified at emit. Stop when enough is known: integer coverage + on-topic consistency + new-term saturation, with simhash-discounted redundancy so copies cannot inflate coverage; no rival column's mirrored docs mention it (grep 2026-08-24) [crawl4ai-adaptive] Adoption: LIB-WIRED importers=3 nonval=2 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Boilerplate block scoring (text and link density)Measured:bd_fit_text exists in runtime/nx_block_density.nx, verified at emit. Kohlschuetter block features as permil scores with weights and threshold as conf rows, feeding the index a fit text instead of nav and footer [kohlschuetter2010] [crawl4ai-fit] Adoption: LIB-WIRED importers=5 nonval=4 — fully adopted (top of its ladder). | ● | ◉ | ● | ● | ● |
Frontier priority and aging (best-first pull, rotating window)Measured:wc_priority exists in runtime/_hdl_build/nx_web_crawl_step.nx, verified at emit. Shallow-first and query-term ranking of the pull window plus a persisted rotation cursor so no pending row starves behind dead-host rows (measured 93 pct of the window); rival codes = the depth and queue options in their crawl APIs, not re-read this pass [crawl4ai-readme] Adoption: RUN-BY:fork:nx_hostctl — fully adopted (top of its ladder). | ● | ◐ | ◐ | ◐ | ◐ |
Domain map (sitemap index recursion, feeds, robots path mining, CT-log subdomains, Wayback CDX)Measured:dm_scan exists in runtime/_hdl_build/nx_domain_map.nx, verified at emit. Every URL under a domain before crawling it, composed from the sitemap, feed, robots and CDX incumbents; Firecrawl and Apify document a map or sitemap step, coded Part pending a mirror read [crawl4ai-readme] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ◐ | ◐ | ○ | ○ |
Soft-404 fingerprintMeasured:dm_soft404 exists in runtime/_hdl_build/nx_domain_map.nx, verified at emit. A known-bad path's status, length and title fingerprint so a 200 that is really not-found never poisons discovery or host health [crawl4ai-readme] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ○ | ○ | ○ | ○ |
Adaptive per-host pacing (Retry-After, robots delay, persisted backoff that halves on success)Measured:pace_decay exists in runtime/_hdl_build/nx_crawl_pace.nx, verified at emit. Rate limiting persists across processes and honours the host's own numbers; rivals pace inside their managed cloud, opaque to the caller [rfc9309] Adoption: PROMOTED-UNREGISTERED — PARTIAL: a real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ◐ | ◐ | ◐ | ◐ |
Sovereign -- own language, compiler, TLS, index (zero SaaS dep)Measured exceed:ss_fold_cp in runtime/nx_seg_store.nx, verified at emit. All four are paid SaaS on big-cloud; Nishi owns the stack [zyte-api] Adoption: LIB-WIRED importers=441 nonval=321 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Runs on your hardware, zero cost, integer-deterministicMeasured exceed:idf_q10 in runtime/nx_intlog.nx, verified at emit. No per-request billing, no vendor lock, reproducible Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Person · product · place — not yet measured for this domain
knowledge/compare/webscraping.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain webscraping, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [firecrawl-docs] Firecrawl. Introduction -- search the web, scrape any page, and interact with it through one API (scrape to clean markdown, crawl, extract with a schema). Vendor documentation. publisher · read in our library
knowledge/fetched/cmp_webscraping_firecrawl.html· pinh5080b1c4b7fb8810091ce9bc9168bc551eec2c13e5aff79422f59d0e417f574f· accessed 2026-08-18 · vendor-docGrounds: The Firecrawl column: Best on "HTML content extraction" (clean markdown extract), "JS-rendered / headless scraping" and "LLM structured extraction" (schema / JSON extract) -- these are the product's documented core verbs, which is why the notes name Firecrawl as leader on those three rows. - [apify-actors] Apify. Actors -- develop, run and share serverless cloud programs; build web scraping and automation tools and publish them on the Apify platform. Vendor documentation. publisher · read in our library
knowledge/fetched/cmp_webscraping_apify.html· pinhd14dbeba5d90eb7996ee9e02a18c82e5b13765005e6b0450e7f4f59c1483ef7a· accessed 2026-08-18 · vendor-docGrounds: The Apify column: Best on "Managed scheduling / scale" (the note says Apify actors + cloud scale) and "JS-rendered / headless scraping" -- Actors are the documented unit of managed, scheduled, cloud-run scrapers the row grades against. - [brightdata-residential] Bright Data. Introduction to Residential proxies -- the residential proxy network product documentation. publisher · read in our library
knowledge/fetched/cmp_webscraping_brightdata.html· pinh4b7879eb665b048a082d933b78fa064374984e7c433bd7d62d418a346a5fb596· accessed 2026-08-18 · vendor-docGrounds: The Bright Data column: Best on "Proxy rotation / residential IPs" (the note: Bright Data leads with a residential pool) and "Anti-bot / CAPTCHA handling" -- the residential-network documentation is the capability the row codes, and its absence in Nishi (_ABSENT_ nx_proxy_pool) is a declared gap, not an oversight. - [zyte-api] Zyte. Zyte API: Get Started -- the managed fetch and extraction API (browser rendering and ban handling behind one endpoint). Vendor documentation. publisher · read in our library
knowledge/fetched/cmp_webscraping_zyte.html· pinhbbb44a4c57d5c40e27a0f3a77f4ea34c1776f8ee4e86d32cd054b0c4db4b77df· accessed 2026-08-18 · vendor-docGrounds: The Zyte column codes (Yes on fetch, crawl, extraction, headless and anti-bot rows; Best on "Managed scheduling / scale"): Zyte API is the documented managed endpoint those codes describe, and like the other three it is paid SaaS -- the foil for the "Sovereign -- own language, compiler, TLS, index (zero SaaS dep)" exceed row. - [warc11-iso28500] International Internet Preservation Consortium. The WARC Format 1.1 (ISO 28500:2017), the web-archive record format specification. publisher · read in our library
knowledge/fetched/cmp_webscraping_warc11.html· pinh0c0bd8cc4b351eddd2b04aade67b2ddf9737961e087371998e01ab8deacde39d· accessed 2026-08-18 · published-standardGrounds: The "WARC archive parse" row (nx_warc_index): the record types, headers and framing nx_warc_index parses are this specification; the row's all-zero rival codes are honest because live scrapers do not read archives, and the row note says so. - [commoncrawl] Common Crawl Foundation. Get Started -- the open web-crawl archive (WARC/WAT/WET) and its CDX capture index. Vendor page. publisher · read in our library
knowledge/fetched/cmp_webscraping_commoncrawl.html· pinh7ff73e446d2367d1e2d8478a3b0e0e9205042a3012f319f20c1f3be0c8ef4570· accessed 2026-08-18 · datasetGrounds: The "Common Crawl ingest (index-only)" and "CDX capture-index query" rows: nx_cc_ingest queries the CDX index and range-fetches WARC records from this archive (the Memex-style index-only ingest the note names) instead of crawling the majors' scale itself. - [manku2007] Manku, G.S., Jain, A., Das Sarma, A. Detecting Near-Duplicates for Web Crawling. WWW 2007. publisher · read in our library
knowledge/fetched/cmp_webscraping_manku2007.html· pinhbc8355bd8703aafa5f9bed15c80bb36dea262e1937727ed6e8a96caff12e9179· accessed 2026-08-18 · published-paperGrounds: The "Near-duplicate detection" row (symbol simhash in nx_web_ingest): 64-bit simhash fingerprints compared by Hamming distance is this paper's method for crawl-time near-duplicate detection; the row's Part codes for the SaaS scrapers reflect that only Bright Data documents mature dedup. - [rfc9309] Koster, M., Illyes, G., Zeller, H., Sassman, L. RFC 9309: Robots Exclusion Protocol. IETF, September 2022. publisher · read in our library
knowledge/fetched/cmp_webscraping_rfc9309.html· pinhbbaa280636d3e38f0a1a130a5ce40e9753db1ed18c860945b0bf795ff8dc1a87· accessed 2026-08-18 · published-standardGrounds: The "Frontier-persisted crawler" and "Sovereign HTTP fetch (TLS)" rows: the robots.txt rules a crawl frontier must honour are this standard, the published bar every column's Yes code (and nx_web_crawl_step) is held to. - [kohlschuetter2010] Kohlschuetter, C., Fankhauser, P., Nejdl, W. Boilerplate Detection using Shallow Text Features. WSDM 2010, ACM. doi:10.1145/1718487.1718542. publisher · read in our library
knowledge/fetched/cmp_webscraping_kohlschuetter2010.html· pinh2e81e88b012c07dbe8add5ac6c1b558d2bf8c3b87adc2ff7248bafa74060b42b· accessed 2026-08-18 · published-paperGrounds: The "Boilerplate block scoring" row (nx_block_density): block-level text density and link density as boilerplate features is this paper's result and the published method the row is graded against; the older "HTML content extraction" row is a tag-stripping extractor and cites it only as the bar it does not yet meet. - [crawl4ai-readme] unclecode. Crawl4AI: Open-source LLM Friendly Web Crawler and Scraper -- README (79,320 stars, pushed 2026-08-24). Source repository documentation. publisher · read in our library
knowledge/fetched/cmp_webscraping_crawl4ai_readme.md· pinh157214ef86639f896087ca220056ec190f36692a8e17affec9417f638d57e718· accessed 2026-08-24 · source-readGrounds: The concept source for the "Frontier priority and aging", "Domain map" and "Soft-404 fingerprint" rows: the README documents best-first deep crawling with scorers and filters, DomainMapper's eight discovery sources with soft-404 detection, and a memory-adaptive dispatcher with a fairness timeout -- the policies these rows lift into sovereign organs; its own bytes are Python over Playwright, which is why it is a concept source and not a column. - [crawl4ai-adaptive] unclecode. Crawl4AI documentation: Adaptive Web Crawling -- coverage, consistency and saturation as a three-layer sufficiency score with a confidence stop. publisher · read in our library
knowledge/fetched/cmp_webscraping_crawl4ai_adaptive.md· pinh1013777b589438f0f4a73d8053ca895b62373a8b9ac7610d9674b92e6e3cdb90· accessed 2026-08-24 · source-readGrounds: The "Adaptive stop -- crawl sufficiency" row: the three metrics and the confidence-threshold stop are this document's method; the sovereign row keeps the shape in integer permil and discounts near-duplicate pages, which the source's pairwise-Jaccard consistency rewards. - [crawl4ai-fit] unclecode. Crawl4AI documentation: Fit Markdown with Pruning and BM25 -- per-node text density, link density and tag weights against a fixed or dynamic threshold. publisher · read in our library
knowledge/fetched/cmp_webscraping_crawl4ai_fit.md· pinh96a5945694d8a86ac025d2661b748b2c5fc24a7d74d151e66aab8331eef1af09· accessed 2026-08-24 · source-readGrounds: The "Boilerplate block scoring" row: the pruning filter's text-density and link-density features are the Kohlschuetter shallow features applied per node; the sovereign row keeps the two published features, moves the 0.48 threshold and the weights into conf rows, and grades against a banked fixture instead of a fixed number.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/webscraping.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers