Nishi FamilyCompare › Web Scraping and Extraction

Nishi Compare · measured, not asserted

Web Scraping and Extraction

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

Nishi vs Firecrawl and Apify and Bright Data and Zyte

Layer 1 · Executive

Where we are. Measured 2026-08-19. Sovereign TLS fetch, range fetch, a persisted frontier crawler, extraction, media harvest, near-duplicate detection, Common Crawl, WARC and CDX and a full-text index are present; the exceeds are sovereignty and zero-cost determinism. Absent: headless rendering, LLM extraction, managed scale, egress rotation and bot-wall handling -- the last two bounded by ethics (owned egress only, no CAPTCHA solving).

Where we need to go. Render JavaScript pages, extract structure with every field grounded, schedule on the swarm, and handle egress and bot walls within ethics -- never trading the sovereign, zero-SaaS stack for reach.

The unit. 1 u = one measured session-leg. Calibration from landed rungs: the mangagen panel compositor went from existing substrate to shipped and live-verified in ONE leg (2026-08-13); the citations rung went from 3 to 55 domains in one leg across seven seats (2026-08-18); a greenfield engine with a bite-proven gate has measured 2 to 4 legs. Estimates recalibrate as rungs land and PR7 actuals write back.
Cost to render: 2.5 u. Through M0.
Cost to extract and scale: 5.5 u. Through M1.
Cost to reach within ethics: 8 u. Everything below.
Cost to crawl smart (crawl4ai-class policies, sovereign): 14 u. Through M5: the six 2026-08-24 rungs lifted from the crawl4ai concept source.

18 of 23 capabilities measured|2 of them measured exceeds|5 open|coverage 782/1000|adoption 13 full / 5 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883283 domain=webscraping target_version=0.1 rungs=10 done=6 open=4 finish=2 ranker=nx_dr_ocm

#StageRungPriorityDerivation
FFINISHDomain map with soft-404 fingerprint (R8) dm_scanREGISTERED-UNAUTHORISEDregistered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked)
FFINISHPacing decay (R9) pace_decayPROMOTED-UNREGISTEREDa real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked)
#10.1Headless rendering (R0) hs_render360v=9 m=1 c=25
#2laterEgress rotation on owned hosts (R3) pp_rotate900v=9 m=1 c=10
#3laterScheduling at scale on the swarm (R2) ssch_plan800v=12 m=1 c=15
#4laterLLM structured extraction (R1) lx_extract400v=6 m=1 c=15

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Headless rendering (R0)hs_renderThe sovereign browser's JS execution renders SPA pages into the crawl; shared contract with the search domain's JS-rendered row; the browser census FETCH-EMPTY records flipOrgan2.5 u
LLM structured extraction (R1)lx_extractSchema-guided JSON extraction on the sovereign seat with every field grounded to a source span; a field without a span is refused, not invented; precision on a banked page set PRE-DECLAREDLocal model1.5 u
Scheduling at scale on the swarm (R2)ssch_planCrawl and extraction jobs planned and leased through the swarm job journal with robots and pace rules as data; a stalled job is re-leased, never lostOrgan1.5 u
Egress rotation on owned hosts (R3)pp_rotateRotation across sovereign-owned egress endpoints only (never residential botnets); the pace gate stays in the path and the policy is rowsOrgan1 u
Bot-wall classification (R4)abt_classifyThe research-fetch bot-wall witness generalised: Cloudflare challenge, cookie walls, Anubis and go-away recognised from captured specimens and reported as a verdict, never solved by a CAPTCHA farm; robots and terms are respected by constructionOrgan1.5 u
Crawl sufficiency (adaptive stop) (R5)cs_confidenceCoverage, consistency and saturation in integer permil over the pages fetched so far, weights and stop bars as conf rows, simhash-clustered duplicates discounted; the crawler and the research fetch stop when the bar is met and print which of the three moved; gate bite-proven on a fixture that must NOT stop early and one that must [@crawl4ai-adaptive]Organ1.5 u
Boilerplate block scoring (R6)bd_fit_textPer-block text and link density scored in permil with weights and threshold as conf rows; the crawler indexes the fit text; accept rule: on the banked fixture every article sentence survives and the nav and footer blocks are dropped, and a page with no boilerplate is byte-identical [@crawl4ai-fit]Organ1 u
Frontier priority and aging (R7)wc_priorityThe pull window is ranked shallow-first with a query-term bonus and rotated by a persisted cursor so every pending row gets a turn; accept rule: on a fixture frontier no row waits more than one full rotation and the ranking is a pure gate-tested function [@crawl4ai-readme]Organ1 u
Domain map with soft-404 fingerprint (R8)dm_scanSitemap index recursion with gzip, feeds, robots path mining, crt.sh subdomains and Wayback CDX composed into one appended-as-decided census whose partition sums; a soft-404 fingerprint taken from a known-bad path classifies probes; robots and pace stay in the path [@crawl4ai-readme]Organ2 u
Pacing decay (R9)pace_decayA throttled host's backoff halves on each success instead of resetting to base, so one 200 after a 429 cannot re-trigger the limit; the pace gate's teeth state the declared policyOrgan0.5 u

Milestones

MilestoneRungsCumulative
M0 · RenderR02.5 u
M1 · Extract and scaleR1,R25.5 u
M2 · Reach within ethicsR3,R48 u
M3 · Know when to stopR59.5 u
M4 · Extract and prioritiseR6,R711.5 u
M5 · Map and paceR8,R914 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiFirecrawlApifyBright DataZyte
Sovereign HTTP fetch (TLS)Measured: nx_https_fetch_follow exists in runtime/nx_https_fetch_follow.nx, verified at emit. All fetch over HTTP; Nishi owns its TLS Adoption: LIB-WIRED importers=232 nonval=211 — fully adopted (top of its ladder).
Byte-range fetchMeasured: nx_https_fetch_range exists in runtime/nx_https_fetch_follow.nx, verified at emit. Partial / resumable fetch Adoption: LIB-WIRED importers=232 nonval=211 — fully adopted (top of its ladder).
Frontier-persisted crawlerMeasured: nx_web_crawl_step exists in runtime/_hdl_build/nx_web_crawl_step.nx, verified at emit. Resumable BFS crawl frontier [rfc9309] Adoption: RUN-BY:fork:nx_hostctl — fully adopted (top of its ladder).
HTML content extractionMeasured: nx_html_to_text exists in runtime/nx_html_to_text.nx, verified at emit. Firecrawl leads (clean markdown extract) [kohlschuetter2010] Adoption: LIB-WIRED importers=56 nonval=41 — fully adopted (top of its ladder).
Media / asset harvestMeasured: nx_media_harvest exists in runtime/nx_media_harvest.nx, verified at emit. Image / video / stream URL extract Adoption: LIB-WIRED importers=6 nonval=5 — fully adopted (top of its ladder).
Near-duplicate detectionOpen — no implementing organ is measured for this axis yet. Content dedup; Bright Data mature [manku2007]
Common Crawl ingest (index-only)Measured: nx_cc_ingest exists in runtime/_hdl_build/nx_cc_ingest.nx, verified at emit. CDX + range-fetch, Memex-style [commoncrawl] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-UNAUTHORISED
WARC archive parseMeasured: nx_warc_index exists in runtime/nx_warc_index.nx, verified at emit. Web-archive format; live scrapers do not [warc11-iso28500] Adoption: LIB-GATE-ONLY importers=1 — PARTIAL: imported only by validation organs (gates, tests, benches): wire it into a shipping program.
adoption LIB-GATE-ONLY importers=1
CDX capture-index queryMeasured: nx_cdx_parse exists in runtime/nx_cdx_parse.nx, verified at emit. Wayback / Common Crawl capture index Adoption: LIB-WIRED importers=8 nonval=6 — fully adopted (top of its ladder).
Full-text index of scraped dataMeasured: nx_search_inverted exists in runtime/nx_search_inverted.nx, verified at emit. Nishi indexes; platforms mostly deliver raw Adoption: LIB-WIRED importers=38 nonval=32 — fully adopted (top of its ladder).
JS-rendered / headless scrapingOpen — watching runtime/nx_headless_scrape.nx : hs_render, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Firecrawl / Apify lead; Nishi static-HTML only
watching hs_render
Proxy rotation / residential IPsOpen — watching runtime/nx_proxy_pool.nx : pp_rotate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Bright Data leads (residential pool); Nishi none [brightdata-residential]
watching pp_rotate
Anti-bot / CAPTCHA handlingMeasured: abt_classify exists in runtime/nx_antibot.nx, verified at emit. Rivals bypass or solve; Nishi CLASSIFIES from specimen rows (tier, vendor, marker) and names the vendor, never solves [crawl4ai-readme] Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder).
LLM structured extractionOpen — watching runtime/nx_llm_extract.nx : lx_extract, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Firecrawl leads (schema / JSON extract); Nishi none [firecrawl-docs]
watching lx_extract
Managed scheduling / scaleOpen — watching runtime/nx_scrape_scheduler.nx : ssch_plan, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Apify actors + cloud scale; Nishi operator-stepped [apify-actors]
watching ssch_plan
Adaptive stop -- crawl sufficiency (coverage, consistency, saturation)Measured: cs_confidence exists in runtime/nx_crawl_sufficiency.nx, verified at emit. Stop when enough is known: integer coverage + on-topic consistency + new-term saturation, with simhash-discounted redundancy so copies cannot inflate coverage; no rival column's mirrored docs mention it (grep 2026-08-24) [crawl4ai-adaptive] Adoption: LIB-WIRED importers=3 nonval=2 — fully adopted (top of its ladder).
Boilerplate block scoring (text and link density)Measured: bd_fit_text exists in runtime/nx_block_density.nx, verified at emit. Kohlschuetter block features as permil scores with weights and threshold as conf rows, feeding the index a fit text instead of nav and footer [kohlschuetter2010] [crawl4ai-fit] Adoption: LIB-WIRED importers=5 nonval=4 — fully adopted (top of its ladder).
Frontier priority and aging (best-first pull, rotating window)Measured: wc_priority exists in runtime/_hdl_build/nx_web_crawl_step.nx, verified at emit. Shallow-first and query-term ranking of the pull window plus a persisted rotation cursor so no pending row starves behind dead-host rows (measured 93 pct of the window); rival codes = the depth and queue options in their crawl APIs, not re-read this pass [crawl4ai-readme] Adoption: RUN-BY:fork:nx_hostctl — fully adopted (top of its ladder).
Domain map (sitemap index recursion, feeds, robots path mining, CT-log subdomains, Wayback CDX)Measured: dm_scan exists in runtime/_hdl_build/nx_domain_map.nx, verified at emit. Every URL under a domain before crawling it, composed from the sitemap, feed, robots and CDX incumbents; Firecrawl and Apify document a map or sitemap step, coded Part pending a mirror read [crawl4ai-readme] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-UNAUTHORISED
Soft-404 fingerprintMeasured: dm_soft404 exists in runtime/_hdl_build/nx_domain_map.nx, verified at emit. A known-bad path's status, length and title fingerprint so a 200 that is really not-found never poisons discovery or host health [crawl4ai-readme] Adoption: REGISTERED-UNAUTHORISED — PARTIAL: registered, no cap ever minted; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-UNAUTHORISED
Adaptive per-host pacing (Retry-After, robots delay, persisted backoff that halves on success)Measured: pace_decay exists in runtime/_hdl_build/nx_crawl_pace.nx, verified at emit. Rate limiting persists across processes and honours the host's own numbers; rivals pace inside their managed cloud, opaque to the caller [rfc9309] Adoption: PROMOTED-UNREGISTERED — PARTIAL: a real binary nobody can call over MCP: /api/tools/register it; no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption PROMOTED-UNREGISTERED
Sovereign -- own language, compiler, TLS, index (zero SaaS dep)Measured exceed: ss_fold_cp in runtime/nx_seg_store.nx, verified at emit. All four are paid SaaS on big-cloud; Nishi owns the stack [zyte-api] Adoption: LIB-WIRED importers=441 nonval=321 — fully adopted (top of its ladder).
Runs on your hardware, zero cost, integer-deterministicMeasured exceed: idf_q10 in runtime/nx_intlog.nx, verified at emit. No per-request billing, no vendor lock, reproducible Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The coverage above is capability presence measured against source — not depth, scale, or polish, where mature rivals may lead. Exceeds are claimed only where a mechanism backs them. Every open gap is a watch contract: it names the organ and symbol that closes it, and this page flips the cell itself when that workstream ships.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/webscraping.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain webscraping, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [firecrawl-docs] Firecrawl. Introduction -- search the web, scrape any page, and interact with it through one API (scrape to clean markdown, crawl, extract with a schema). Vendor documentation. publisher · read in our library knowledge/fetched/cmp_webscraping_firecrawl.html · pin h5080b1c4b7fb8810091ce9bc9168bc551eec2c13e5aff79422f59d0e417f574f · accessed 2026-08-18 · vendor-docGrounds: The Firecrawl column: Best on "HTML content extraction" (clean markdown extract), "JS-rendered / headless scraping" and "LLM structured extraction" (schema / JSON extract) -- these are the product's documented core verbs, which is why the notes name Firecrawl as leader on those three rows.
  2. [apify-actors] Apify. Actors -- develop, run and share serverless cloud programs; build web scraping and automation tools and publish them on the Apify platform. Vendor documentation. publisher · read in our library knowledge/fetched/cmp_webscraping_apify.html · pin hd14dbeba5d90eb7996ee9e02a18c82e5b13765005e6b0450e7f4f59c1483ef7a · accessed 2026-08-18 · vendor-docGrounds: The Apify column: Best on "Managed scheduling / scale" (the note says Apify actors + cloud scale) and "JS-rendered / headless scraping" -- Actors are the documented unit of managed, scheduled, cloud-run scrapers the row grades against.
  3. [brightdata-residential] Bright Data. Introduction to Residential proxies -- the residential proxy network product documentation. publisher · read in our library knowledge/fetched/cmp_webscraping_brightdata.html · pin h4b7879eb665b048a082d933b78fa064374984e7c433bd7d62d418a346a5fb596 · accessed 2026-08-18 · vendor-docGrounds: The Bright Data column: Best on "Proxy rotation / residential IPs" (the note: Bright Data leads with a residential pool) and "Anti-bot / CAPTCHA handling" -- the residential-network documentation is the capability the row codes, and its absence in Nishi (_ABSENT_ nx_proxy_pool) is a declared gap, not an oversight.
  4. [zyte-api] Zyte. Zyte API: Get Started -- the managed fetch and extraction API (browser rendering and ban handling behind one endpoint). Vendor documentation. publisher · read in our library knowledge/fetched/cmp_webscraping_zyte.html · pin hbbb44a4c57d5c40e27a0f3a77f4ea34c1776f8ee4e86d32cd054b0c4db4b77df · accessed 2026-08-18 · vendor-docGrounds: The Zyte column codes (Yes on fetch, crawl, extraction, headless and anti-bot rows; Best on "Managed scheduling / scale"): Zyte API is the documented managed endpoint those codes describe, and like the other three it is paid SaaS -- the foil for the "Sovereign -- own language, compiler, TLS, index (zero SaaS dep)" exceed row.
  5. [warc11-iso28500] International Internet Preservation Consortium. The WARC Format 1.1 (ISO 28500:2017), the web-archive record format specification. publisher · read in our library knowledge/fetched/cmp_webscraping_warc11.html · pin h0c0bd8cc4b351eddd2b04aade67b2ddf9737961e087371998e01ab8deacde39d · accessed 2026-08-18 · published-standardGrounds: The "WARC archive parse" row (nx_warc_index): the record types, headers and framing nx_warc_index parses are this specification; the row's all-zero rival codes are honest because live scrapers do not read archives, and the row note says so.
  6. [commoncrawl] Common Crawl Foundation. Get Started -- the open web-crawl archive (WARC/WAT/WET) and its CDX capture index. Vendor page. publisher · read in our library knowledge/fetched/cmp_webscraping_commoncrawl.html · pin h7ff73e446d2367d1e2d8478a3b0e0e9205042a3012f319f20c1f3be0c8ef4570 · accessed 2026-08-18 · datasetGrounds: The "Common Crawl ingest (index-only)" and "CDX capture-index query" rows: nx_cc_ingest queries the CDX index and range-fetches WARC records from this archive (the Memex-style index-only ingest the note names) instead of crawling the majors' scale itself.
  7. [manku2007] Manku, G.S., Jain, A., Das Sarma, A. Detecting Near-Duplicates for Web Crawling. WWW 2007. publisher · read in our library knowledge/fetched/cmp_webscraping_manku2007.html · pin hbc8355bd8703aafa5f9bed15c80bb36dea262e1937727ed6e8a96caff12e9179 · accessed 2026-08-18 · published-paperGrounds: The "Near-duplicate detection" row (symbol simhash in nx_web_ingest): 64-bit simhash fingerprints compared by Hamming distance is this paper's method for crawl-time near-duplicate detection; the row's Part codes for the SaaS scrapers reflect that only Bright Data documents mature dedup.
  8. [rfc9309] Koster, M., Illyes, G., Zeller, H., Sassman, L. RFC 9309: Robots Exclusion Protocol. IETF, September 2022. publisher · read in our library knowledge/fetched/cmp_webscraping_rfc9309.html · pin hbbaa280636d3e38f0a1a130a5ce40e9753db1ed18c860945b0bf795ff8dc1a87 · accessed 2026-08-18 · published-standardGrounds: The "Frontier-persisted crawler" and "Sovereign HTTP fetch (TLS)" rows: the robots.txt rules a crawl frontier must honour are this standard, the published bar every column's Yes code (and nx_web_crawl_step) is held to.
  9. [kohlschuetter2010] Kohlschuetter, C., Fankhauser, P., Nejdl, W. Boilerplate Detection using Shallow Text Features. WSDM 2010, ACM. doi:10.1145/1718487.1718542. publisher · read in our library knowledge/fetched/cmp_webscraping_kohlschuetter2010.html · pin h2e81e88b012c07dbe8add5ac6c1b558d2bf8c3b87adc2ff7248bafa74060b42b · accessed 2026-08-18 · published-paperGrounds: The "Boilerplate block scoring" row (nx_block_density): block-level text density and link density as boilerplate features is this paper's result and the published method the row is graded against; the older "HTML content extraction" row is a tag-stripping extractor and cites it only as the bar it does not yet meet.
  10. [crawl4ai-readme] unclecode. Crawl4AI: Open-source LLM Friendly Web Crawler and Scraper -- README (79,320 stars, pushed 2026-08-24). Source repository documentation. publisher · read in our library knowledge/fetched/cmp_webscraping_crawl4ai_readme.md · pin h157214ef86639f896087ca220056ec190f36692a8e17affec9417f638d57e718 · accessed 2026-08-24 · source-readGrounds: The concept source for the "Frontier priority and aging", "Domain map" and "Soft-404 fingerprint" rows: the README documents best-first deep crawling with scorers and filters, DomainMapper's eight discovery sources with soft-404 detection, and a memory-adaptive dispatcher with a fairness timeout -- the policies these rows lift into sovereign organs; its own bytes are Python over Playwright, which is why it is a concept source and not a column.
  11. [crawl4ai-adaptive] unclecode. Crawl4AI documentation: Adaptive Web Crawling -- coverage, consistency and saturation as a three-layer sufficiency score with a confidence stop. publisher · read in our library knowledge/fetched/cmp_webscraping_crawl4ai_adaptive.md · pin h1013777b589438f0f4a73d8053ca895b62373a8b9ac7610d9674b92e6e3cdb90 · accessed 2026-08-24 · source-readGrounds: The "Adaptive stop -- crawl sufficiency" row: the three metrics and the confidence-threshold stop are this document's method; the sovereign row keeps the shape in integer permil and discounts near-duplicate pages, which the source's pairwise-Jaccard consistency rewards.
  12. [crawl4ai-fit] unclecode. Crawl4AI documentation: Fit Markdown with Pruning and BM25 -- per-node text density, link density and tag weights against a fixed or dynamic threshold. publisher · read in our library knowledge/fetched/cmp_webscraping_crawl4ai_fit.md · pin h96a5945694d8a86ac025d2661b748b2c5fc24a7d74d151e66aab8331eef1af09 · accessed 2026-08-24 · source-readGrounds: The "Boilerplate block scoring" row: the pruning filter's text-density and link-density features are the Kohlschuetter shallow features applied per node; the sovereign row keeps the two published features, moves the 0.48 threshold and the weights into conf rows, and grades against a banked fixture instead of a fixed number.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/webscraping.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers