code wiki / (root) / nx_crawl_doc_test.nx

nx_crawl_doc_test.nx

buildroot/runtime/nx_crawl_doc_test.nx

3859 B76 linesdepth 6pulls 11 transitivereach 0 importersview sourcekind gate/prooftopic crawl
docsdependenciesstructsconstsfunctions

about

nx_crawl_doc_test.nx -- KAT for the crawl back-half on real CONTENT. Native sovereign lane; exit 0 = pass, N = assertion N failed. Proves the dedup lesson resolved: SimHash on extracted PAGE CONTENT (not titles) correctly treats a scraped copy (same body, different boilerplate) as a near-duplicate while an unrelated page stays distinct, and the kept pages index + retrieve.

dependencies 3 imports · 0 importers

fx.nx nx_str.nx nx_crawl_doc.nx nx_crawl_doc_test.nx

imports: fx.nxnx_str.nxnx_crawl_doc.nx

imported by: nobody (leaf or entry point)

call flow from main pre-order; caps 40 nodes / depth 6 declared; ↻ = already shown

main nx_crawl_extract nx_html_to_text nx_html_to_text_x h2t_state_new sys_mmap sys_mmap ↻ h2t_scan_tag h2t_skip_to_gt h2t_scan_name h2t_is_name h2t_is_block_tag h2t_name_eq_lit h2t_lc h2t_lc ↻ h2t_is_suppress_tag h2t_name_eq_lit ↻ h2t_emit_newline h2t_name_eq_lit ↻ h2t_emit_raw h2t_find_href h2t_lc ↻ h2t_is_name ↻ h2t_decode_entity h2t_decode_numeric h2t_is_name ↻ nx_html_entity_lookup h2t_emit_codepoint h2t_emit_text_byte h2t_is_ws h2t_emit_text_byte ↻ sys_munmap nx_simhash_fingerprint nx_inv_is_token_char nx_inv_hash_bytes_lower nx_str_len nx_simhash_hamming nx_bits_popcount64_soft nx_crawl_should_index nx_simhash_min_hamming

structs

none

consts

none

functions

13func main() -> i64