nx_corpus_fetch.nx
buildroot/runtime/nx_corpus_fetch.nx
about
nx_corpus_fetch.nx -- Modelwright M1: sovereign corpus-at-scale builder.
Fetches public-text pages via the sovereign fetcher (nx_https_fetch_follow),
strips HTML to plain text (nx_html_to_text), concatenates into one training
corpus, and writes knowledge/corpus/corpus.txt. This is rung M1 of training
Modelwright FROM SCRATCH: the toy LM trained on 21KB; this scales the corpus.
100% sovereign (nx_cc->nxasm, own TLS-1.3 + own HTML stripper). The only
non-Nishi inputs are the CA-root DATA (/tmp/mozilla_certdata.txt) + the
fetched public text itself -- exactly the "what we have to" DATA category.
expect_exit: 0 license_tier: ORIGINAL
nx_safety_envelope:
intended_use: build a sovereign training corpus from public web text
sil_target: SIL1
verdict: NOT_YET_EVALUATED
genealogy_id: corpus_concat lineage_id: nx_corpus_fetch_v1
dependencies 5 imports · 0 importers
imports: nx_syscalls.nxnx_x509_trust_store.nxnx_trust_store_load_from_certdata.nxnx_https_fetch_follow.nxnx_html_to_text.nx
imported by: nobody (leaf or entry point)
call flow from main pre-order; caps 40 nodes / depth 6 declared; ↻ = already shown
structs
| none |
consts
| 24 | const K_MAGIC_8388608: i64 = 8388608 |
| 25 | const K_MAGIC_4194304: i64 = 4194304 |
| 26 | const K_MAGIC_33554432: i64 = 33554432 |
functions
| 28 | func w(s: *u8) -> i64 { var n: i64=0; while s[n]!=(0 as u8){n=n+1} sys_write(1,s,n); return 0 } |
| 29 | func wn(v: i64) -> i64 |
| 41 | func grab(url: *u8, store: *TrustStore, raw: *u8, txt: *u8, corpus: *u8, coff: *i64) -> i64 |
| 59 | func main() -> i64 |