nx_llm_forward_profile.nx
buildroot/runtime/nx_llm_forward_profile.nx
about
nx_llm_forward_profile.nx -- clean per-decode-forward timing on the
REAL model, to find where a decode token's time actually goes.
The 34-token bench is too noisy on this host (WSL bounces + multi-
session load) to attribute sub-20% changes. This times ONE
forward_v4 decode call at a time (no sampler/detok/penalty/loop
overhead), after a warmup prefill so the dequant cache + lm_head
pool are hot. Combined with the isolated lm_head number
(nx_matmul_t_pool_gate: ~143ms pooled at m=1,n=151936), it splits
the decode forward into lm_head vs the 24-layer block stack.
Prints each step's us (watch for growth = KV-cache-length effect)
and the average. No vocab/bpe load (forward needs neither).
expect_exit: 0
dependencies 16 imports · 0 importers
diagram shows first 10 each side; +6 more imports, +0 more importers in the complete lists below.
imports: nx_syscalls.nxnx_itoa_lib.nxnx_tier.nxnx_gguf.nxnx_gguf_load.nxnx_gguf_meta.nxnx_f32.nxnx_f32_kv_cache.nxnx_f32_lazy_weight.nxnx_f32_llama_block.nxnx_f32_llama_block_v4.nxnx_f32_llama_stack_v4.nxnx_f32_llama_layer_lazy_load.nxnx_f32_llm.nxnx_f32_llm_v4.nxnx_f32_llm_read_dims.nx
imported by: nobody (leaf or entry point)
call flow from main pre-order; caps 40 nodes / depth 6 declared; ↻ = already shown
structs
| none |
consts
| none |
functions
| 34 | func pf_puts(s: *u8) -> i64 { var n: i64=0; while s[n]!=(0 as u8){n=n+1} sys_write(1,s,n); return 0 } |
| 39 | func pf_num(v: i64) -> i64 { nxi_out(v); return 0 } |
| 40 | func pf_nl() -> i64 { pf_puts("\n" as *u8); return 0 } |
| 42 | func main(argc: i64, argv: *i64) -> i64 |