nx_p256_fieldmul_mulx_bench.nx
buildroot/runtime/nx_p256_fieldmul_mulx_bench.nx
about
nx_p256_fieldmul_mulx_bench.nx -- field-mul-level speed: production p256_field_mul (u256_mul_wide
8x32 + Solinas) vs p256_field_mul_mulx_s (fused __mul256_wide + the SAME Solinas, reused scratch).
The Solinas reduction is IDENTICAL in both, so the delta isolates the multiply. Dependent-chain
harness (DCE/hoist-proof). HONEST caveat printed: the mulx variant reuses scratch (the optimized
form); the production field_mul mmaps its product buffer per call -- so this reflects a real
drop-in optimized field_mul, and the ratio shows how the 53x primitive win survives Amdahl once
the common (unchanged) reduce cost is included. expect_exit: 0 license_tier: ORIGINAL
dependencies 1 imports · 0 importers
imports: nx_p256_fieldmul_mulx.nx
imported by: nobody (leaf or entry point)
call flow from main pre-order; caps 40 nodes / depth 6 declared; ↻ = already shown
structs
| none |
consts
| 9 | const N_MAGIC_1000000000: i64 = 1000000000 |
| 11 | const N_ITER: i64 = 3000000 |
functions
| 13 | func bp(s: *u8) -> i64 { var n: i64 = 0; while s[n] != (0 as u8) { n = n + 1 } sys_write(1, s, n); return 0 } called by 1: main |
| 14 | func bn(v: i64) -> i64 called by 1: main |
| 20 | func now_ns(ts: *i64) -> i64 { __syscall(SYS_CLOCK_GETTIME, 1, ts as i64, 0, 0, 0, 0); return ts[0] * N_MAGIC_1000000000 + ts[1] } called by 1: main |
| 22 | func main() -> i64 |