Nishi FamilyCompare › Supervision and Service Monitoring

Nishi Compare · measured, not asserted

Supervision and Service Monitoring

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

Nishi vs systemd and supervisord and Prometheus and Netdata

Layer 1 · Executive

Where we are. Measured 2026-08-19. Supervision and never-brick promotion are one live circuit with guards, health-gated deploy and rollback, liveness audit and capability-pinned remote ops. Observability depth is the gap the matrix names, and since it was scored (2026-07-10) the estate has shipped pieces the rungs compose rather than invent: nx_netobs (series, probes, a page), nx_resmon (per-process RSS and swap with a trend reader over resmon.log), nx_resgov (memory caps enforced every minute), nx_slo (an importable SLO library), nx_notice and the vizsla notifier. None is yet wired into the supervision surface, which is why the rows stay absent.

Where we need to go. See the fleet on one served surface, act on it (alerts, SLOs, logs), then add the init-system depth systemd is judged on and the anomaly and federation axes Netdata and Prometheus own -- composing the observability organs that already exist.

The unit. 1 u = one measured session-leg. Calibration from landed rungs: the mangagen panel compositor went from existing substrate to shipped and live-verified in ONE leg (2026-08-13); the citations rung went from 3 to 55 domains in one leg across seven seats (2026-08-18); a greenfield engine with a bite-proven gate has measured 2 to 4 legs. Estimates recalibrate as rungs land and PR7 actuals write back.
Cost to see the fleet: 4 u. Through M0: a UI, per-service metrics and history, all composition of existing organs.
Cost to act on it: 7 u. Through M1.
Cost to the full surface: 15.5 u. Everything below.

8 of 23 capabilities measured|2 of them measured exceeds|15 open|coverage 347/1000|adoption 6 full / 2 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883477 domain=supervisor target_version=0.1 rungs=11 done=0 open=11 finish=0 ranker=nx_dr_ocm

#StageRungPriorityDerivation
#10.1Per-service resource metrics (R1) rmx_sample1400v=14 m=1 c=10
#20.1Health UI (R0) hu_page933v=14 m=1 c=15
#30.1Metrics history (R2) ts_append400v=6 m=1 c=15
#4laterAlert routing (R3) nl_route_alert1500v=15 m=1 c=10
#5laterSLO report (R4) hc_slo_report1500v=3 m=1 c=2
#6laterDependency-ordered startup (R6) hc_dep_order1142v=8 m=1 c=7
#7laterLog aggregation (R5) la_tail_index933v=14 m=1 c=15
#8laterResource-limit enforcement named (R7) rg_enforce700v=7 m=1 c=10
#9laterAnomaly detection (R9) ao_detect350v=7 m=1 c=20
#10laterFederated monitoring (R10) fm_federate350v=7 m=1 c=20
#11laterSocket and timer activation (R8) sa_listen100v=2 m=1 c=20

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Health UI (R0)hu_pageA served monitoring page composing what already exists -- mgmt_snap.json, nx_netobs no_page and no_svg_series, resmon.log -- with what-to-look-at-first ordering; graded by the UI judge against the Netdata bar; closes the dashboard row and U1 to U4 which share this symbolOrgan1.5 u
Per-service resource metrics (R1)rmx_sampleCPU, memory and IO per supervised service sampled on the beat (nx_resmon already meters RSS and swap per process; CPU and IO join it) and written as rows; the partition of services sumsOrgan1 u
Metrics history (R2)
after R1
ts_appendAn append-only time-series store over the per-service samples with a bounded reader; nx_sizeguard budgets it; resmon.log and netobs series are the first two writersOrgan1.5 u
Alert routing (R3)nl_route_alertA health RED or budget breach routes through the notice ledger to the vizsla reminder engine and email digest already proven for deadlines; a planted RED reaches the digest and a GREEN does notOrgan1 u
SLO report (R4)
after R2
hc_slo_reportnx_slo (availability permille, error budget, latency percentile) computed per service from the history and printed by hostctl status; a breach is a verdict, not a log lineOrgan0.5 u
Log aggregation (R5)la_tail_indexThe per-daemon /tmp tails indexed into one searchable store on the beat; a query returns the line and its daemon and epochOrgan1.5 u
Dependency-ordered startup (R6)hc_dep_orderServices declare dependencies as data and hostctl starts them in topological order with a loud cycle refusal; a planted cycle is refusedOrgan1.5 u
Resource-limit enforcement named (R7)
after R1
rg_enforceThe resgov memory caps already live are extracted into a named enforcement function and extended to CPU and IO limits as conf rows; a process over its declared limit is governed and the action is loggedOrgan1 u
Socket and timer activation (R8)
after R6
sa_listenA daemon may declare socket activation so hostctl holds the listener and spawns on first connection; measured idle cost drops to zero for the declared servicesOrgan2 u
Anomaly detection (R9)
after R2
ao_detectPer-series anomaly flags from the history (rate-of-change and seasonal baselines) that ABSTAIN with too little history; a planted spike is flagged and a flat series is notOrgan2 u
Federated monitoring (R10)
after R2
fm_federateLaptop, NAS and workers report into one history with per-host partition; the UI shows the fleet, not one boxOrgan2 u

Milestones

MilestoneRungsCumulative
M0 · See the fleetR0,R1,R24 u
M1 · Act on itR3,R4,R57 u
M2 · Init-system depthR6,R7,R811.5 u
M3 · BeyondR9,R1015.5 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishisystemdsupervisordPrometheusNetdata
Process supervision with crash-restart guardsMeasured: hc_guard_one exists in runtime/_hdl_build/nx_hostctl.nx, verified at emit. LIVE on the NAS: supervise + per-service guards; a guard restart of nx_relate_daemon was captured in this session's nx_status output; systemd is the init-system bar [systemd-service] Adoption: LIVE-DAEMON — fully adopted (top of its ladder).
Never-brick deploy watchdog (health-gated promote + auto-rollback)Measured: func ma_do_deploy exists in runtime/_hdl_build/nx_mgmt_api.nx, verified at emit. DEPLOYED-GREEN proven across mgmt/edge/office arcs; the field bars supervise OR observe -- none promote with rollback (that is Argo-class territory, see /compare/deploy) Adoption: LIVE-DAEMON — fully adopted (top of its ladder).
Health snapshot + service inventory APIMeasured: ss_cat exists in runtime/_hdl_build/nx_mgmt_snapshot.nx, verified at emit. mgmt_snap.json written each poll; GET /api/services returned the live 14-service inventory over MCP in this session [supervisord] Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder).
Deployed-capability liveness audit (BUILT is not LIVE)Measured: audit_row exists in runtime/_hdl_build/nx_reader_liveness.nx, verified at emit. The RACI monitor activity exists because a built reader shipped dead; this organ proves served-and-wired, not just compiled; Prometheus blackbox probing is the partial peer Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it.
adoption SOURCE-ONLY
Accountability-routed monitoring (RACI lineage)Measured: lr_resolve exists in runtime/_hdl_build/nx_lineage_raci.nx, verified at emit. Monitoring findings resolve to the ONE accountable role; the field routes to dashboards, not owners Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it.
adoption BUILT-UNPROMOTED
Live-derived maturity grading (anti-staleness by design)Measured: em_domain_level exists in runtime/_hdl_build/nx_ecomat_lib.nx, verified at emit. Grades re-derive from evidence files on every read -- a snapshot grade is impossible by construction; exposed live as an MCP tool Adoption: LIB-WIRED importers=21 nonval=13 — fully adopted (top of its ladder).
Web health dashboardOpen — watching runtime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. A /health fleet UI exists in the containers arc but no serving symbol verified in THIS fleet's organs -- counted absent honestly; Netdata auto-dashboards are the bar; Prometheus needs Grafana
watching hu_page
Time-series metrics databaseOpen — watching runtime/nx_tsdb.nx : ts_append, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus TSDB is the bar [prometheus-overview]; our snapshots are point-in-time, no history
watching ts_append
Alerting rules + notification routingOpen — watching runtime/_hdl_build/nx_notice.nx : nl_route_alert, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Alertmanager is the bar [alertmanager]; our vizsla reminder engine fires medical deadlines but is NOT wired to service alerts -- the named next rung
watching nl_route_alert
Log aggregation + searchOpen — watching runtime/nx_log_aggregate.nx : la_tail_index, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. journald is the bar; our logs are per-daemon /tmp tails read via status
watching la_tail_index
Per-service resource metrics (CPU mem IO)Open — watching runtime/nx_resmetrics.nx : rmx_sample, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata per-second resource visibility is the bar [netdata-docs]; systemd cgroup accounting partial
watching rmx_sample
Anomaly detection / AIOpsOpen — watching runtime/nx_aiops.nx : ao_detect, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata ships ML anomaly bits [netdata-ml]; nothing sovereign here yet
watching ao_detect
SLO / error-budget trackingOpen — watching runtime/_hdl_build/nx_hostctl.nx : hc_slo_report, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus recording rules approximate it; we track gate GREEN, not error budgets [sre-slo]
watching hc_slo_report
Socket / timer activationOpen — watching runtime/nx_socket_activate.nx : sa_listen, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd socket activation is the bar [systemd-socket]; our daemons are always-on
watching sa_listen
Dependency-ordered startup graphOpen — watching runtime/_hdl_build/nx_hostctl.nx : hc_dep_order, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd unit graphs are the bar; hostctl starts a flat ordered list
watching hc_dep_order
Resource-limit enforcementOpen — watching runtime/nx_resgov.nx : rg_enforce, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd cgroup limits are the bar; rule 21 is awareness, not enforcement
watching rg_enforce
Multi-host federated monitoringOpen — watching runtime/nx_fed_monitor.nx : fm_federate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus federation is the bar [prom-federation]; laptop+NAS+workers exist but each is watched alone
watching fm_federate
U1 visual design of the monitoring surfaceOpen — watching runtime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. U-axes vs the NAMED product Netdata v2.10.3 dashboard (banked 2026-07-10); we ship no monitoring UI to grade
watching hu_page
U2 interaction quality (drill-down, filtering, latency)Open — watching runtime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata per-second interactive charts are the bar; nothing of ours to grade
watching hu_page
U3 information design (what to look at first)Open — watching runtime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata auto-organizes by host and service; our status output is a raw console dump
watching hu_page
U4 polish and cohesionOpen — watching runtime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Honest zero until a sovereign monitoring surface exists
watching hu_page
Promotion and monitoring as ONE never-brick circuitMeasured exceed: ma_do_deploy_status in runtime/_hdl_build/nx_mgmt_api.nx, verified at emit. Deploy gates on health, watchdog auto-rolls-back, guards respawn, liveness audits -- one auditable loop in one substrate; the field splits this across four tools and none rolls back Adoption: LIVE-DAEMON — fully adopted (top of its ladder).
Capability-token remote ops with fixed-arg pinningMeasured exceed: tea_run_argv_from in runtime/nx_tool_exec_allow.nx, verified at emit. Remote status/health run pinned allowlisted subs -- a hostile argv is IGNORED (proven 2026-07-08); no ambient authority anywhere; the field uses SSH keys and bearer tokens with full shell reach Adoption: LIB-WIRED importers=8 nonval=3 — fully adopted (top of its ladder).
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. The Nishi supervisor is a REAL running supervision loop with a property none of the field bars carry: promotion and monitoring are ONE never-brick circuit -- deploys are health-gated with automatic rollback, crash guards respawn daemons live on the NAS (a guard restart was captured in this very session's status output), and the liveness auditor exists precisely because BUILT-but-not-LIVE once shipped a dead end. Remote operations ride capability tokens with fixed-arg pinning, not ambient authority. Against the field it is honestly thin on observability depth: no time-series database, no alerting rules, no log aggregation, no anomaly detection, no SLO tracking, no resource-limit enforcement, no federation -- Prometheus and Netdata own those axes. The climb: a sovereign metrics store feeding the /health surface, alert routing through the vizsla reminder engine that already fires deadlines, then per-service resource metering.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/supervisor.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain supervisor, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [systemd-service] systemd project. systemd.service(5) -- Service unit configuration: Restart=, RestartSec, StartLimitBurst and the ExecStart/Type lifecycle (man page as mirrored by man7.org). publisher · read in our library knowledge/fetched/cmp_supervisor_systemd-service5.html · pin h6739e3cd7dd45b902c5fd07d4361f74578b947791f872700ed8bb084571032dd · accessed 2026-08-18 · vendor-docGrounds: The Process supervision with crash-restart guards row where systemd is the init-system bar (Best): Restart= plus start-rate limiting is the documented mechanism hostctl supervise (PID + SERVING probe + crash-loop backoff) is graded against, and the Dependency-ordered startup graph and Resource-limit enforcement rows that name systemd unit graphs and cgroup limits as the bar.
  2. [systemd-socket] systemd project. systemd.socket(5) -- Socket unit configuration: socket-based activation, ListenStream and Accept= (man page as mirrored by man7.org). publisher · read in our library knowledge/fetched/cmp_supervisor_systemd-socket5.html · pin h2a5afec02e7666f911244e888dd1700b45e0d2c36259054b9292b144e21bff1f · accessed 2026-08-18 · vendor-docGrounds: The Socket / timer activation row (_ABSENT_ nx_socket_activate): systemd socket activation is the bar; our daemons are always-on, which is what this row honestly marks 0.
  3. [supervisord] Supervisor project. Supervisor: A Process Control System (supervisord.org) -- supervisord daemon, supervisorctl, program autorestart and startretries, XML-RPC interface and web UI. publisher · read in our library knowledge/fetched/cmp_supervisor_supervisord.html · pin h6521253bca91b8b99bc52f670351da5bf3d04632ff574fcfa09e847275593dcf · accessed 2026-08-18 · vendor-docGrounds: The supervisord column: Process supervision with crash-restart guards (Yes), Health snapshot + service inventory API (supervisorctl status), Web health dashboard (Part) and Alerting rules + notification routing (Part via event listeners) -- the classic userland process controller the Nishi supervise loop is peer to.
  4. [prometheus-overview] Prometheus Authors. Prometheus Overview (prometheus.io/docs/introduction/overview): multi-dimensional time-series data model, PromQL, pull-based scraping, local TSDB and Alertmanager integration. publisher · read in our library knowledge/fetched/cmp_supervisor_prometheus.html · pin h67aef8e1953b5999ee3e69eacc8d956caa0a233b19f117b29f1ee3081fbedea0 · accessed 2026-08-18 · vendor-docGrounds: The Prometheus column: Time-series metrics database (Prometheus TSDB is the bar for the _ABSENT_ nx_tsdb row), SLO / error-budget tracking (recording rules approximate it), Deployed-capability liveness audit (blackbox probing is the partial peer) and Web health dashboard (needs Grafana).
  5. [alertmanager] Prometheus Authors. Alerting Overview -- Alertmanager (prometheus.io/docs/alerting/latest/overview): alerting rules in Prometheus, grouping, inhibition, silencing and notification routing to receivers. publisher · read in our library knowledge/fetched/cmp_supervisor_alertmanager.html · pin hdc5e41383bee1d16eba4d4366a7240a9305cb949608c4c3f72d6a48f16ce3065 · accessed 2026-08-18 · vendor-docGrounds: The Alerting rules + notification routing row (_ABSENT_ nx_alerting): Alertmanager is the bar; the vizsla reminder engine fires medical deadlines but is NOT wired to service alerts -- the named next rung.
  6. [prom-federation] Prometheus Authors. Federation (prometheus.io/docs/prometheus/latest/federation): hierarchical and cross-service federation of time series between Prometheus servers. publisher · read in our library knowledge/fetched/cmp_supervisor_prom-federation.html · pin hf5c25a4f97a6717b323866e71fc999dccb1c62d33d699def2bbd0e62d76a8a94 · accessed 2026-08-18 · vendor-docGrounds: The Multi-host federated monitoring row (_ABSENT_ nx_fed_monitor): Prometheus federation is the bar; laptop + NAS + workers exist but each is watched alone.
  7. [netdata-docs] Netdata Inc. Welcome to Netdata (learn.netdata.cloud): per-second real-time metrics collection, auto-generated dashboards, agent-parent streaming and Netdata Cloud. publisher · read in our library knowledge/fetched/cmp_supervisor_netdata.html · pin h269df330a39c45855862bbbaf70bcc4cecf683c63f85759542d56cced5890885 · accessed 2026-08-18 · vendor-docGrounds: The Netdata column: Per-service resource metrics (per-second visibility is the bar), Web health dashboard (auto-dashboards are the bar) and the U1-U4 UI/PRODUCT axes graded against the NAMED Netdata v2.10.3 dashboard -- U1 visual design, U2 interaction quality, U3 information design, U4 polish and cohesion.
  8. [netdata-ml] Netdata Inc. Machine Learning Anomaly Detection (learn.netdata.cloud/docs/netdata-ai/anomaly-detection): unsupervised k-means models trained per metric on the agent, anomaly bit per sample, Anomaly Advisor. publisher · read in our library knowledge/fetched/cmp_supervisor_netdata-ml.html · pin h3b1a5c33fcf67079acf25b361f0a900b995f5ece3e3d49f12cb0fc578415a1a2 · accessed 2026-08-18 · vendor-docGrounds: The Anomaly detection / AIOps row (_ABSENT_ nx_aiops): Netdata ships ML anomaly bits; nothing sovereign here yet.
  9. [sre-slo] Beyer, Jones, Petoff, Murphy (eds). Site Reliability Engineering: How Google Runs Production Systems, chapter 4 Service Level Objectives (sre.google/sre-book). O'Reilly, 2016. publisher · read in our library knowledge/fetched/cmp_supervisor_sre-slo.html · pin h449fb54ce65e05102fac46c797a08b86c5ab93ade2c70c63ae25e7648d687d7d · accessed 2026-08-18 · published-courseGrounds: The SLO / error-budget tracking row (_ABSENT_ nx_slo): SLI, SLO and error budget are this chapter's definitions; we track gate GREEN, not error budgets, and Prometheus recording rules only approximate it.

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/supervisor.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers