Nishi Family › Compare › Supervision and Service Monitoring
Nishi Compare · measured, not asserted
Supervision and Service Monitoring
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs systemd and supervisord and Prometheus and Netdata
Where we are. Measured 2026-08-19. Supervision and never-brick promotion are one live circuit with guards, health-gated deploy and rollback, liveness audit and capability-pinned remote ops. Observability depth is the gap the matrix names, and since it was scored (2026-07-10) the estate has shipped pieces the rungs compose rather than invent: nx_netobs (series, probes, a page), nx_resmon (per-process RSS and swap with a trend reader over resmon.log), nx_resgov (memory caps enforced every minute), nx_slo (an importable SLO library), nx_notice and the vizsla notifier. None is yet wired into the supervision surface, which is why the rows stay absent.
Where we need to go. See the fleet on one served surface, act on it (alerts, SLOs, logs), then add the init-system depth systemd is judged on and the anomaly and federation axes Netdata and Prometheus own -- composing the observability organs that already exist.
8 of 23 capabilities measured|2 of them measured exceeds|15 open|coverage 347/1000|adoption 6 full / 2 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883477 domain=supervisor target_version=0.1 rungs=11 done=0 open=11 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | 0.1 | Per-service resource metrics (R1) rmx_sample | 1400 | v=14 m=1 c=10 |
| #2 | 0.1 | Health UI (R0) hu_page | 933 | v=14 m=1 c=15 |
| #3 | 0.1 | Metrics history (R2) ts_append | 400 | v=6 m=1 c=15 |
| #4 | later | Alert routing (R3) nl_route_alert | 1500 | v=15 m=1 c=10 |
| #5 | later | SLO report (R4) hc_slo_report | 1500 | v=3 m=1 c=2 |
| #6 | later | Dependency-ordered startup (R6) hc_dep_order | 1142 | v=8 m=1 c=7 |
| #7 | later | Log aggregation (R5) la_tail_index | 933 | v=14 m=1 c=15 |
| #8 | later | Resource-limit enforcement named (R7) rg_enforce | 700 | v=7 m=1 c=10 |
| #9 | later | Anomaly detection (R9) ao_detect | 350 | v=7 m=1 c=20 |
| #10 | later | Federated monitoring (R10) fm_federate | 350 | v=7 m=1 c=20 |
| #11 | later | Socket and timer activation (R8) sa_listen | 100 | v=2 m=1 c=20 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Health UI (R0) | hu_page | A served monitoring page composing what already exists -- mgmt_snap.json, nx_netobs no_page and no_svg_series, resmon.log -- with what-to-look-at-first ordering; graded by the UI judge against the Netdata bar; closes the dashboard row and U1 to U4 which share this symbol | Organ | 1.5 u |
| Per-service resource metrics (R1) | rmx_sample | CPU, memory and IO per supervised service sampled on the beat (nx_resmon already meters RSS and swap per process; CPU and IO join it) and written as rows; the partition of services sums | Organ | 1 u |
| Metrics history (R2) after R1 | ts_append | An append-only time-series store over the per-service samples with a bounded reader; nx_sizeguard budgets it; resmon.log and netobs series are the first two writers | Organ | 1.5 u |
| Alert routing (R3) | nl_route_alert | A health RED or budget breach routes through the notice ledger to the vizsla reminder engine and email digest already proven for deadlines; a planted RED reaches the digest and a GREEN does not | Organ | 1 u |
| SLO report (R4) after R2 | hc_slo_report | nx_slo (availability permille, error budget, latency percentile) computed per service from the history and printed by hostctl status; a breach is a verdict, not a log line | Organ | 0.5 u |
| Log aggregation (R5) | la_tail_index | The per-daemon /tmp tails indexed into one searchable store on the beat; a query returns the line and its daemon and epoch | Organ | 1.5 u |
| Dependency-ordered startup (R6) | hc_dep_order | Services declare dependencies as data and hostctl starts them in topological order with a loud cycle refusal; a planted cycle is refused | Organ | 1.5 u |
| Resource-limit enforcement named (R7) after R1 | rg_enforce | The resgov memory caps already live are extracted into a named enforcement function and extended to CPU and IO limits as conf rows; a process over its declared limit is governed and the action is logged | Organ | 1 u |
| Socket and timer activation (R8) after R6 | sa_listen | A daemon may declare socket activation so hostctl holds the listener and spawns on first connection; measured idle cost drops to zero for the declared services | Organ | 2 u |
| Anomaly detection (R9) after R2 | ao_detect | Per-series anomaly flags from the history (rate-of-change and seasonal baselines) that ABSTAIN with too little history; a planted spike is flagged and a flat series is not | Organ | 2 u |
| Federated monitoring (R10) after R2 | fm_federate | Laptop, NAS and workers report into one history with per-host partition; the UI shows the fleet, not one box | Organ | 2 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M0 · See the fleet | R0,R1,R2 | 4 u |
| M1 · Act on it | R3,R4,R5 | 7 u |
| M2 · Init-system depth | R6,R7,R8 | 11.5 u |
| M3 · Beyond | R9,R10 | 15.5 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | systemd | supervisord | Prometheus | Netdata |
|---|---|---|---|---|---|
Process supervision with crash-restart guardsMeasured:hc_guard_one exists in runtime/_hdl_build/nx_hostctl.nx, verified at emit. LIVE on the NAS: supervise + per-service guards; a guard restart of nx_relate_daemon was captured in this session's nx_status output; systemd is the init-system bar [systemd-service] Adoption: LIVE-DAEMON — fully adopted (top of its ladder). | ● | ◉ | ● | ○ | ○ |
Never-brick deploy watchdog (health-gated promote + auto-rollback)Measured:func ma_do_deploy exists in runtime/_hdl_build/nx_mgmt_api.nx, verified at emit. DEPLOYED-GREEN proven across mgmt/edge/office arcs; the field bars supervise OR observe -- none promote with rollback (that is Argo-class territory, see /compare/deploy) Adoption: LIVE-DAEMON — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Health snapshot + service inventory APIMeasured:ss_cat exists in runtime/_hdl_build/nx_mgmt_snapshot.nx, verified at emit. mgmt_snap.json written each poll; GET /api/services returned the live 14-service inventory over MCP in this session [supervisord] Adoption: LIB-WIRED importers=2 nonval=1 — fully adopted (top of its ladder). | ● | ● | ● | ◐ | ● |
Deployed-capability liveness audit (BUILT is not LIVE)Measured:audit_row exists in runtime/_hdl_build/nx_reader_liveness.nx, verified at emit. The RACI monitor activity exists because a built reader shipped dead; this organ proves served-and-wired, not just compiled; Prometheus blackbox probing is the partial peer Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ○ | ○ | ◐ | ○ |
Accountability-routed monitoring (RACI lineage)Measured:lr_resolve exists in runtime/_hdl_build/nx_lineage_raci.nx, verified at emit. Monitoring findings resolve to the ONE accountable role; the field routes to dashboards, not owners Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ○ | ○ | ○ | ○ |
Live-derived maturity grading (anti-staleness by design)Measured:em_domain_level exists in runtime/_hdl_build/nx_ecomat_lib.nx, verified at emit. Grades re-derive from evidence files on every read -- a snapshot grade is impossible by construction; exposed live as an MCP tool Adoption: LIB-WIRED importers=21 nonval=13 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Web health dashboardOpen — watchingruntime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. A /health fleet UI exists in the containers arc but no serving symbol verified in THIS fleet's organs -- counted absent honestly; Netdata auto-dashboards are the bar; Prometheus needs Grafana | ○ | ○ | ◐ | ◐ | ◉ |
Time-series metrics databaseOpen — watchingruntime/nx_tsdb.nx : ts_append, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus TSDB is the bar [prometheus-overview]; our snapshots are point-in-time, no history | ○ | ○ | ○ | ◉ | ● |
Alerting rules + notification routingOpen — watchingruntime/_hdl_build/nx_notice.nx : nl_route_alert, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Alertmanager is the bar [alertmanager]; our vizsla reminder engine fires medical deadlines but is NOT wired to service alerts -- the named next rung | ○ | ◐ | ◐ | ◉ | ● |
Log aggregation + searchOpen — watchingruntime/nx_log_aggregate.nx : la_tail_index, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. journald is the bar; our logs are per-daemon /tmp tails read via status | ○ | ◉ | ◐ | ○ | ◐ |
Per-service resource metrics (CPU mem IO)Open — watchingruntime/nx_resmetrics.nx : rmx_sample, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata per-second resource visibility is the bar [netdata-docs]; systemd cgroup accounting partial | ○ | ◐ | ○ | ● | ◉ |
Anomaly detection / AIOpsOpen — watchingruntime/nx_aiops.nx : ao_detect, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata ships ML anomaly bits [netdata-ml]; nothing sovereign here yet | ○ | ○ | ○ | ◐ | ◉ |
SLO / error-budget trackingOpen — watchingruntime/_hdl_build/nx_hostctl.nx : hc_slo_report, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus recording rules approximate it; we track gate GREEN, not error budgets [sre-slo] | ○ | ○ | ○ | ◐ | ○ |
Socket / timer activationOpen — watchingruntime/nx_socket_activate.nx : sa_listen, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd socket activation is the bar [systemd-socket]; our daemons are always-on | ○ | ◉ | ○ | ○ | ○ |
Dependency-ordered startup graphOpen — watchingruntime/_hdl_build/nx_hostctl.nx : hc_dep_order, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd unit graphs are the bar; hostctl starts a flat ordered list | ○ | ◉ | ◐ | ○ | ○ |
Resource-limit enforcementOpen — watchingruntime/nx_resgov.nx : rg_enforce, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. systemd cgroup limits are the bar; rule 21 is awareness, not enforcement | ○ | ◉ | ◐ | ○ | ○ |
Multi-host federated monitoringOpen — watchingruntime/nx_fed_monitor.nx : fm_federate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prometheus federation is the bar [prom-federation]; laptop+NAS+workers exist but each is watched alone | ○ | ○ | ○ | ◉ | ◐ |
U1 visual design of the monitoring surfaceOpen — watchingruntime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. U-axes vs the NAMED product Netdata v2.10.3 dashboard (banked 2026-07-10); we ship no monitoring UI to grade | ○ | ○ | ○ | ◐ | ◉ |
U2 interaction quality (drill-down, filtering, latency)Open — watchingruntime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata per-second interactive charts are the bar; nothing of ours to grade | ○ | ○ | ○ | ◐ | ◉ |
U3 information design (what to look at first)Open — watchingruntime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netdata auto-organizes by host and service; our status output is a raw console dump | ○ | ○ | ○ | ◐ | ◉ |
U4 polish and cohesionOpen — watchingruntime/nx_health_ui.nx : hu_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Honest zero until a sovereign monitoring surface exists | ○ | ○ | ○ | ◐ | ◉ |
Promotion and monitoring as ONE never-brick circuitMeasured exceed:ma_do_deploy_status in runtime/_hdl_build/nx_mgmt_api.nx, verified at emit. Deploy gates on health, watchdog auto-rolls-back, guards respawn, liveness audits -- one auditable loop in one substrate; the field splits this across four tools and none rolls back Adoption: LIVE-DAEMON — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Capability-token remote ops with fixed-arg pinningMeasured exceed:tea_run_argv_from in runtime/nx_tool_exec_allow.nx, verified at emit. Remote status/health run pinned allowlisted subs -- a hostile argv is IGNORED (proven 2026-07-08); no ambient authority anywhere; the field uses SSH keys and bearer tokens with full shell reach Adoption: LIB-WIRED importers=8 nonval=3 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Person · product · place — not yet measured for this domain
knowledge/compare/supervisor.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain supervisor, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [systemd-service] systemd project. systemd.service(5) -- Service unit configuration: Restart=, RestartSec, StartLimitBurst and the ExecStart/Type lifecycle (man page as mirrored by man7.org). publisher · read in our library
knowledge/fetched/cmp_supervisor_systemd-service5.html· pinh6739e3cd7dd45b902c5fd07d4361f74578b947791f872700ed8bb084571032dd· accessed 2026-08-18 · vendor-docGrounds: The Process supervision with crash-restart guards row where systemd is the init-system bar (Best): Restart= plus start-rate limiting is the documented mechanism hostctl supervise (PID + SERVING probe + crash-loop backoff) is graded against, and the Dependency-ordered startup graph and Resource-limit enforcement rows that name systemd unit graphs and cgroup limits as the bar. - [systemd-socket] systemd project. systemd.socket(5) -- Socket unit configuration: socket-based activation, ListenStream and Accept= (man page as mirrored by man7.org). publisher · read in our library
knowledge/fetched/cmp_supervisor_systemd-socket5.html· pinh2a5afec02e7666f911244e888dd1700b45e0d2c36259054b9292b144e21bff1f· accessed 2026-08-18 · vendor-docGrounds: The Socket / timer activation row (_ABSENT_ nx_socket_activate): systemd socket activation is the bar; our daemons are always-on, which is what this row honestly marks 0. - [supervisord] Supervisor project. Supervisor: A Process Control System (supervisord.org) -- supervisord daemon, supervisorctl, program autorestart and startretries, XML-RPC interface and web UI. publisher · read in our library
knowledge/fetched/cmp_supervisor_supervisord.html· pinh6521253bca91b8b99bc52f670351da5bf3d04632ff574fcfa09e847275593dcf· accessed 2026-08-18 · vendor-docGrounds: The supervisord column: Process supervision with crash-restart guards (Yes), Health snapshot + service inventory API (supervisorctl status), Web health dashboard (Part) and Alerting rules + notification routing (Part via event listeners) -- the classic userland process controller the Nishi supervise loop is peer to. - [prometheus-overview] Prometheus Authors. Prometheus Overview (prometheus.io/docs/introduction/overview): multi-dimensional time-series data model, PromQL, pull-based scraping, local TSDB and Alertmanager integration. publisher · read in our library
knowledge/fetched/cmp_supervisor_prometheus.html· pinh67aef8e1953b5999ee3e69eacc8d956caa0a233b19f117b29f1ee3081fbedea0· accessed 2026-08-18 · vendor-docGrounds: The Prometheus column: Time-series metrics database (Prometheus TSDB is the bar for the _ABSENT_ nx_tsdb row), SLO / error-budget tracking (recording rules approximate it), Deployed-capability liveness audit (blackbox probing is the partial peer) and Web health dashboard (needs Grafana). - [alertmanager] Prometheus Authors. Alerting Overview -- Alertmanager (prometheus.io/docs/alerting/latest/overview): alerting rules in Prometheus, grouping, inhibition, silencing and notification routing to receivers. publisher · read in our library
knowledge/fetched/cmp_supervisor_alertmanager.html· pinhdc5e41383bee1d16eba4d4366a7240a9305cb949608c4c3f72d6a48f16ce3065· accessed 2026-08-18 · vendor-docGrounds: The Alerting rules + notification routing row (_ABSENT_ nx_alerting): Alertmanager is the bar; the vizsla reminder engine fires medical deadlines but is NOT wired to service alerts -- the named next rung. - [prom-federation] Prometheus Authors. Federation (prometheus.io/docs/prometheus/latest/federation): hierarchical and cross-service federation of time series between Prometheus servers. publisher · read in our library
knowledge/fetched/cmp_supervisor_prom-federation.html· pinhf5c25a4f97a6717b323866e71fc999dccb1c62d33d699def2bbd0e62d76a8a94· accessed 2026-08-18 · vendor-docGrounds: The Multi-host federated monitoring row (_ABSENT_ nx_fed_monitor): Prometheus federation is the bar; laptop + NAS + workers exist but each is watched alone. - [netdata-docs] Netdata Inc. Welcome to Netdata (learn.netdata.cloud): per-second real-time metrics collection, auto-generated dashboards, agent-parent streaming and Netdata Cloud. publisher · read in our library
knowledge/fetched/cmp_supervisor_netdata.html· pinh269df330a39c45855862bbbaf70bcc4cecf683c63f85759542d56cced5890885· accessed 2026-08-18 · vendor-docGrounds: The Netdata column: Per-service resource metrics (per-second visibility is the bar), Web health dashboard (auto-dashboards are the bar) and the U1-U4 UI/PRODUCT axes graded against the NAMED Netdata v2.10.3 dashboard -- U1 visual design, U2 interaction quality, U3 information design, U4 polish and cohesion. - [netdata-ml] Netdata Inc. Machine Learning Anomaly Detection (learn.netdata.cloud/docs/netdata-ai/anomaly-detection): unsupervised k-means models trained per metric on the agent, anomaly bit per sample, Anomaly Advisor. publisher · read in our library
knowledge/fetched/cmp_supervisor_netdata-ml.html· pinh3b1a5c33fcf67079acf25b361f0a900b995f5ece3e3d49f12cb0fc578415a1a2· accessed 2026-08-18 · vendor-docGrounds: The Anomaly detection / AIOps row (_ABSENT_ nx_aiops): Netdata ships ML anomaly bits; nothing sovereign here yet. - [sre-slo] Beyer, Jones, Petoff, Murphy (eds). Site Reliability Engineering: How Google Runs Production Systems, chapter 4 Service Level Objectives (sre.google/sre-book). O'Reilly, 2016. publisher · read in our library
knowledge/fetched/cmp_supervisor_sre-slo.html· pinh449fb54ce65e05102fac46c797a08b86c5ab93ade2c70c63ae25e7648d687d7d· accessed 2026-08-18 · published-courseGrounds: The SLO / error-budget tracking row (_ABSENT_ nx_slo): SLI, SLO and error budget are this chapter's definitions; we track gate GREEN, not error budgets, and Prometheus recording rules only approximate it.
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/supervisor.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers