Nishi Family › Compare › Developer Guardrails and the Agentic Control Plane
Nishi Compare · full-field SOTA · measured, not asserted
Developer Guardrails and the Agentic Control Plane
Nishi vs the full field — every axis measured or researcher-sourced, grouped by category; each strip shows the whole field at a glance.
Gates, Teeth, Probes, Telemetry, Accelerators + the MCP swarm/hive/hub-spoke layer. 11 competitors, 26 axes. Nishi cells name the organ that proves them; where nothing proves it the cell is No, not Part.
Measured stakes (quantitative)
Nishi MEASURED 2026-08-14 by nx_capsearch's own corpus count (considered=1028); every one is ocap-gated and callable by an AI seat with an attenuated token. Peer columns reflect published MCP-server availability as of 2026-08 [mcp-spec], not tool counts, which are not comparable across products
GitHubActions MCP server (2025+)GitLabCI MCP serverJenkins plugins, no MCPSonarQube limitedSemgrep limitedSnyk limitedDatadog MCP serverOpenTelemetry n/a specBackstage pluginsLaunchDarkly SDK/APIKubernetes kubectl-ai era
Nishi MEASURED 2026-08-07 by nx_gatesubj + nx_biteall over the whole tree. Published deliberately: 2,287 gate sources nothing invokes is the estate's largest honest gap, and 0 invocation gaps means every gate anything runs is deployed and byte-identical to source
GitHubActions n/a (workflow files)GitLabCI n/aJenkins n/aSonarQube rulesetsSemgrep rulesetsSnyk rulesetsDatadog monitorsOpenTelemetry n/aBackstage n/aLaunchDarkly flagsKubernetes admission ctl
Nishi MEASURED 2026-08-14 via nx_cron_watch: each beat's evidence log carries a heartbeat and a max-age; a missed beat becomes STALE rather than silently absent. Peers schedule work but rarely ship the freshness assertion itself
GitHubActions schedulesGitLabCI schedulesJenkins timersSonarQube n/aSemgrep n/aSnyk n/aDatadog monitorsOpenTelemetry n/aBackstage n/aLaunchDarkly n/aKubernetes CronJob
Nishi never-brick promote: health-checked, .prev-banked, auto-rollback, verified live this session (DEPLOYED-GREEN after a daemon self-deploy). LaunchDarkly's flag flip is the fastest human-facing undo in the field and is graded Best for that reason
GitHubActions manual/actionGitLabCI manualJenkins manualSonarQube n/aSemgrep n/aSnyk n/aDatadog n/aOpenTelemetry n/aBackstage n/aLaunchDarkly instant flag flipKubernetes rolling undo
Quality Gates (the blockers)
Every compiler promotion passes nx_cc_equiv_gate (10 differential rows + a self-host fixpoint) and a 49-cell gauntlet before the canary door; hosted CI is the industrial bar for breadth and parallelism
NO BRANCHES to protect: the substrate is content-addressed and promotion is gated at the artifact, not the ref [gh-branch-protection]. Honest No -- and the sovereign git host is DOWN and unsupervised (filed), so this is a real gap, not a design flourish
Nothing implements a freeze window. LaunchDarkly-class kill switches are the bar. Filed, not designed away
/api/deploy runs a pre-deploy gate and reports DEPLOY-SAFE with a blocker count before promoting; the promote door separately refuses on a digest mismatch, a capability-loss diff, or a backwards artifact
expect_sha256 is REQUIRED on promote: it refuses unless the staged bytes are exactly the ones the caller verified, which is the only check that catches a same-size substitution. Sigstore-class signing is the nearest peer practice [sigstore-docs]
Teeth (the enforcers)
No canonical formatter exists -- gofmt is the bar and this is a filed rung, not a stance
nx_cwe_scan + nx_secret_scan_gate + nx_srclint exist and the secret gate is currently RED and uninvoked (filed); Semgrep/Snyk are the bar for rule breadth and CVE currency
The deepest teeth here and no peer's equivalent: the compiler refuses float-into-integer stores at all six store sites, out-of-range constant shifts, discarded pure expressions, non-exhaustive matches, wrong call arity and both directions of pointer/integer confusion -- each with a 5W plus H message naming the fix. A linter suggests; this refuses. EVIDENCE (this cell carried none until 2026-08-14): nx_langdiag_gate runs the whole language-diagnostic probe corpus, 11 rows GREEN, with in-tree witnesses per refusal (nx_probe_ub_shift.nx for the constant over-shift); each refusal stamps a greppable capability slug at the point of friction -- capability=shift-count-range is emitted from nx_parse.nx:3853 and asserted by the gate, so the claim is checkable at a line number rather than described
nx_langdiag_gate asserts each refusal names ITS OWN rule or capability slug, bite-proven by planting a slug the rule never prints (exit code matched, reason did not). Most suites assert only that something failed -- blind by construction, and this estate had that hole until 2026-08-14
Zero third-party packages by doctrine, so nothing scans them; source carries license_tier headers and nx_licgate exists. Graded No because the CHECK is absent, not because the risk is
nx_gate_bite plants known-bad mutants and demands RED; a gate that has only ever passed is unverified. Essentially absent from mainstream CI practice
No SBOM is produced. Zero third-party packages makes the inventory trivial to state and that is exactly why nothing states it -- an unstated inventory is not an attested one. Syft/Snyk-class output is the bar [spdx-spec]
Nothing scans a shipped artifact for known CVEs. nx_nvd_ingest pulls the feed but no organ joins it to what we run
No OPA/rego-class policy engine; guardrails are hand-written organs, so a new policy is a build rather than a rule. Kubernetes admission control is the bar [k8s-admission] and the 2026 MCP literature specifically names AI-updated global policy as the pattern to reach for
Nothing defines an SLO or an error budget, so nothing can burn one [sre-slo]; resource envelopes are asserted per change but not tracked against a target
Gates write durable verdict logs and a dead-man switch flags a stale beat, but no alert reaches a human -- a RED at 3am waits for a seat to read it
Absent. Resource counters are sampled per beat; there is no CPU/heap profile attributable to a code path
Absent by posture as much as by gap -- the surfaces ship zero third-party script and no telemetry beacon, so field performance is inferred from synthetic checks only
Probes (the observers)
nx_page_verify is browser-grade (TLS handshake, asset fetch, PNG deep-decode, a11y-lite) and runs on the publish beat; Datadog Synthetics is the bar for geographic breadth
Health-checked restart plus a supervisor guard and hostctl sentinel on a 1-minute beat; Kubernetes liveness probes are the bar [k8s-probes]
Deploy waits on a health check and rolls back, but there is no separate readiness signal that holds traffic; Kubernetes is the bar [k8s-probes]
The language-probe corpus runs as a NishiLang gate that asserts exit code AND the required diagnostic substring, with acceptance rows as the discrimination control and an UNRESOLVED third state when the subject is missing. It replaced hand-written shell scripts on 2026-08-14
Telemetry (the monitors)
CORRECTED 2026-08-14: this row previously claimed a per-minute beat. MEASURED against the live clocksched- plane: nx_procchurn 300s and nx_resgov 600s were on the clock, nx_resmon had NO ROW AT ALL, and no cadence was verified for nx_ctxtop. A COMPARE ROW IS A CLAIM, NOT A MEASUREMENT. Now declared: resmonbeat 300s (matching procchurn, its closest analogue -- both are /proc censuses) and memvelbeat 240s. Its leak axis is now VELOCITY-based (nx_memvel's sustained-grower count over N windows) rather than the VmSize==VmPeak SNAPSHOT, which cannot express growth at all, flagged the arena-once leak-FREE design as a leak, and therefore held this organ RED PERMANENTLY at 17 suspects against a red threshold of 6 while the true sustained count was 0. Stale or absent velocity data degrades the axis to UNOBSERVABLE, which casts no vote in either direction. resmon.log is append-only and unbounded, so it is now watched by nx_sizeguard at the sibling's proven budget (1 MiB / 20000 lines) -- fail loud before a reader truncates, never rotate, because nx_jrnlguard correctly treats a shrinking journal as a clobber. Measured attribution proved NAS churn was DSM, not the estate. Datadog is the bar
Every gate writes its own timestamped verdict log because a RED that lands in a vacuum is not a measurement; the actlog journal mines tool/verb frequency
No trace context propagation exists. OpenTelemetry is the standard [otel-spec] and this is a genuine absence
The loop no peer ships: capability-limit refusals stamp greppable capability slugs at the moment of friction, /api/build appends them to a durable demand journal, and the mine re-ranks the build order from real usage. Refusal to roadmap with no human in the circuit. EVIDENCE (this cell carried none until 2026-08-14): the journal is knowledge/status/lang_demand.jrnl -- epoch, slug, target; append-only, bounded at 8 rows per build, and it can never fail a build. LIVE-BITTEN 2026-08-13: a single refused build wrote its 3 rows the moment it failed, which is the whole loop demonstrated end to end rather than asserted
Accelerators (the enablers)
Absent: risky daemons are tested on a spare port by hand. Filed
Config-driven switches exist (opt-in debug emission, an opt-out crash guard, conf-gated lanes) but there is no flag service with targeting or instant kill; LaunchDarkly is the bar
nx_catalog answers is-it-there-and-is-it-WIRED across the full chain -- source, built, staged, promoted, registered, authorised, invoked -- with a weakest-link verdict. Backstage is the field bar for humans; this one is machine-readable and it is how the unregistered addr2line binary was found
The agentic control plane (MCP swarm / hive / hub-and-spoke)
The entire build-gate-promote-deploy loop is driven by an AI seat over MCP [mcp-spec] in this very session; peers expose an MCP server over a product, not a whole sovereign ops plane
X-Nishi-Cap ocap tokens [capmyths03]: attenuate-only, scoped per tool, expiring, revocable by nonce, every delegation audited to a consent log, and fixed-arg pinning so a hostile argv cannot escalate a pinned tool. The MCP ecosystem's own 2026 literature names secrets brokering and agent authorisation as its top unsolved risk. EVIDENCE (this cell carried none until 2026-08-14): least-authority was PROVEN live 2026-07-08 -- a read-scoped cap presented to nx_mgmt is denied with reason 4 while the same cap serves its own tools, and fixed-arg pinning was demonstrated by handing nx_status a hostile [selfswap] argv which it ignored in favour of the pinned sub. Every delegation is appended to cap_consent.log, and revocation is by nonce against cap_revoked.list
One upsert call updates the hub from the sovereign registry; 48 domain spokes re-measure and republish on a 6-hourly beat, and an orphan census reconciles docroot directories against the roster with parts that must sum. EVIDENCE (this cell carried none until 2026-08-14): the census MEASURED census dirs=71 matrix=50 radar=7 hand=14 ORPHANS=0 parts-sum OK on the 2026-08-14 beat -- a printed partition that reconciles, which is the difference between claiming no sprawl and showing that the parts add up to the population
Watch contracts named for a not-yet-existing symbol are measured every emit, so a page flips from OPEN to LANDED the moment the organ ships -- proven twice with real ships, zero hand edits, and the contracts are stored as sovereign plane rows that double as the PM intake queue
A swarm family exists (admit, coord, place, job, heal, shardserve) and seat coordination runs over a workstream plane with leases and checkin/checkout; Kubernetes scheduling is the bar for maturity
Person · product · place — not yet measured for this domain
knowledge/compare/devguardrails.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain devguardrails, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [mcp-spec] Model Context Protocol Specification, revision 2025-06-18: JSON-RPC 2.0 hosts/clients/servers, tools, resources, prompts, and the Security and Trust and Safety principles (user consent, tool safety, least-privilege guidance). publisher · read in our library
knowledge/fetched/cmp_devguardrails_mcp-spec.html· pinh57308609c176192e9a3561fe8d8115235aab6bb72234463b8ff66c7e67ff84af· accessed 2026-08-18 · published-standardGrounds: The "Agent-callable guardrail tools exposed over MCP" and "Guardrails callable by an AI agent over MCP" rows: the protocol every column's MCP-server cell is graded against, and its own security section names consent and tool-authorization as implementor obligations -- which the ocap-token row answers by construction. - [otel-spec] OpenTelemetry Specification (opentelemetry.io/docs/specs/otel): the vendor-neutral standard for traces, metrics, logs and context propagation. publisher · read in our library
knowledge/fetched/cmp_devguardrails_otel-spec.html· pinhfb7ab28ec4e393f3d5c1ffe9173e2fc5d38bbe46aa6725144218fea9314c0c2f· accessed 2026-08-18 · published-standardGrounds: The OpenTelemetry column and the "Distributed tracing across services" row (Nishi n): OTel context propagation is the standard that row names as a genuine absence here. - [k8s-probes] Kubernetes documentation: Configure Liveness, Readiness and Startup Probes -- kubelet-driven liveness restarts and readiness gating of traffic. publisher · read in our library
knowledge/fetched/cmp_devguardrails_k8s-probes.html· pinh2c91059571135379613836c808ab83040714818126088e3dd3b200f5accd1c93· accessed 2026-08-18 · vendor-docGrounds: The Kubernetes column on "Liveness checks on running services" (B) and "Readiness gating before traffic" (B): the probe semantics those Best codes cite; ours has health-checked restart but no separate readiness signal that holds traffic. - [k8s-admission] Kubernetes documentation: Admission Controllers Reference -- validating and mutating admission that intercepts API requests before persistence. publisher · read in our library
knowledge/fetched/cmp_devguardrails_k8s-admission.html· pinh901cf891e255ac593d5865bcd24d95ba29279c79f6bb56290c28b8a75183e6fe· accessed 2026-08-18 · vendor-docGrounds: The "Policy-as-code evaluated at admission" row: Kubernetes admission control is the named bar (B); our guardrails are hand-written organs, so a new policy is a build rather than a rule. - [sre-slo] Beyer, Jones, Petoff, Murphy (eds.). Site Reliability Engineering, chapter Service Level Objectives (Jones, Wilkes, Murphy, Smith): SLIs, SLOs and the error budget as the rate at which SLOs may be missed. publisher · read in our library
knowledge/fetched/cmp_devguardrails_sre-slo.html· pinh449fb54ce65e05102fac46c797a08b86c5ab93ade2c70c63ae25e7648d687d7d· accessed 2026-08-18 · published-courseGrounds: The "SLOs with error budgets driving release decisions" row (Nishi n): the definition of SLO and error budget that row measures against; nothing here defines an SLO so nothing can burn a budget. - [sigstore-docs] Sigstore documentation: keyless signing with cosign, Fulcio certificate authority and the Rekor transparency log for software artifacts. publisher · read in our library
knowledge/fetched/cmp_devguardrails_sigstore-docs.html· pinh008286126a95de87c85a13073dbc269d3c0c62afdbdc244c98eec9d7f2460728· accessed 2026-08-18 · vendor-docGrounds: The "Artifact identity verified before promotion" row: the re-grade note names Sigstore-class signing as the nearest peer practice to expect_sha256-required promotion; this is what that peer is. - [spdx-spec] SPDX (System Package Data Exchange) Specification 3.0.1, The Linux Foundation: the software bill of materials data model and serialization. publisher · read in our library
knowledge/fetched/cmp_devguardrails_spdx-spec.html· pinha905ca11b54700955ba92b4dbdf9d35553852b82bfddf2fabc06cabc194e120a· accessed 2026-08-18 · published-standardGrounds: The "Software bill of materials emitted per build" row (Nishi n): the SBOM format the Syft/Snyk-class bar emits; an unstated zero-dependency inventory is not an attested one, and this is the attestation shape. - [gh-branch-protection] GitHub Docs: About protected branches -- required status checks, required pull-request reviews and approvals, linear history, force-push restrictions. publisher · read in our library
knowledge/fetched/cmp_devguardrails_gh-branch-protection.html· pinhf18b62d250903fa0059c6b04fa7ff32354e770b389fa4de64076cdf03a3453e7· accessed 2026-08-18 · vendor-docGrounds: The GitHubActions column on "Merge/branch protection with required approvals" (B): the vendor's own definition of the control the row grades Nishi n against -- no branches to protect, promotion gated at the artifact instead. - [capmyths03] Miller, Yee, Shapiro. Capability Myths Demolished. Johns Hopkins University Systems Research Laboratory technical report SRL2003-02, 2003 (mirror: Agoric papers). publisher · read in our library
knowledge/fetched/cmp_devguardrails_capmyths03.pdf· pinhb6a3e04e60d7ef08d32900143f8e93acbdcb62e2b63160b604591d7a021f7f42· accessed 2026-08-18 · published-paperGrounds: The "Least-authority capability tokens for agent access" row (Nishi B): the object-capability model whose seven properties (attenuation, revocability, confused-deputy resistance) the X-Nishi-Cap attenuate-only, nonce-revocable, fixed-arg-pinned tokens implement; the ACL-vs-capability distinction the row's peers lack.
Generated by nx_swcompare_sota from knowledge/compare/devguardrails.sota — quantitative axes measured/sourced; researcher-fed (nx_swcompare_research). Zero JS, zero trackers.
Where we are. Shipped and measured, and genuinely ahead of the field: the whole build-gate-promote-deploy loop is agent-callable over MCP under attenuate-only capability tokens [capmyths03], deploys are never-brick with health-checked promote and automatic rollback, the compare surface re-measures and republishes itself on a beat with an orphan census whose parts must sum, and the bounded-read discipline is real -- nx_fs declares its own truncation in the payload and a PostToolUse warner restates the bound so a partial read cannot be quoted as a census. That envelope behaviour is the estate at its best and nothing in this plan weakens it. What is MISSING is the layer underneath: the guard plane that enforces all of this cannot state its own health, the seat is held to laws whose enforcing organs it is not granted, and this domain -- the one that measures the agentic control plane -- had no plan at all, so the ranker structurally refused it and the tool plane was the only major surface in the estate with no computed build order. MEASURED 2026-08-22 by full enumeration of the compare data directory, 349 of 349 entries with truncated=0: devguardrails carried a refs file and a sota sheet and neither a matrix nor a plan. cleanserve is in the same state and is named here so the next reader does not have to re-derive it.
Where we need to go. Make the seat-facing tool plane as honest about itself as the surfaces it guards. Every rung below converts one currently-invisible failure into a row that is emitted whether or not anyone asks: a hook that times out is counted against attempts rather than logged into a denominator-free void, a law that mandates an organ also grants it, a per-call tax is bounded and named, and this domain becomes rankable so the order of the remaining work is computed from the boards rather than chosen by a seat.
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| This domain becomes rankable (DG0) | dg_plan_admitted | LANDED by this file. devguardrails had a refs register and a sota sheet but no plan, proven by full enumeration of buildroot knowledge compare at 349 of 349 entries with truncated=0, so nx_compare_rank refused it and no build order existed for the tool plane. HONEST LIMIT declared rather than hidden: the rank proof itself is BLOCKED, because nx_compare_rank is not in this seat's capability set -- which is precisely rung DG3, so the blocker for proving this rung is itself a rung on this board | Organ | 0 u |
| The guard plane states its own health (DG1) | wg_attempt_counter | The hook guard writes its evidence log ONLY at condemnation sites -- Write-GuardLog is called at the SKIP-ORGAN and TIMEOUT-ORGAN branches and nowhere else, and the script's own header enumerates every logged verdict as a refusal. MEASURED 2026-08-22 over the whole file: 1067 recorded events across 2026-08-10 to 2026-08-22, partitioning exactly as 474 TIMEOUT-ORGAN plus 216 SKIP-ORGAN plus 159 SKIP-VM plus 152 SKIP plus 44 RESET-ORGAN plus 12 TIMEOUT-VM plus 4 TIMEOUT plus 3 RESET plus 2 LAUNCH-FAIL plus 1 RESET-VM, which sums to 1067 with no residual. There is no success call site, so 1067 is a COUNT OF REFUSALS AND NOT A RATE, and hook health is structurally uncomputable -- the estate's own law that every aggregate assertion binds to its denominator, violated in the plane that enforces the estate's laws. ACCEPT: attempts are counted alongside refusals so a rate exists, two consecutive windows reconcile, and a deliberately forced timeout raises attempts and refusals by exactly one each | Organ | 1 u |
| The timeout worklist collapsed to its two causes (DG2) after DG1 | wg_stop_budget | A count without a worklist is not actionable, and a worklist that is not collapsed to causes adjudicates one edit N times. MEASURED 2026-08-22 over the whole log: the 474 organ timeouts are 286 nx_worklog_stop and 134 nx_memplane_run, so TWO Stop-hook organs are 420 of 474 or about 886 permil, and every other subject is 24 or fewer. A live instance was open during this very session, with nx_truncwarn breaker-open and its age climbing from 77 to 139 seconds while being repeatedly skipped -- so the truncation warner the operator had just seen fire was suppressed minutes later. Per-organ extraction differs from the verdict tally by one row because the two counts read different fields, stated rather than smoothed. ACCEPT: both dominant organs complete inside their budget on a real turn, measured on the turn and not on a fixture, or the slow leg moves off the blocking path -- and the fix is proven by the breaker for those keys staying closed across a full session | Organ | 1.5 u |
| The seat can reach the laws it is held to (DG3) | cap_law_floor | The measurement law requires that absence is never asserted from a filtered read and that nx_absent is the organ that proves it. MEASURED 2026-08-22: nx_absent returns capability denied to this seat, and so does nx_compare_rank, the roadmap ranker. The seat is therefore mandated to prove absence with a tool it is not granted, and told to work from the ranked board while unable to ask for the ranking. This is the governor-privilege inversion the estate already banked against a build governor, recurring on the absence-prover and the ranker, and a read-only census is not a privileged act. THIS IS THE FOURTH INDEPENDENT REDISCOVERY OF ONE UNFIXED DEFECT AND NOT A NEW FINDING, so the priors are named here to stop the next reader repeating the investigation: 1785028986 measured register-then-use as a two-actor operation with a manual mint step, 1785436247 measured the entire body and anatomy organ family absent from the standing allow-set, and 1786115553 measured the discovery tools themselves ungranted. That last row also records the CONSEQUENCE this operator report independently reproduced from the outside -- a seat that hits a cap wall on its first probe reaches for shell, which the row names as a concrete driver of the estate's shell-usage rate, so the felt badness of the tool surface and this capability gap are the same defect seen from two ends. The fix shape is known and is not a wider token: the any-grants presenter in nx_tools_api already collects the body capability, the X-Nishi-Cap header, the Bearer header and the query capability and lets the FIRST ONE THAT GRANTS win, and minting everything at once was already refused as allow-too-long at 60 tools, so the remedy is a third scoped read-only capability alongside the two that exist. The refusal text is otherwise exemplary -- it names the exact mint call and tells the caller not to fall back to shell [@mcp-spec]. ACCEPT: a seat holding the standard read capability can run the absence-prover and the ranker, a write or admin operation still refuses from that same capability, and the negative control is that an unminted capability is still denied | Organ | 1 u |
| The per-call hook tax is bounded and named (DG4) | wg_wrap_costwarden | Every other hook on this host runs under the wsl guard, which bounds it five ways and fails open. The PreToolUse hook matching every nishi MCP call does NOT -- it is raw inline PowerShell with no guard wrapper and no timeout, spawning a process and taking a named mutex on EVERY tool call. MEASURED this session: two Thread failed to start errors and one No stderr output from that hook, and the same host exhaustion reached the seat's own tooling when a Bash call failed with uv_spawn while 330 processes were live. ACCEPT: the hook runs under the same bounded launcher as its siblings, a spawn failure degrades to silence rather than to an error banner in the operator's transcript, and the cost of the hook per call is measured and published rather than assumed | Organ | 0.5 u |
| Watch contracts for the tool plane (DG5) after DG0 | dg_matrix_admit | This domain renders from a sota sheet only, so it has no matrix and therefore no watch contracts, and a rung here cannot flip itself when its organ ships -- every status on this board is a hand assertion until it does. Admit a devguardrails.matrix whose gap rows name each open symbol on this plan as an absent-symbol watch contract, so the page re-measures itself on the beat like every other hive domain. ACCEPT: each open rung symbol on this plan resolves to exactly one matrix row, a planted symbol flips its row from open to landed on the next emit with zero hand edits, and the flip is recorded as the receipt and never quoted as the proof of the rung's done-rule | Organ | 1 u |
| Gate adoption adjudicated rather than rebuilt (DG6) after DG5 | ga_adjudicate | The sota sheet already publishes the estate's largest honest gap, and publishing it is not closing it: 2317 gate sources exist, 135 are deployed, 30 are actually invoked, and there are 0 invocation gaps -- so every gate anything runs is deployed and byte-identical to source, and the problem is 2287 sources nothing invokes plus 105 deployed binaries nobody calls. A deployed uninvoked gate is a measurement the estate pays for and does not collect, and it degrades to a false sense of coverage because its existence is counted while its verdict is not. The remedy is adjudication, wire or retire, and explicitly NOT the mass rebuild the 94-percent-undeployed headline invites, which the record already forbids after two of eight sampled rebuilds lost capability. ACCEPT: every deployed uninvoked gate carries a decision of wired or retired with its reason, retirement is reversible through the retire-path organ and never a delete, and the invoked count moves while the invocation-gap count stays at 0 | Organ | 2 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| D0 · The guard plane can state its own health | DG1,DG2 | 2.5 u |
| D1 · The seat can reach its own laws | DG3,DG4 | 1.5 u |
| D2 · The tool plane is rankable and adopted | DG5,DG6 | 3 u |
Risk register
| Risk | Likelihood x impact | Mitigation |
|---|---|---|
| Counting hook attempts adds a write to the hottest path in the session and could itself become the tax it measures | possible x medium | DG1 increments a counter in the file the guard already opens at its existing call sites rather than adding a new artifact or a new process, so the attempt path gains no spawn; if the counter cannot be written the guard still fails open and the absent count reads as UNOBSERVABLE rather than as zero. |
| Granting the absence-prover and the ranker to the read capability widens the read surface | certain x low | DG3 grants exactly two read-only census tools and nothing that writes, promotes or deploys, and the negative control is kept: the same capability must still be refused for a write or admin operation, so least-authority is preserved rather than traded away [@capmyths03]. |
| Adjudicating 2287 gate sources is read as a mandate to build them and becomes a campaign that hammers the host | possible x high | DG6 scopes to the 105 deployed uninvoked binaries, which is a decision list and not a build queue, and states in the rung that the mass rebuild is forbidden; an unbuildable or abandoned subject is retired reversibly rather than hidden inside a queue that then never finishes. |
| A guard-plane health rate is published and then read as a service level with no target behind it | possible x medium | The rate ships as a count and a denominator only. Nothing here defines an SLO or an error budget and the sota sheet already grades that row as absent [@sre-slo], so the number is presented as an observation and the target stays an open gap rather than an implied one. |
Watch contracts (measured)
| Axis | Organ | Symbol | Status | Note |
|---|---|---|---|---|
| The guard plane states its own health | runtime/nx_hookguard.nx | wg_attempt_counter | WATCHING decl | OPEN -- rung DG1. Cells copied from the sota row SLOs with error budgets driving release decisions, which this plan's own risk row already binds to this rung. The hook guard writes its evidence log ONLY at condemnation sites, so its 1067 recorded events across 2026-08-10 to 2026-08-22 are a COUNT OF REFUSALS AND NOT A RATE and hook health is structurally uncomputable. ACCEPT: attempts are counted alongside refusals so a rate exists, two consecutive windows reconcile, and a forced timeout raises attempts and refusals by exactly one each. ORGAN PATH IS A GREENFIELD CONTRACT, NOT AN OVERSIGHT: the guard today is laptop-local PowerShell and nx_hookguard does not exist, PROVEN not assumed by glob over buildroot/runtime returning matches=0 with corpus_complete=1 |
| The timeout worklist collapsed to its two causes | runtime/_hdl_build/nx_worklog_stop.nx | wg_stop_budget | WATCHING decl | OPEN -- rung DG2. Cells copied from the sota row Distributed tracing across services, the field capability that attributes a slow path to its cause. The 474 organ timeouts are 286 nx_worklog_stop and 134 nx_memplane_run, so TWO Stop-hook organs are about 886 permil of them and every other subject is 24 or fewer. The organ named here is the dominant subject and it is real and on disk. ACCEPT: both dominant organs complete inside their budget ON A REAL TURN and not on a fixture, or the slow leg moves off the blocking path, proven by the breaker for those keys staying closed across a full session |
| The seat can reach the laws it is held to | runtime/nx_tools_api.nx | cap_law_floor | WATCHING decl | OPEN -- rung DG3, and the governor-privilege inversion. Cells copied from the sota row Policy-as-code evaluated at admission, where Kubernetes admission control is the bar. The seat is mandated to prove absence with nx_absent and to work from the ranked board, and is granted neither. THIS IS THE FOURTH INDEPENDENT REDISCOVERY OF ONE UNFIXED DEFECT, priors 1785028986 and 1785436247 and 1786115553, named so the next reader does not repeat the investigation. The organ is real and on disk and is the right home: its any-grants presenter already lets the first capability that grants win, so the remedy is a third scoped read-only capability beside the two that exist, never a wider token |
| The per-call hook tax is bounded and named | runtime/nx_hookguard.nx | wg_wrap_costwarden | WATCHING decl | OPEN -- rung DG4. Cells copied from the sota row Continuous profiling in production, the field capability that attributes cost to a code path. The PreToolUse hook matching every nishi MCP call is raw inline PowerShell with no guard wrapper and no timeout, spawning a process and taking a named mutex on EVERY tool call, where every sibling hook runs bounded and fails open. ACCEPT: the hook runs under the same bounded launcher as its siblings, a spawn failure degrades to silence rather than an error banner in the operator transcript, and the per-call cost is measured and published rather than assumed |
| Watch contracts for the tool plane | runtime/nx_compare_regen.nx | dg_matrix_admit | WATCHING decl | OPEN -- rung DG5, the rung this file is. Cells copied from the sota row Hive queue workstream contracts that flip themselves, where the estate holds Best and every rival is honestly zero -- and this domain was the one not using the mechanism it leads the field on. THE DATA HALF LANDED 2026-08-25 and is what made this domain rankable at all. The row STAYS OPEN on purpose: the rung's own accept rule additionally demands that a planted symbol flips its row with zero hand edits, and that proof has not been run. DECLARED LIMIT: this symbol names a data admission rather than a func any organ exports, so the mechanical watch cannot witness it and it will read open until an organ declares it -- a watch contract pointing at a structurally unimplementable symbol can never flip, and saying so is cheaper than letting the next reader discover it |
| Gate adoption adjudicated rather than rebuilt | runtime/_hdl_build/nx_gateadjudicate.nx | ga_adjudicate | WATCHING decl | OPEN -- rung DG6, the estate's largest published honest gap. Cells copied from the sota row Service catalogue with ownership and status, where Backstage is the bar. 2317 gate sources exist, 135 are deployed, 30 are actually invoked and there are 0 invocation gaps, so every gate anything runs is deployed and byte-identical to source and the problem is 2287 sources nothing invokes plus 105 deployed binaries nobody calls. The remedy is adjudication, wire or retire, and explicitly NOT the mass rebuild the 94-percent headline invites, which the record forbids after two of eight sampled rebuilds lost capability. ACCEPT: every deployed uninvoked gate carries a decision with its reason, retirement is reversible through the retire-path organ and never a delete, and the invoked count moves while the invocation-gap count stays at 0 |
watch rows=6 landed=0 watching=6 present=0 missing=0 absent=0 (partition sums)