Nishi Family › Compare › Failure Modes
Nishi Compare · measured, not asserted
Failure Modes
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
The published failure canon -- Yuan's error-handling result, gray failure, metastable failure, silent data corruption at fleet scale, crash-only recovery, the CWE error-handling family, SRE alerting and error budgets, OpenTelemetry error semantics, chaos engineering, and the 2025-2026 agent-failure taxonomies -- reconciled into NINE classes and then MEASURED over our own 70,100-row failure record.
Where we are. Measured 2026-08-20 full population over the 70,100-row law bank. STRONG and proven: a nine-class taxonomy reconciled from the published canon rather than invented; a classifier whose partition PRINTS AND RECONCILES its own sum, whose every count carries a worklist naming the signature that fired, and which emits NUMBERS AND NO VERDICT because an uncalibrated classifier that votes is a false-alarm generator with an authoritative name; census instruments for the three classes that emit no signal at all; an absence prover that refuses to claim absence from a partial read; and a capability-loss oracle on every promotion. WEAK and equally proven: no alerting discipline of any kind, no SLO, no error budget, no chaos or fault injection, no runtime error-status vocabulary, no differential-observability check, and no error-handler lint -- the one gap where a rival already leads at full strength, because the field has already MEASURED that 92 percent of catastrophic failures live in exactly that code [yuan14]. Incidence trend over the 34,949 dated rows, per 1000 rows of each month: LIMIT chronic and largest every month at 60 then 71 then 65; SILENT-BY-OMISSION growing 16 then 36 then 44; ABSENT growing 7 then 11 then 18; LOUD-AND-IGNORED growing fastest at 0.4 then 1.6 then 4.7; METASTABLE the ONLY class trending down, 7.7 then 5.4 then 2.7. No class is extinct.
Where we need to go. Move this estate from CENSUS-ONLY failure detection to a position where the classes that DO emit a signal are scored like detectors and the classes that do NOT are enumerated by mechanism rather than by a seat remembering to look. The exceed we already hold is that our instruments REFUSE and ABSTAIN; extend that shape to the runtime axis instead of adding dashboards. Two classes on this board are ours and must stay honestly labelled as ours: ABSENT and FABRICATED-EVIDENCE-THAT-PASSES-ITS-OWN-GATE.
17 of 26 capabilities measured|8 of them measured exceeds|9 open|coverage 653/1000|adoption 12 full / 5 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883568 domain=failmodes target_version=- rungs=11 done=3 open=8 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | later | Error-budget burn-rate alerting (FM6) fc_burn_rate | 2000 | v=2 m=1 c=1 |
| #2 | later | Partial failure declared at the interface (FM10) fc_partial_iface | 2000 | v=2 m=1 c=1 |
| #3 | later | Metastable failure: detect the sustaining effect, not the trigger (FM5) fc_sustain_probe | 1000 | v=3 m=1 c=3 |
| #4 | later | Chaos: a steady-state hypothesis with a bounded blast radius (FM7) fc_steady_state | 1000 | v=3 m=1 c=3 |
| #5 | later | Silent data corruption: a fleet screen for a wrong answer (FM9) fc_sdc_screen | 666 | v=2 m=1 c=3 |
| #6 | later | Error-handler static lint: the Yuan three-rule class (FM1) sl_errpath | 200 | v=2 m=1 c=10 |
| #7 | later | Fetch provenance journal: requested url to final url to mirror to hash (FM2) rf_prov_journal | 100 | v=1 m=1 c=10 |
| #8 | later | Claim support: the mirror supports the row, not merely exists (FM8) cr_claim_support | 100 | v=1 m=1 c=10 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Error-handler static lint: the Yuan three-rule class (FM1) after - | sl_errpath | A CATCH-UP RUNG: the ONE row on this board where a rival already leads at full strength, verified in the mirror as CWE-390 Automated Static Analysis at Effectiveness High [@cwe390]. The ranker places it tied-third on rival-deficit and that placement is CORRECT ON ITS OWN TERMS -- deficit 2 against FM5 and FM7 at 3 -- because the case for it is ABSOLUTE PREVALENCE, which the competitive model structurally cannot see and which must NEVER be smuggled in by inflating a rival code. The measured case: almost all (92 percent) of catastrophic failures come from incorrect handling of non-fatal errors EXPLICITLY SIGNALLED in software, 58 percent are catchable by simple testing of the error-handling code alone, and three mechanical rules over handlers -- empty or log-only handler, abort on an over-caught exception, TODO or FIXME in the handler -- would have prevented 33 percent, finding new bugs in nine systems that already ran 400 FindBugs rules [@yuan14]. Extend nx_srclint, do NOT build a second linter. ACCEPT RULE declared in advance: on a runtime-assembled fixture of 8 NishiLang handlers -- 2 empty, 2 log-only, 2 with a TODO, and 2 CORRECT handlers that must NOT fire -- the organ returns exactly the 6 planted sites and zero of the 2 controls, and its false-positive rate over the whole tree is published as a number rather than asserted to be zero. A detector that finds nothing has not found nothing until it has found something planted. | Organ | 1 u |
| Fetch provenance journal: requested url to final url to mirror to hash (FM2) after - | rf_prov_journal | Closes a LIVE, MEASURED defect: 13 of 22 published claims on one domain were absent from their own mirrors while the refs gate passed all of them, because a silently repointed remote is a stale pin that satisfies every existing tooth. Journal the redirect chain at FETCH time and join each ref row against it. ACCEPT RULE declared in advance: the tooth must have THREE states, with UNPROVEN as its own bucket, and it must ship RATCHETED not armed -- every pre-existing mirror predates the journal, so arming fleet-wide turns 68 domains RED at once and produces the permanently-red detector everyone learns to ignore. Bite proof: a fixture whose recorded final-url differs from its requested url must go RED, and an unjournalled legacy row must read UNPROVEN and NOT RED. | Organ | 1 u |
| Alert quality scored as a classifier (FM3) after - | fc_alert_score | DONE 2026-08-21 -- SHIPPED and PROVEN. The ruler lives in nx_verdictlog_lib so the census verb and its referee cannot drift about what a verdict is; nx_verdictlog_gate is GREEN with per-tooth gv_check, NAMED neg-controls, an anti-vacuity tooth first and a positive control, and it was killed in four independent directions by mutants that each died on exactly the tooth naming them. FULL POPULATION: 880 verdict streams and 193,489,903 bytes in knowledge/status, partition reconciles to 880, LATCHED declared a SEPARATE AXIS because it overlaps the partition, distribution published BEFORE any bar and no bar exists anywhere -- every predicate is exact. ACCEPT RULE MET AS WRITTEN: nine currently-shipping gates have fired on EVERY RUN of their entire recorded history, COOPSCHEDGATE 306 of 306 through rebuild_contentloss 269 of 269, which is a constant and not a signal. FALSE-POSITIVE RATE MEASURED AGAINST A FULLY ENUMERATED REAL CONTROL, AND IT INDICTED THIS DETECTOR RATHER THAN THE FLEET: 10 of the 50 offenders v1 named were streams whose verdict DIALECT the ruler had itself declared unreadable and was still asserting LATCHED -- 200 per thousand -- root-fixed so an unreadable dialect is unscored on EVERY axis and not only the convenient ones; honest count 40. DEPLOYMENT STATE DECLARED RATHER THAN IMPLIED: the gate is promoted, registered and has run the production census; the nx_failclass binary carrying this verb is BUILT and STAGED and NOT promoted, refused by build admission on a loaded box, which is a governor obeyed and not a step skipped. Our LOUD-AND-IGNORED class is the fastest-growing thing on our own record, up roughly twelvefold since June, and we have no discipline for it at all. The field's is fully specified: score every rule on precision, recall, detection time and reset time [@sre-burnrate], and admit it only if it detects an otherwise undetected condition that is urgent, actionable and user-visible [@sre-monitoring]. ACCEPT RULE declared in advance: run it over the EXISTING detector fleet and publish the four numbers per detector; the rung is accepted only if at least one currently-shipping detector is shown to score badly enough to justify changing or retiring it. A scorer that flatters every incumbent has measured nothing. | Organ | 1 u |
| Gray failure: compare two vantages, never one (FM4) after - | fc_two_vantage | DONE 2026-08-21 -- SHIPPED and PROVEN. Built on the SAME shared ruler as FM3, so one definition of a verdict serves both rungs and they cannot drift apart. THE ACCEPT RULE IS ENFORCED STRUCTURALLY RATHER THAN BY DISCIPLINE: exactly one place in the comparator can return AGREE or DISAGREE, and control reaches it only after both observations exist, both carry a timestamp, both classify to a declared state and their separation is inside the declared window; every earlier exit returns an UNOBSERVABLE that NAMES its failing conjunct. Both out-of-window neg-controls are carried in the SAME run -- one where the two vantages agree and one where they differ -- so neither answer can be reached by luck, and a missing or self-abstaining vantage abstains rather than acquits. THE WINDOW IS DERIVED AND AUDITABLE: it is read off a bound the estate already declares and then CHECKED against each stream's own observed inter-run interval, printing WINDOW-NARROWER-THAN-OBSERVED-CADENCE when the arithmetic cannot work, because a cadence and a bound are a decidable pair. HONEST SCOPE: the production subject table declares ONE pair and that shortness IS the measurement -- almost everything this estate records is observed once, from the observer's own side, so most subjects cannot be pointed at this detector at all. Gray failure is DIFFERENTIAL OBSERVABILITY -- the app sees unhealthy while the observer sees healthy -- and the remedy is to move from singular heartbeat detection to multi-dimensional health monitoring that approximates the app's own view [@huang17]. We have already shipped the textbook instance: a supervisor whose probe reported failure and repeatedly killed a service while a direct request to the same port returned 200. ACCEPT RULE declared in advance: the organ must report DISAGREE only when the two vantages are observed within a bounded window of each other, and must emit UNOBSERVABLE rather than AGREE when either vantage is missing -- an axis that cannot see must abstain, not acquit. | Organ | 1 u |
| Metastable failure: detect the sustaining effect, not the trigger (FM5) after - | fc_sustain_probe | A system that enters a bad state which persists even when the trigger is removed; the root cause is the sustaining effect and the prescription is a characteristic metric with a declared safe range plus hidden capacity measured by triggered stress test [@bronson21]. SRE names the same shape as positive feedback [@sre-cascading] with retry amplification bounded by a budget under 10 percent and three attempts per request [@sre-overload], and the unjittered storm growing with the square of N [@aws-jitter]. Our witness is banked: a redundant supervisor whose every wedge-kill discarded a partially-completed startup and put a fresh instance at the back of the same lock queue. ACCEPT RULE declared in advance: replay that banked incident and the probe must fire on it, AND must NOT fire on a same-shaped load spike that recovered on its own once the load fell -- distinguishing spike from sustained overload is the entire discriminator. | Organ | 1.5 u |
| Error-budget burn-rate alerting (FM6) after FM3 | fc_burn_rate | REPOINT, NOT A SECOND RULER: `supervisor` owns SLO and error-budget tracking as watch contracts and `devguardrails` owns the release-decision row, both citing the same chapter [@sre-slo]. What is missing is the burn-rate FORM: alert on the rate of budget consumption over a long window for precision and a short window so the alert stops when the burn stops [@sre-burnrate], with the budget measured by a neutral instrument and bound to a mechanical consequence [@sre-errorbudget]. ACCEPT RULE declared in advance: this rung is REFUSED if it stands up a second SLO definition anywhere -- it must consume the supervisor row's, or it does not ship. | Organ | 0.5 u |
| Chaos: a steady-state hypothesis with a bounded blast radius (FM7) after - | fc_steady_state | Define steady state as a MEASURABLE OUTPUT, vary real-world events, run against a control group, and set out to DISPROVE the hypothesis [@chaos-principles], sampling stimuli from past postmortems and automating because confidence in a past pass DECAYS [@chaos-basiri2016]. This rung is what finally exercises the recovery path, which is the defect crash-only software was designed to remove by making recovery the ONLY path so that recovery code runs at every start [@candea03]. ACCEPT RULE declared in advance: blast radius is bounded BY CONSTRUCTION and the first experiment runs against a fixture, never the serving root; and the never-brick guarantee is unconditional -- no fault injection may touch persistent hardware state. | Organ | 1.5 u |
| Claim support: the mirror supports the row, not merely exists (FM8) after FM2 | cr_claim_support | Provenance (FM2) proves the mirror is OF the url. It does NOT prove the mirror SUPPORTS the claim, and that second gap had to be closed by hand today, 22 claims at a time. Honest scoping: this is a SEMANTIC check and may never be fully mechanical. ACCEPT RULE declared in advance: ship the decidable half only -- every numeric figure in a row note must appear as a literal in its cited mirror, or the row is flagged UNWITNESSED (a third state, not RED). An abstract page cannot witness a number from the paper body. | Organ | 1 u |
| Silent data corruption: a fleet screen for a wrong answer (FM9) after - | fc_sdc_screen | The hardware floor under our whole SILENT class. SDCs are not captured by CPU error reporting and surface as application bugs on a host with clean event and kernel logs [@dixit21]; Google frames the same defect as mercurial cores, ranks wrong answers that are NEVER DETECTED as the top risk, and uses recidivism on the same core as the discriminator [@hochschild21]. Both prescribe scheduled screening with RANDOMIZED data inputs because corruptions are data dependent. ACCEPT RULE declared in advance: reuse an end-to-end checksum we already compute rather than adding a second pass over the same bytes, and schedule as a low-priority beat that yields to build admission -- a screen that hammers the array is a bug even when it works. | Organ | 1.5 u |
| Partial failure declared at the interface (FM10) after - | fc_partial_iface | Carried DELIBERATELY as a rung that may never be built, because the honest boundary of this whole board belongs on it. Partial failure and indeterminacy cannot be fixed as implementation quality: they must be declared at the TYPE level, with cause-reporting, an undetermined-cause recovery path, idempotency and request IDs placed IN THE INTERFACE [@waldo94]. The related agent-side shape is per-call and structural [@toolfailbench-2026]. ACCEPT RULE declared in advance: this rung is accepted by a WRITTEN INTERFACE CHANGE, never by a detector; if a seat proposes a runtime detector for it, that is evidence the rung was misunderstood. | Organ | 0.5 u |
| Refusal shape: every refusal names its subject and its remedy (FM11) after - | fc_refusal_shape | DONE 2026-08-20 -- SHIPPED and PROVEN: the ruler lives in nx_refusal_shape_lib so the census verb and its gate cannot drift apart, nx_refusal_shape_gate is 46 of 46 GREEN and declared in knowledge/organ_gate.conf, and the bar is a SET OF NAMES self-baselined at 9,039 offenders of a 14,494-block population whose partition reconciles. The accept rule below was met as written and the false-positive rate was MEASURED against a fully enumerated real control rather than asserted: 4 of 17 flagged rows in nx_tools_api.nx, 235 per thousand, with zero known-good guards flagged and the two residual classes named. The mechanizable half of the tenth cell, LOUD-AND-CORRECT, added 2026-08-20 after the detector census found 3 of ~25 seed instances were good behaviour rather than defects. A taxonomy with no cell for a correct guard will get one FIXED, which is the worst possible remediation. Emitter-side only: does the refusal name its SUBJECT and a REMEDY [@sre-monitoring] [@otel-error-type]. ACCEPT RULE declared in advance: run it over the estate's EXISTING refusal messages and publish the pass rate; it is accepted only if it PASSES the three known-correct guards (the nx_fs read truncation envelope, the build-admission refusal, the capability-denied mint recipe) AND FAILS a crafted refusal that names no remedy. A shape checker that flags a known-good guard is worse than none, because someone will act on it. The caller-heeded half is explicitly OUT OF SCOPE for this rung: it is census-only over the call record and belongs with the trajectory scanner, since a refusal callers routinely retry through is LOUD-AND-IGNORED however well written it is | Organ | 0.5 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M1 · The signalling classes get detectors | FM1,FM3,FM4,FM5 | 4.5 u |
| M2 · The evidence classes get provenance | FM2,FM8 | 2 u |
| M3 · The runtime classes get exercised | FM6,FM7,FM9 | 3.5 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | GoogleSRE | OpenTelemetry | CWE-MITRE | NetflixChaos |
|---|---|---|---|---|---|
Failure taxonomy measured over our own failure recordMeasured exceed:fc_scan in runtime/nx_failclass.nx, verified at emit. MEASURED FULL POPULATION 2026-08-20 over 70,100 law rows and 22,491,252 bytes. Partition PRINTS AND RECONCILES its own sum: SILENTOMIT 2546 SILENTDISCARD 203 LOUDWRONG 220 LOUDIGNORED 171 LIMIT 5089 METASTABLE 252 GRAY 91 ABSENT 1053 FABRICATED 310 MULTI 882 UNCLASSIFIED 59283 = 70100. No competitor classifies its OWN failure record against a published taxonomy -- the field publishes taxonomies and studies other people's postmortems [mast-cemri2025] Adoption: LIVE — fully adopted (top of its ladder). | ◉ | ○ | ○ | ● | ○ |
Word-bounded matching with UNCLASSIFIED as its own bucketMeasured:fc_find_word exists in runtime/nx_failclass.nx, verified at emit. OpenTelemetry is Best and we adopted its shape: error.type must be low cardinality and ships `_OTHER` as a built-in fallback so a novel failure surfaces rather than vanishing [otel-error-type]. Word boundaries are load-bearing here and were measured -- bare substring cap hits 7802 rows, the WORD cap hits 2992, and the 3204-row residue is capability and capture and capsearch and capacity. A substring matcher would have published a LIMIT rate inflated by 160 percent Adoption: LIVE — fully adopted (top of its ladder). | ● | ○ | ◉ | ○ | ○ |
Three-state verdict: an axis that cannot see ABSTAINSMeasured:gv_need exists in runtime/nx_gate_verdict.nx, verified at emit. OpenTelemetry Best: span status is Unset or Ok or Error with a total order Ok greater than Error greater than Unset, and instrumentation libraries SHOULD NOT set Ok -- the default posture of automatic instrumentation is SILENCE, not acquittal [otel-span-status]. That is our own abstain-not-acquit law written into a wire format by an independent body Adoption: LIB-WIRED importers=1506 nonval=108 — fully adopted (top of its ladder). | ● | ● | ◉ | ○ | ○ |
Per-tooth counters: declared equals executed by constructionMeasured exceed:gv_check in runtime/nx_gate_verdict.nx, verified at emit. A hand-rolled tally can print passed 22 of 20, and a tooth that silently stops running lowers BOTH numbers and still reads GREEN. No named competitor makes declared-equals-executed structural Adoption: LIB-WIRED importers=1506 nonval=108 — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Capability-loss oracle refuses a promote that deletes runsOpen — no implementing organ is measured for this axis yet. Every promotion passes a diff that NAMES each lost run. The closest thing in the field is a fitness function, and none of the four columns imposes this proof obligation on a deploy | ○ | ○ | ○ | ○ | ○ |
Absence proven with declared coverage, or refused outrightMeasured exceed:ab_has in runtime/nx_absent.nx, verified at emit. Exit 0 ABSENT-PROVEN only when matches=0 AND coverage_complete=1 AND corpus_complete=1 with no truncation marker; exit 3 UNPROVEN otherwise, in which case NO conclusion is available in either direction. This admission itself was proven that way. PRESENCE needs one witness, ABSENCE needs exhaustive coverage Adoption: LIVE — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Dead-path census: defined and never called, whole treeMeasured:uw_defcomplete exists in runtime/_hdl_build/nx_unwired.nx, verified at emit. CWE names dead code as a weakness class and SAST is Best at finding it in one language. Ours is whole-tree over 7132 sources and 27678 functions with a NAME-SET ratchet, because a count-only ratchet on a shared tree reports a regression without saying whose Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ○ | ○ | ◉ | ○ |
Gate-subject resolution: what does this detector actually watchMeasured exceed:gs_execsubj in runtime/_hdl_build/nx_gatesubj.nx, verified at emit. Resolves each gate's subject FROM THE GATE'S OWN SOURCE with paths STAT'd not assumed: 907 permil of 2316 gates resolved against 8 permil declared. A detector whose subject nobody can name cannot be audited Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ◉ | ○ | ○ | ○ | ○ |
Limit guard with a fourth UNMEASURED stateMeasured:sg_verdict exists in runtime/nx_sizeguard.nx, verified at emit. SRE is Best: saturation is one of the four golden signals and the errors signal already contains our silent class -- an error counts when it is implicit, for example an HTTP 200 success response coupled with the wrong content [sre-monitoring]. Ours watches BYTES and LINES with the worst axis deciding and a distinct UNMEASURED verdict Adoption: LIVE — fully adopted (top of its ladder). | ● | ◉ | ○ | ○ | ○ |
Envelope audit: does the truncation marker exist at allMeasured exceed:ea_has in runtime/_hdl_build/nx_envelope_audit.nx, verified at emit. A cap reached in silence becomes a measurement nobody knows is partial. No competitor audits its own tools for whether they DECLARE their truncation Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ◉ | ○ | ○ | ○ | ○ |
Loss classification carries a named reasonMeasured:lc_classify exists in runtime/_hdl_build/nx_lossclass.nx, verified at emit. OpenTelemetry Best via error.type as a bounded class label [otel-error-type]. A count without a worklist is not actionable, and a worklist without the reason is one step short Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked). | ● | ● | ◉ | ○ | ○ |
Leak as a derivative: two samples in time, horizon publishedMeasured:mv_scan exists in runtime/_hdl_build/nx_memvel.nx, verified at emit. A level cannot express a leak. Ours samples five windows and requires growth in EVERY one, so a zero proves NO FAST LEAK and not NO LEAK -- the horizon is published beside the verdict because a horizon that is not published reads as completeness Adoption: LIVE — fully adopted (top of its ladder). | ● | ◉ | ○ | ○ | ○ |
Source lint for a test that cannot failMeasured exceed:sl_lhs in runtime/_hdl_build/nx_srclint.nx, verified at emit. The cursor-sentinel idiom, where a loop exits by clobbering the variable that holds the answer. A vacuous-test detector is the only thing that can find a bug whose symptom is a PASSING TEST -- and the field's own measurement agrees the class is real: Yuan et al. found five empty catch blocks in their own checker's source [yuan14] Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ◉ | ○ | ○ | ● | ○ |
Agent trajectory anti-patterns over the work recordMeasured exceed:tj_is_verify in runtime/_hdl_build/nx_trajscan.nx, verified at emit. Deterministic and judge-free: verification-skip, search-loop, oracle-edit, retry-echo. This is the estate's answer to the agent-failure taxonomies [msft-agentic-2026], whose 2026 revision states plainly that detection requires behavioral analysis across the full session. It matters that ours is judge-free: no judge configuration across five models and five prompt conditions exceeds AUROC 0.65 at detecting false success, even when handed the full ground-truth specification, because judges anchor on confident closing language [false-success-2026] Adoption: LIVE — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Citation register measured, never trustedMeasured exceed:main in runtime/nx_compare_refs_gate.nx, verified at emit. Ten teeth over 68 registers and 667 rows: schema, key uniqueness, url shape, mirror existence, content pins and inline cite marks each carry their own tooth, with two bite-proven negative controls. This domain's own rows were added only after the baseline was confirmed GREEN, so any failure it now reports is ours Adoption: GATE:LIVE trial=- — fully adopted (top of its ladder). | ◉ | ○ | ○ | ○ | ○ |
Error-handler static rules: the Yuan three-rule classOpen — watchingruntime/_hdl_build/nx_srclint.nx : sl_errpath, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. A CATCH-UP RUNG, and the ranker is right to price it as one -- this is the ONE row on this board where a rival already leads at full strength. VERIFIED IN THE MIRROR 2026-08-20, not asserted: CWE-390 carries an Automated Static Analysis detection method at Effectiveness High, so the field detects discarded-error paths today and we do not [cwe390]. The prize is large in ABSOLUTE terms and small in RIVAL-DEFICIT terms, which is exactly why it ranks below rungs where two rivals lead: Yuan et al. measured that almost all (92 percent) of catastrophic failures come from incorrect handling of non-fatal errors EXPLICITLY SIGNALLED in software, 58 percent catchable by simple testing of the error-handling code, and three mechanical rules over handlers -- empty or log-only handler, abort on an over-caught exception, TODO or FIXME in the handler -- would have prevented 33 percent, finding new bugs in nine systems already running 400 FindBugs rules [yuan14]. AN ABSOLUTE-PREVALENCE ARGUMENT IS NOT COLUMN-SHAPED AND MUST NEVER BE SMUGGLED IN BY INFLATING A RIVAL CODE: if this rung is to outrank, it does so through a declared sponsor row carrying verbatim grounds, never through this matrix [cwe755] | ○ | ○ | ○ | ◉ | ○ |
Alert quality scored as a classifierMeasured:fc_alert_score exists in runtime/nx_failclass.nx, verified at emit. SHIPPED AND MEASURED 2026-08-21. fc_alert_score in nx_failclass over the shared ruler nx_verdictlog_lib, refereed by nx_verdictlog_gate. FULL POPULATION, NEVER A SAMPLE: knowledge/status yields 880 verdict streams and 193,489,903 bytes, and the partition PRINTS AND RECONCILES ITS OWN SUM -- SCORED-INFORMATIVE 55, CONSTANT-NONGREEN 28, CONSTANT-GREEN 87, UNKNOWN-VOCAB 76, SNAPSHOT 433, NO-VERDICT 201, UNREADABLE 0 = 880, with 1,275 files skipped by the declared extension list and that scope printed beside the total so every count is a FLOOR whose direction is stated. LATCHED is carried as a SEPARATE AXIS, never a partition member, because a stream can be SCORED-INFORMATIVE and latched at once and folding an overlapping class into a partition breaks its sum. THE FOUR SRE NUMBERS ARE PUBLISHED PER DETECTOR AND EACH IS THREE-STATE, and that is not a softening of the rule but the only honest reading of the evidence: a verdict log records what a detector SAID and never whether it was right, so a precision figure computed from these files would be a constant wearing the shape of a measurement. Precision reads UNINFORMATIVE-CONSTANT where the output never varies and UNOBSERVABLE-NO-ADJUDICATION otherwise; recall reads UNVERIFIED-NEVER-FIRED or UNOBSERVABLE-NO-GROUND-TRUTH; detection time is the largest observed interval between runs and is LABELLED A BOUND; reset is the observed recovery latency, or LATCHED, or NO-EPISODE. NO THRESHOLD EXISTS ANYWHERE -- every predicate is exact, so there is no bar to tune and none to flatter -- and THE DISTRIBUTION IS PUBLISHED BEFORE ANY BAR: across 246 scored series the firing rate spreads 177 streams at 0 per-mil, 41 across the middle bins and 28 at 1000, so the signal demonstrably does not saturate. THE ACCEPT RULE IS MET AND IT IS NOT CLOSE. Nine currently-shipping gates have fired on EVERY RUN of their entire recorded history with an identical reason: COOPSCHEDGATE 306 of 306, shipcheck 320 of 320, nndev 311 of 311, papers_beat 315 of 315, ale_format 309 of 309, driver_exceed 306 of 306, doctor_import 305 of 305, account_admin_census 303 of 303, rebuild_contentloss 269 of 269 -- four of them read directly and confirmed to carry verdict=RED with reason=emit-failed. A detector that has never once been green is a constant, and a constant carries no information no matter how urgent its wording. FALSE-POSITIVE RATE MEASURED AGAINST A REAL CONTROL, AND IT FOUND A DEFECT IN THIS DETECTOR RATHER THAN IN THE FLEET: v1 named 50 offenders, every one enumerated at corpus_complete=1 and reconciled by two independent counters, and hand-adjudicating them surfaced ONE WHOLE KIND of false positive -- 10 streams whose verdict DIALECT the ruler had already declared unreadable were still being asserted LATCHED, so uiq_sentinel, which spells success verdict=CLEAN across 339 healthy runs, was named an offender. 200 per thousand, fixed at the ROOT so an unreadable dialect is unscored on every axis rather than only the convenient ones, and the honest offender count is 40. A first hypothesis about a second false-positive class was REFUTED by reading the streams: virtio_net, coopsched and telemetry look like bring-up transcripts by filename and are in fact NETHSGATE, COOPSCHEDGATE and DRVF1GATE. SRE is Best here and remains so: it treats an alerting rule as a DETECTOR and scores it on precision, recall, detection time and reset time [sre-burnrate] and admits it through five actionability questions [sre-monitoring]. What this rung buys is that our own fleet is now scored that way and the worklist names every stream with its numbers on the row An alerting rule is treated as a DETECTOR and measured like one on precision, recall, detection time and reset time [sre-burnrate], and admitted through five questions of which the first is whether the rule detects an otherwise undetected condition that is urgent and actionable and user-visible [sre-monitoring]. As of 2026-08-21 we DO score them this way and the claim that we did not is retracted here rather than left standing; LOUD-AND-IGNORED remains the fastest-growing class on our own record, and the 40 named offenders are the first worklist we have ever had for it Adoption: LIVE — fully adopted (top of its ladder). | ● | ◉ | ○ | ○ | ○ |
Error-budget burn-rate alertingOpen — watchingruntime/nx_failclass.nx : fc_burn_rate, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. REPOINT NOT DUPLICATE: `supervisor` already carries SLO and error-budget tracking as watch contracts and `devguardrails` carries the release-decision row, both citing the same chapter [sre-slo]. What is missing HERE is the burn-rate form -- alert on the RATE of budget consumption over two windows, a long one for precision and a short one so the alert stops when the burn stops [sre-burnrate], with the budget itself measured by a neutral instrument and bound to a mechanical consequence [sre-errorbudget] | ○ | ◉ | ○ | ○ | ○ |
Fetch provenance: requested url to final url to mirror to hashOpen — watchingruntime/nx_research_fetch.nx : rf_prov_journal, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. MEASURED FAILURE, 2026-08-20: 13 of 22 published claims on one domain were ABSENT from their own mirrors and the refs gate passed all of them, because a silently REPOINTED remote is a stale pin that satisfies every check -- the fetch succeeded, the file is real, and the content moved underneath the citation. Journalling the redirect chain at FETCH time closes the class structurally. Must ship RATCHETED, not armed: every pre-existing mirror predates the journal, so arming it fleet-wide turns 68 domains RED at once and teaches everyone to ignore the gate | ○ | ○ | ○ | ○ | ○ |
| Claim support | |||||
the mirror SUPPORTS the row, not merely existsOpen — watchingruntime/nx_compare_refs_gate.nx : cr_claim_support, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The harder half, and nobody in the four columns does it either. Provenance proves the mirror is OF the url; it does NOT prove the mirror SUPPORTS the claim. Today that gap was closed by hand, 22 claims at a time, over four mirrors read byte by byte in sequential 20 KB windows. An abstract page cannot witness a number from the paper body, and a URL path fragment is not a statement of fact | ○ | ○ | ○ | ○ | ○ |
| Gray failure | |||||
two vantages compared, not oneMeasured:fc_two_vantage exists in runtime/nx_failclass.nx, verified at emit. SHIPPED AND MEASURED 2026-08-21. fc_two_vantage in nx_failclass over the same shared ruler nx_verdictlog_lib that FM3 uses, so there is exactly ONE definition of what a verdict is and the two rungs cannot drift apart about it. THE ACCEPT RULE IS ENFORCED STRUCTURALLY, NOT BY DISCIPLINE: there is exactly one place in the comparator where AGREE or DISAGREE can be returned, and control reaches it only after both observations exist, both carry a timestamp, both classify to a declared state, and their separation is inside the declared window. Every earlier exit returns an UNOBSERVABLE that NAMES ITS FAILING CONJUNCT -- A-UNREADABLE, B-UNREADABLE, A-NO-TIMESTAMPED-VERDICT, B-NO-TIMESTAMPED-VERDICT, OUTSIDE-WINDOW, or A-VANTAGE-ABSTAINS -- because a compound assertion that will not name its failing conjunct is a false-alarm generator and the reader always guesses the alarming one. Refereed by nx_verdictlog_gate over eight runtime-assembled subjects whose outcome partition RECONCILES to 8, and the two neg-controls that matter both fire: an out-of-window pair is UNOBSERVABLE and never AGREE even when both vantages say the same thing, and never DISAGREE when they differ, with BOTH cases present in the same run so neither answer can be reached by luck. A missing vantage ABSTAINS and names which side is missing; a vantage that itself abstains does not become a DISAGREE. THE WINDOW IS DERIVED, NOT CHOSEN, AND IT IS AUDITABLE: it is read off a bound the estate already declares -- nx_cron_watch carries api-contract-probe at max_age 1800 seconds, the larger declared bound of the pair -- and the organ then CHECKS that number against each stream's own observed inter-run interval and prints WINDOW-NARROWER-THAN-OBSERVED-CADENCE when the arithmetic cannot work, because a cadence and a bound are a decidable pair and a row that can never be co-observed is UNOBSERVABLE by construction rather than by accident. HONEST SCOPE, STATED RATHER THAN LEFT TO LOOK LIKE COVERAGE: the production subject table declares ONE pair -- the internal /api/health probe against a client-side HTTP fetch of the served sites -- and that shortness IS the measurement. Almost everything this estate records is observed once, from the observer's own side, so most subjects cannot be pointed at this detector at all; adding one is a data admission in knowledge/verdictlog.conf and never a code change. AND THE DETECTOR'S OWN PATH CARRIES THE DEFECT IT HUNTS, which is said here rather than discovered later: it reads two RECORDINGS, so a vantage that stops recording appears as UNOBSERVABLE and never as DISAGREE -- the conservative direction, but not the same as seeing it, which is exactly Huang's point that the heartbeat traverses the working link [huang17] Adoption: LIVE — fully adopted (top of its ladder). | ● | ◉ | ○ | ○ | ○ |
Metastable failure: the sustaining effect outlives the triggerOpen — watchingruntime/nx_failclass.nx : fc_sustain_probe, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Bronson et al. define the class as a system that enters a bad state which persists even when the trigger is removed, and locate the root cause in the SUSTAINING EFFECT rather than the trigger; the prescription is a characteristic metric with a declared safe range, plus hidden capacity measured by triggered stress test [bronson21]. SRE names the same shape as a cascading failure that grows over time as a result of positive feedback [sre-cascading], with retry amplification bounded by a retry budget under 10 percent and three attempts per request [sre-overload], and the arithmetic of an unjittered retry storm growing with the square of N [aws-jitter]. Our own witness: a redundant supervisor whose every wedge-kill discarded a partially-completed startup and put a fresh instance at the back of the same lock queue, converting a slow start into an unrecoverable one | ○ | ◉ | ○ | ○ | ● |
| Chaos | |||||
a steady-state hypothesis with a bounded blast radiusOpen — watchingruntime/nx_failclass.nx : fc_steady_state, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Netflix is Best. The discipline is to define steady state as a MEASURABLE OUTPUT, vary real-world events, run against a control group, and set out to DISPROVE the hypothesis rather than confirm it [chaos-principles], sampling stimuli from past postmortems and automating because confidence in a past pass DECAYS [chaos-basiri2016]. Precision note carried deliberately: the IEEE paper states FOUR principles and Minimize Blast Radius is website-only, so citing five to the paper is wrong. Nothing here injects a fault on purpose | ○ | ● | ○ | ○ | ◉ |
Silent data corruption: the fleet screen for a wrong answerOpen — watchingruntime/nx_failclass.nx : fc_sdc_screen, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The hardware floor under our whole SILENT class, and it does not announce: SDCs are not captured by CPU error reporting and propagate up as application-level problems, on a host with clean system event logs and clean kernel logs, detected only by an application data inconsistency [dixit21]. Google frames the same defect as mercurial cores and ranks its symptoms in increasing order of risk, with wrong answers that are NEVER DETECTED at the top, using recidivism on the same core as the discriminator [hochschild21]. Both prescribe scheduled fleet-wide screening with randomized data inputs, because corruptions are data dependent. We run none | ○ | ◉ | ○ | ○ | ○ |
Every refusal names its subject and its remedyMeasured:fc_refusal_shape exists in runtime/nx_failclass.nx, verified at emit. SHIPPED AND MEASURED 2026-08-20. fc_refusal_shape in nx_failclass over the shared ruler nx_refusal_shape_lib, refereed by nx_refusal_shape_gate at 46 of 46 GREEN; both organs promoted, _offc-twinned and registered, and the gate declared in knowledge/organ_gate.conf so the ship loop resolves it by DECLARATION rather than by guessing a name. FULL POPULATION, NEVER A SAMPLE: 18,637 sources and 117,259,362 bytes yield 114,420 literal blocks and a population of 14,494 refusal blocks, and the partition PRINTS AND RECONCILES ITS OWN SUM -- SHAPED 674, NO-REMEDY 5,016, NO-SUBJECT 352, BARE 3,671, UNKNOWN 534, TEST-ASSERTION 2,963, USAGE-BANNER 1,284, offenders 9,039 which is 624 per thousand of the population. THE LAST TWO BUCKETS EXIST BECAUSE THE FIRST RUN'S HITS WERE READ RATHER THAN SUPPRESSED: v1 flagged 920 per thousand, and reading the flagged rows found gate tooth text and call-grammar banners in there, neither of which is a refusal a caller reads, so each became its own named bucket that sums into the partition and is excluded from the offender count. THE UNIT IS THE COMPOSED EMISSION, NOT THE LITERAL, and that is forced by the evidence: the truncation envelope is built from three pieces with its remedy in a DIFFERENT literal, so a per-literal checker would flag a known-good guard. FALSE-POSITIVE RATE MEASURED AGAINST A REAL CONTROL, NOT ASSERTED: every one of the 28 refusal blocks in nx_tools_api.nx was adjudicated by hand at coverage_complete=1 -- 4 of 17 flagged rows are false positives, 235 per thousand, and ZERO of the 11 SHAPED rows is a known-good guard wrongly flagged. The two residual false-positive classes are NAMED rather than legislated away: a non-message literal block such as served markup or a route table, and a success-path note that merely mentions a refusal word. ACCEPT RULE MET AS WRITTEN: all three known-correct guards score SHAPED and a crafted remedy-less refusal is FLAGGED, bite-proven with 13 mutants of which 12 die on exactly the tooth that names them, the thirteenth being a defence-in-depth survivor whose whole-invariant twin kills it. THE BAR IS A SET OF NAMES, self-baselined at 9,039 in knowledge/status/refusal_shape.baseline and printed in every verdict, so adoption cannot turn the fleet red on day one. TWO DEFECTS THIS RUNG FOUND IN ITSELF AND FIXED AT THE ROOT: a fixture the defect could not fail, caught only because a mutant survived; and TWO RULERS for one concept, caught only because the census said 9,039 offenders while the ratchet baselined 13,286 over the same file in the same run. THE MECHANIZABLE HALF OF A TENTH CELL ADDED 2026-08-20: LOUD-AND-CORRECT, a refusal that is accurate, complete and names its own remedy, which is NOT a failure and must never be counted as one. Declared a SEPARATE AXIS, never a tenth partition member, and the proof it must be is one instance that scores on BOTH axes at once: nx_fs read declares its truncation cap AND its remedy in its last line (emitter CORRECT) and the reader filtered that line out and filed it as a critical defect (caller IGNORED). Testing the tempting single-axis rule -- correct if the caller obeyed -- puts that instance in LOUD-AND-IGNORED, excluding the very case that motivated the cell. GoogleSRE is Best: every page must be actionable and every rule passes five admission questions [sre-monitoring]. OpenTelemetry is partial: it binds the Description to the Error status and holds error.type to low cardinality, so the reason travels with the failure and cannot be attached to a pass [otel-error-type]. The caller-heeded half is deliberately NOT in this row -- it is census-only over the call record, because a refusal callers routinely retry through is LOUD-AND-IGNORED no matter how well written it is Adoption: LIVE — fully adopted (top of its ladder). | ● | ◉ | ● | ○ | ○ |
Partial failure declared at the interface, not patched at runtimeOpen — watchingruntime/nx_failclass.nx : fc_partial_iface, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Waldo et al. argue the class is unfixable as implementation quality: partial failure and indeterminacy must be declared at the TYPE level, with cause-reporting, an undetermined-cause recovery path, idempotency and request IDs put IN THE INTERFACE [waldo94]. This row exists to keep an honest boundary on this whole board -- some failure classes are interface-design defects and no runtime detector will ever find them. Related and equally unmeasured here: tool-use failures in agent systems, where the diagnosis is per-call and structural [toolfailbench-2026] | ○ | ● | ● | ○ | ○ |
Person · product · place — not yet measured for this domain
knowledge/compare/failmodes.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain failmodes, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [yuan14] Ding Yuan and Yu Luo and Xin Zhuang and Guilherme Renna Rodrigues and Xu Zhao and Yongle Zhang and Pranay U. Jain and Michael Stumm (University of Toronto), Simple Testing Can Prevent Most Critical Failures - An Analysis of Production Failures in Distributed Data-Intensive Systems, OSDI 2014 publisher · read in our library
knowledge/fetched/cmp_failmodes_yuan14.pdf· pinh2cd4463c773dae22259ed7c112f287b2126b669f75cc6777cc36b8cc8f964e09· accessed 2026-08-20 · published-paperGrounds: Error-handler static rules the Yuan three-rule class -- URL IS THE PUBLISHER-CANONICAL USENIX COPY because that is what was mirrored. The author-hosted eecg.toronto.edu copy is TLS-1.2-only and failed the certificate pipeline twice (rc=-209), and it is the copy whose body was read in full via local extraction of 113675 bytes over 1030 lines - [huang17] Peng Huang and Chuanxiong Guo and Lidong Zhou and Jacob R. Lorch and Yingnong Dang and Murali Chintalapati and Randolph Yao (Microsoft Research and Johns Hopkins), Gray Failure - The Achilles Heel of Cloud-Scale Systems, HotOS 2017 publisher · read in our library
knowledge/fetched/cmp_failmodes_huang17.pdf· pinh65a4c67155481286cb334f5d6e90fb136de62cbe685a0d9adc2bd48eafa1d90a· accessed 2026-08-20 · published-paperGrounds: Gray failure two vantages compared not one -- the first fetch failed transiently under box load (rc=-2) and a Wayback copy was taken, then a retry of this canonical URL succeeded and returned a BYTE-IDENTICAL file at the same 547251 bytes, so the canonical url stands. Body was also read in full, all 6 pages, via local extraction - [waldo94] Jim Waldo and Geoff Wyant and Ann Wollrath and Sam Kendall (Sun Microsystems Laboratories), A Note on Distributed Computing, Sun Labs Technical Report TR-94-29, 1994 publisher · read in our library
knowledge/fetched/cmp_failmodes_waldo94.pdf· pinheaf8aa45d7df9c68db9396a81ebd81a888a3b649d7b78e653304fa7ef2f9b7e8· accessed 2026-08-20 · published-paperGrounds: Partial failure declared at the interface not patched at runtime -- URL IS A WAYBACK CAPTURE OF THE ORIGINAL SUN TECH REPORT because that is what was mirrored. The gatech copy explicitly negotiates TLS-1.2 suite 0xc030 with no TLS 1.3 and failed rc=-209 twice, and a Harvard copy returned an Akamai Access Denied bot wall saved as a 421-byte pdf which was REJECTED not recorded. Body was read in full, all 14 pages, via local extraction - [bronson21] Nathan Bronson and Abutalib Aghayev and Aleksey Charapko and Timothy Zhu, Metastable Failures in Distributed Systems, HotOS 2021 publisher · read in our library
knowledge/fetched/cmp_failmodes_bronson21.pdf· pinh3556cdef28e57967af3dbd5d436bc7c94e9bf032df25ee298c48934a17ee1342· accessed 2026-08-20 · published-paperGrounds: Metastable failure the sustaining effect outlives the trigger - [candea03] George Candea and Armando Fox (Stanford University), Crash-Only Software, HotOS 2003 publisher · read in our library
knowledge/fetched/cmp_failmodes_candea03.pdf· pinh64f17ab7407da0671ab660b7008a6f91924ca5a2b8e0b6af1d61bf55527ba155· accessed 2026-08-20 · published-paperGrounds: Chaos a steady-state hypothesis with a bounded blast radius - [dixit21] Harish Dattatraya Dixit and Sneha Pendharkar and Matt Beadon and Chris Mason and Tejasvi Chakravarthy and Bharath Muthiah and Sriram Sankar (Facebook), Silent Data Corruptions at Scale, arXiv 2102.11245, 2021 publisher · read in our library
knowledge/fetched/cmp_failmodes_dixit21.html· pinhe3f146481cdb526e86b8dc5d619d9c9c7f5a2cc0ad212e00ce97a78c0577ab81· accessed 2026-08-20 · published-paperGrounds: Silent data corruption the fleet screen for a wrong answer - [hochschild21] Peter H. Hochschild and Paul Turner and Jeffrey C. Mogul and Rama Govindaraju and Parthasarathy Ranganathan and David E. Culler and Amin Vahdat (Google), Cores that do not count, HotOS 2021 publisher · read in our library
knowledge/fetched/cmp_failmodes_hochschild21.pdf· pinh482b7318d0ba39c31a7362ecbfdfffc5009a2de2a7a1fae5d055190c4dd86181· accessed 2026-08-20 · published-paperGrounds: Silent data corruption the fleet screen for a wrong answer - [sre-monitoring] Rob Ewaschuk with Betsy Beyer (Google), Monitoring Distributed Systems, chapter 6 of Site Reliability Engineering, O Reilly Media 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-monitoring.html· pinhea5268bca492024a399730734132c7513b4412b10098a207f2e3090eeae1a6fe· accessed 2026-08-20 · published-courseGrounds: Alert quality scored as a classifier - [sre-errorbudget] Marc Alvidrez with Betsy Beyer (Google), Embracing Risk, chapter 3 of Site Reliability Engineering, O Reilly Media 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-errorbudget.html· pinh879d7913327c42bbd9be020f019ed2b61448eddceb1106211ea8bd3eebb6208d· accessed 2026-08-20 · published-courseGrounds: Error-budget burn-rate alerting - [sre-slo] Chris Jones and John Wilkes and Niall Murphy with Cody Smith (Google), Service Level Objectives, chapter 4 of Site Reliability Engineering, O Reilly Media 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-slo.html· pinh449fb54ce65e05102fac46c797a08b86c5ab93ade2c70c63ae25e7648d687d7d· accessed 2026-08-20 · published-courseGrounds: Error-budget burn-rate alerting -- originally REUSED from the devguardrails register to avoid a third fetch, then re-fetched under this domain's own key because the reused mirror predates the fetch-provenance journal and was the ONLY unprovenanced mirror this domain contributed. The reuse was legitimate and the re-fetch proves it: byte-identical at 45445 bytes with the same pin - [sre-burnrate] Steven Thurgood and Jess Frame and Anthony Lenton and Carmela Quinito and Anton Tolchanov and Nejc Trdin with Betsy Beyer (Google), Alerting on SLOs, chapter 5 of The Site Reliability Workbook, O Reilly Media 2018 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-burnrate.html· pinh5a2195ae9280d92489ec6c5be9781ef191d73137514da7c716ebabf5479cfcf5· accessed 2026-08-20 · published-courseGrounds: Alert quality scored as a classifier - [sre-overload] Alejandro Forero Cuervo with Sarah Chavis (Google), Handling Overload, chapter 21 of Site Reliability Engineering, O Reilly Media 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-overload.html· pinh8ca912a82390e7f61e8bbae7baab3a74489f5068d71dee1ff24aed99375e0373· accessed 2026-08-20 · published-courseGrounds: Metastable failure the sustaining effect outlives the trigger - [sre-cascading] Mike Ulrich with Betsy Beyer (Google), Addressing Cascading Failures, chapter 22 of Site Reliability Engineering, O Reilly Media 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_sre-cascading.html· pinhf16f9a582bab016af83c9393e56d7571ba91b49e371bd6b55e7a482c041dddb3· accessed 2026-08-20 · published-courseGrounds: Metastable failure the sustaining effect outlives the trigger - [otel-span-status] OpenTelemetry Authors (Cloud Native Computing Foundation), OpenTelemetry Specification Tracing API - Span Status, SetStatus and StatusCode and RecordException publisher · read in our library
knowledge/fetched/cmp_failmodes_otel-span-status.html· pinh5d8486176fd0bac80447c4c9e470838140de476d054ef114d134930a5781e227· accessed 2026-08-20 · published-standardGrounds: Three-state verdict an axis that cannot see ABSTAINS - [otel-error-type] OpenTelemetry Authors (Cloud Native Computing Foundation), OpenTelemetry Semantic Conventions Attributes Registry - the error namespace publisher · read in our library
knowledge/fetched/cmp_failmodes_otel-error-type.html· pinh223a7fbe9d816055d7067e866286a20657a207f726af4ee6304881d9581503df· accessed 2026-08-20 · published-standardGrounds: Word-bounded matching with UNCLASSIFIED as its own bucket - [cwe390] MITRE, CWE-390 Detection of Error Condition Without Action, Common Weakness Enumeration version 4.20 publisher · read in our library
knowledge/fetched/cmp_failmodes_cwe390.html· pinh833a71fe02802290aaf50cfc8abfd2ef4b62caee0e977020df7ca1a0c7229bb5· accessed 2026-08-20 · published-standardGrounds: Failure taxonomy measured over our own failure record -- the SILENT-BY-DISCARD half of the split - [cwe392] MITRE, CWE-392 Missing Report of Error Condition, Common Weakness Enumeration version 4.20 publisher · read in our library
knowledge/fetched/cmp_failmodes_cwe392.html· pinh02c590a57976a5a14ca447d42175c62671e1318ab99fb4a3ecbd76cb14b3c70d· accessed 2026-08-20 · published-standardGrounds: Failure taxonomy measured over our own failure record -- the SILENT-BY-OMISSION half, and the one CWE in the family with NO Detection Methods section at all - [cwe755] MITRE, CWE-755 Improper Handling of Exceptional Conditions, Common Weakness Enumeration version 4.20 publisher · read in our library
knowledge/fetched/cmp_failmodes_cwe755.html· pinhdc7bfd32dead462fbc8d5c1b37a60fe137d6b3f42c159ec362d091e16e0b2745· accessed 2026-08-20 · published-standardGrounds: Error-handler static rules the Yuan three-rule class -- the parent class of both 390 and 392 - [chaos-principles] Principles of Chaos Engineering, the community statement of the discipline including the advanced principles publisher · read in our library
knowledge/fetched/cmp_failmodes_chaos-principles.html· pinh06bb2df5d5da7442473850d74d911458dba77e0a289b4ebac73c55c4ded1229b· accessed 2026-08-20 · vendor-docGrounds: Chaos a steady-state hypothesis with a bounded blast radius - [chaos-basiri2016] Ali Basiri and Niosha Behnam and Ruud de Rooij and Lorin Hochstein and Luke Kosewski and Justin Reynolds and Casey Rosenthal (Netflix), Chaos Engineering, IEEE Software 2016 publisher · read in our library
knowledge/fetched/cmp_failmodes_chaos-basiri2016.pdf· pinh64a9380e5a5da9b0b04e1b92eaeae90adf746ba671ece57c034f85c02300b73b· accessed 2026-08-20 · published-paperGrounds: Chaos a steady-state hypothesis with a bounded blast radius -- PRECISION NOTE the paper states FOUR principles and Minimize Blast Radius is website-only so citing five to the paper is wrong - [aws-jitter] Marc Brooker (Amazon Web Services), Exponential Backoff And Jitter, AWS Architecture Blog publisher · read in our library
knowledge/fetched/cmp_failmodes_aws-jitter.html· pinh0800c23ad6d06c58b808d32def372f275c04b6f285d7e911c5d7ef16419c1a7d· accessed 2026-08-20 · vendor-docGrounds: Metastable failure the sustaining effect outlives the trigger -- substituted after the AWS Builders Library retries article 301-redirected to a JavaScript shell with no body - [mast-cemri2025] Mert Cemri and colleagues, Why Do Multi-Agent LLM Systems Fail - the MAST taxonomy, arXiv 2503.13657, 2025 publisher · read in our library
knowledge/fetched/cmp_failmodes_mast-cemri2025.html· pinh96f6e43bbd21f6291a5c4e2d8bc9508957413cf11247ee25e4fdc6b2ce633c05· accessed 2026-08-20 · published-paperGrounds: Failure taxonomy measured over our own failure record -- 14 modes in 3 categories whose shares sum to 100.01 so the partition reconciles - [msft-agentic-2026] Microsoft AI Red Team, Updating the taxonomy of failure modes in agentic AI systems - a year of red teaming, Microsoft Security Blog, June 4 2026 publisher · read in our library
knowledge/fetched/cmp_failmodes_msft-agentic-2026.html· pinh89f9b07edac0876fc2b7935920d4affb9192eb4f82e81b868a95fc7300bb5a85· accessed 2026-08-20 · vendor-docGrounds: Agent trajectory anti-patterns over the work record -- body read in full with its seven new modes recorded verbatim, and the mirror independently confirmed by its embedded headline plus datePublished 2026-06-04 plus the Microsoft AI Red Team byline - [false-success-2026] Laksh Advani (University of Colorado), From Confident Closing to Silent Failure - Characterizing False Success in LLM Agents, arXiv 2606.09863, 2026 publisher · read in our library
knowledge/fetched/cmp_failmodes_false-success-2026.html· pinh22ad64d4521f68a53338690c3f2f1cd450091d9b443261d4976ab258eb3700a9· accessed 2026-08-20 · published-paperGrounds: Agent trajectory anti-patterns over the work record -- the measured result that reading a claim cannot detect a false claim - [toolfailbench-2026] Harsh Soni (UC Berkeley), ToolFailBench - Diagnosing Tool-Use Failures in LLM Agents, arXiv 2607.04686, 2026 publisher · read in our library
knowledge/fetched/cmp_failmodes_toolfailbench-2026.html· pinh04ca95ba6b565eef2a3d0034c2d21c92c6f04c9257f108445266c6d4756151a0· accessed 2026-08-20 · published-paperGrounds: Partial failure declared at the interface not patched at runtime
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/failmodes.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers