Nishi FamilyCompare › Autonomous Bug Resolver and Code Developer

Nishi Compare · measured, not asserted

Autonomous Bug Resolver and Code Developer

Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.

The loop that turns the estate's own backlog and the field's benchmarks into verified, shipped fixes -- intake, localise, propose (local no-float tier, ox-alpha through the egress warden, Claude for judged work), reproduce, verify by executable oracle, ship, ledger, publish under a rigor envelope -- measured 2026-08-27 against Claude Code, Codex CLI, OpenHands and Live-SWE-agent and against the August-2026 research record

Layer 1 · Executive

Where we are. MEASURED 2026-08-27 against the catalogue: the self-localising fix loop nx_autofix_auto is BUILT and UNPROMOTED (325,366 B staged, nothing served); its intake gate nx_autofix_intake_gate and its SWE-bench-contract grader nx_swebench_local_gate are PROMOTED and UNREGISTERED; the evolution harness nx_evolve is PROMOTED and UNREGISTERED; the patch-cascade guard, the gated editor, the heal plane and the deep-research runner are registered and dark; the Verified grader nx_swebv_grade is registered and unauthorised. The published scoreboard at /compare/autograde reports 92 percent of our foreign-bug instances resolved under the SWE-bench contract with no denominator on the page, and the July record shows the loop measured 6 of 14 then 7 of 14 with all reverts exact. What IS live and proven: the byte-exact judge and bounded repair loop (coding CB0 and CB1), the mutation harness nx_gate_bite, the promote ruler nx_contentdiff, the ship loop nx_organ_ship with a journal, the failure-record ruler nx_failclass over 880 verdict streams, the egress warden nx_extllm and the external call nx_extllm_call with an egress ledger, and the local no-float Coder tiers. The estate's own boards are the task supply: 4,223 debt rows, 633 roster gates of which 73 read RED or absent or timed out, nine gates that have never gone red, 398 unwired functions, 3,476 inline magic numbers and 271 binaries behind their build -- every one with a machine oracle attached. The counts below are measured at emit.

Where we need to go. OPERATOR ASK (2026-08-27): an autonomous bug resolver and code developer built from what we have and what the August-2026 record shows, using ox-alpha, Claude and the local tools, with /compare as the emitter -- the second project after /compare/gen -- carrying proven evidence on BOTH boards as a mandatory condition, to a standard that would pass a real-world review and doctoral rigor. Deliverable: the loop as organs on the proven lanes, the estate's boards as intake, executable oracles as the only judge, a rigor envelope that refuses any bare number on either board, an external referee on decontaminated multilingual tasks, and a bounded self-improvement path whose proposals never auto-admit.

The unit. 1 u = one measured session-leg (organ plus gate plus live flip), the standing calibration. Local evidence: the trigger-first composer for /compare/gen went from record-mining to deployed in one leg on 2026-08-27; nx_extllm and nx_extllm_call landed in one leg each on 2026-08-25.
Where we are: 20 present rows, 20 open rungs. A loop that exists in pieces, most of them uncallable, graded by a number with no denominator. AHEAD on the oracle: byte-exact judge, mutation harness, promote ruler, failure-record classifier. BEHIND on scope (single function), on scale (1.5B local maker), on external standing (no leaderboard result), and on rigor of publication (no n, no interval, no harness hash).
Cost to an honest loop: 8.5 u. M0: the rigor envelope, the harness manifest, the intake plane, the null controls, the episode ledger, untrusted-input admission and the sandbox root. Version 0.1 is a loop whose every number can be refuted.
Cost to a resolver on the estate backlog: 6.5 u more. M1: reproduction tooth, verify chain, retrieval localiser, sampler tiers and best-of-N. 15 u cumulative = version 1.0, measured by the debt board's own closure count and by recurrence.
Cost to a grade-1 external number: 8 u more. M2: multi-file patch sets, the SWE-rebench V2 and SWE-bench Pro runner, the Aider-Polyglot runner. 23 u cumulative = version 2.0; the first number is whatever it is, published with its interval.
Cost to bounded self-improvement: 9.5 u more. M3: harness evolution as proposals scored on held-out tasks, bounded failure memory, chained maintenance, terms-gated distillation, prevention emission. 32.5 u cumulative = version 3.0.

Research bar. SWE-rebench V2, the post-cutoff multilingual task pool is measured on 32,079 executable real-world tasks across 20 languages from 3,617 repositories with pre-built images, unsound instances filtered by an ensemble of LLM judges [@swerebenchv2]. Theirs: the referee pool the field trains and evaluates on. Ours: AD14 runs a post-cutoff subset bench-only in an isolated host and publishes n, the interval and the harness hash; until then we have no number here and say so.

Research bar. SWE-bench Pro, the held-out long-horizon set is measured on 1,865 problems from 41 repositories in public, held-out and commercial splits; multi-file patches that take a professional hours to days; human-verified and contamination-resistant [@swebenchpro]. Theirs: Live-SWE-agent reports 45.8 percent on it without test-time scaling [@liveswe]. Ours: no number; AD13 multi-file editing is the prerequisite and AD14 the run.

Research bar. ChainSWE, dependent bug chains in one codebase is measured on 304 chronologically chained issues across 54 projects; agents drop by up to 70 percent as chains lengthen [@chainswe]. Theirs: the axis nobody publishes a number on yet. Ours: AD18 publishes resolve rate per chain length over the debt board's own chains, where the same defect was filed five times in three weeks.

Research bar. Holistic Agent Leaderboard, the evaluation infrastructure is measured on 21,730 rollouts across 9 models and 9 benchmarks for about 40,000 dollars, with LLM-aided log inspection that found benchmark-gaming behaviours [@hal]. Theirs: standardised harness, three-dimensional analysis, all logs released. Ours: AD5 ships the episode ledger and AD2 the harness manifest so every run of ours is auditable the same way at zero dollars of inference on the sovereign tier.

26 of 40 capabilities measured|12 of them measured exceeds|14 open|coverage 650/1000|adoption 23 full / 3 partial

Layer 2 · Roadmap

Do this next — computed by the ranker, never chosen by a seat

Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883594 domain=autodev target_version=1.0 rungs=24 done=11 open=13 finish=0 ranker=nx_dr_ocm

#StageRungPriorityDerivation
#11.0Sampler tiers as data (AD11) af_sampler_tier7000v=7 m=2 c=2
#21.0Verify chain with an execution budget (AD9) af_verify_chain4000v=6 m=2 c=3
#31.0Retrieval-backed localiser (AD10) af_localize_index2666v=4 m=2 c=3
#41.0Best-of-N against the free verifier (AD12) fg_best_of_n1333v=4 m=1 c=3
#51.0Reproduction tooth cogenerated with the fix (AD8) af_repro_tooth666v=1 m=2 c=3
#6laterAider-Polyglot runner (AD15) fg_polyglot_bench5333v=16 m=2 c=6
#7laterExternal referee run (AD14) sg_rebench_run2266v=17 m=4 c=30
#8laterMulti-file coordinated patch set (AD13) fg_multi_file1600v=16 m=1 c=10
#9laterBounded failure memory (AD17) fc_failmem1428v=5 m=2 c=7
#10laterChained maintenance (AD18) af_chain_run666v=1 m=2 c=3
#11laterHarness evolution as proposals (AD16) ev_harness_search200v=2 m=3 c=30
#12laterTerms-gated distillation (AD19) xl_distill_corpus200v=1 m=3 c=15
#13laterPrevention emission (AD20) fc_prevent_emit100v=1 m=1 c=10

Critical path — contract, done-rule, executor, cost

RungCloses withDefinition of done (pre-declared)ExecutorEst.
Fix loop v0, self-localising, revert proven exact (AL0)episodeLANDED 2026-07-16 as source and binary; BUILT-UNPROMOTED, which debt row autodev-loop-unpromoted below namesOrgan0 u
SWE-bench contract as the judge (AL1)swb_resolvedLANDED, gate 4 of 4; PROMOTED-UNREGISTEREDGate0 u
Mutation harness, subject rebuilt per mutant (AL2)gb_tryLANDED and LIVE; the oracle-of-the-oracle every rung below leans onOrgan0 u
Egress warden and external call with a ledger row per call (AL3)xl_callLANDED and LIVE 2026-08-25 (extllm XL1 to XL3)Organ0 u
Rigor envelope in the gate base class (AD1)gv_envelopeA published number is a struct -- value, n, a Wilson interval widened by hierarchical bootstrap over the task nesting, seed, harness hash, model pin, run date -- and the compare emit REFUSES a bare percentage for autodev AND gen [@dontpassk] [@statprecipice]. Done rule: the autograde 92-percent row and gen's G5 referee row both publish through it or both read REFUSED; neg-control: a hand-typed percent in api.json is refused by name; the interval width target is a conf row, never a literalGate1.5 u
Harness disclosure manifest (AD2)
after AD1
os_harness_manifestThe ship loop prints and journals a hash over tools, prompts, context builder, verifier, budgets and sampler tier for every result [@harnessthesis]; done rule: two runs with different manifests are refused a direct comparison by the envelope; neg-control: a changed prompt file changes the hashOrgan1 u
Intake plane from every board (AD3)af_intake_boarddebt, gate roster RED, failclass LATCHED, unwired, magic and artifactdrift BEHIND fold into one task plane, each task carrying its oracle and a RED-before receipt; the partition across sources sums to the plane count; chain order preserved [@chainswe]; neg-control: a task whose oracle is GREEN before any fix is refused as unreproducedOrgan1.5 u
Null controls (AD4)
after AD1
af_null_controlEmpty patch, revert patch, verbatim echo and replay of a prior solution run through the same judge on every batch and their scores print beside the resolver's [@statprecipice]; done rule: every null control scores zero on the batch, and a batch where one does not is REFUSED as an oracle defect, never publishedGate0.5 u
Episode ledger (AD5)
after AD1
af_episode_ledgerOne plane row per episode -- task, tier, samples, executions, verdict, tokens, wall, cost -- appended as decided so an interrupted run keeps what it finished; the api.json is emitted from the plane and from nothing else [@hal]Organ1 u
Untrusted-input admission (AD6)af_admit_untrustedDeny by default: an external bug, patch or trajectory enters the tree or any training set only as DATA, provenance-pinned, and only if maintainer-merged upstream; eats debt rows 1784661149 and 1784661346; neg-control: a planted unmerged patch is refused by name; positive control: a pinned merged fix is admitted as data and never executedGate1.5 u
Sandbox root (AD7)af_sandbox_rootEvery candidate builds and runs in an isolated scratch root under a time and resource limit and never touches the serving root [@openhands] [@dgm]; done rule: a candidate that writes outside its root is killed and the write is absent; neg-control: a fork bomb candidate is boundedOrgan1 u
Reproduction tooth cogenerated with the fix (AD8)
after AD3
af_repro_toothA fix ships with a gate row that fails on the pre-fix binary and passes on the post-fix one, and the planted mutant must die under nx_gate_bite [@brtcogen]; done rule: fix rate with cogeneration is not lower than without, with non-overlapping intervals or a declared tieOrgan1.5 u
Verify chain with an execution budget (AD9)
after AD7
af_verify_chainbuild, gate, bite, contentdiff, behaveprobe, magic and unwired ratchets as one pipeline; the execution budget per task class is a conf row derived from measured benefit [@runornot]; done rule: a fix that adds a magic literal or an unwired function is REFUSED; neg-control: a fix that loses a live run is refused by the promote rulerOrgan1.5 u
Retrieval-backed localiser (AD10)
after AD3
af_localize_indexcode index plus failclass plus gate stderr rank candidate functions; hit at 1 measured on a held-out split of the intake plane and published with its interval [@adi]; done rule: hit at 1 exceeds the FNRES-only baseline with non-overlapping intervals or the leg stays UNWIREDOrgan1.5 u
Sampler tiers as data (AD11)
after AD6
af_sampler_tierlocal 1.5B, ox-alpha through the warden, Claude for judged work, chosen per task class by a conf row; the free tier sees U bytes only [@stealtheula] and the Claude tier's output never enters the corpus; done rule: the same task run on each tier lands in the same ledger with the tier namedOrgan1 u
Best-of-N against the free verifier (AD12)
after AD11
fg_best_of_nshared with coding CD2 and extllm XL5: pass at N monotone in N on the intake plane, NONE reported when nothing is GREEN [@ttcagentic] [@swezerohero]; AF_BON_N retires as a literalOrgan1 u
Multi-file coordinated patch set (AD13)
after AD9
fg_multi_fileshared with coding CD5: a task may name several organs, the judge builds the whole closure, a patch that breaks a third importer is refused [@swebenchpro]Organ3 u
External referee run (AD14)
after AD1,AD2,AD7
sg_rebench_runa post-cutoff subset of SWE-rebench V2 and the SWE-bench Pro public split run bench-only in an isolated sandbox host under the untrusted-input rule, graded by our grader, published with n, interval and manifest hash [@swerebenchv2] [@swebenchpro] [@swebenchcase]; done rule: the run publishes or it prints REFUSED with the missing conjunct named; a Verified-only run is a sanity check and never the headlineGate3 u
Aider-Polyglot runner (AD15)
after AD1
fg_polyglot_benchshared with coding CD6; publishes pass rate per language with N and the manifest hashOrgan2 u
Harness evolution as proposals (AD16)
after AD1,AD5,AD14
ev_harness_searchharness parameters are conf rows; the evolution harness proposes edits scored on a HELD-OUT split against a matched-budget test-time-scaling baseline [@metaharness] [@ahe] [@harnessevalrethink] [@lilharness]; a proposal lands on the pm intake plane and never auto-admits; done rule: a proposal is adopted only when it beats the baseline on held-out tasks with non-overlapping intervals; neg-control: a proposal that wins on the search split and loses on held-out is refused and the refusal is journaledOrgan3 u
Bounded failure memory (AD17)
after AD5
fc_failmema per-class failure library with an eviction bound as a conf row, read by the localiser and the prevention emitter [@liveswe]; done rule: the store never exceeds its bound and eviction is journaled; neg-control: a store over bound refuses the writeOrgan1.5 u
Chained maintenance (AD18)
after AD3,AD9
af_chain_rundependent tasks on one organ family without a reset, resolve rate published per chain length with intervals [@chainswe]; done rule: the drop with chain length is a published curve, not a hidden averageOrgan1.5 u
Terms-gated distillation (AD19)
after AD5,AD6
xl_distill_corpusshared with extllm XL9: verified episodes enter the sovereign corpus only from providers whose terms permit it, never from ox-alpha or Claude, and with every external sampler disabled the corpus must still strictly grow over a run [@swezerohero]Organ1.5 u
Prevention emission (AD20)
after AD17
fc_prevent_emitevery resolved class emits a new gate or lint; done rule: recurrence rate per class on the debt board falls after the emitter lands, measured over the whole board and never a sample; neg-control: a class with no emitted oracle shows no fallOrgan2 u

Milestones

MilestoneRungsCumulative
M0 · Honest loopAD1,AD2,AD3,AD4,AD5,AD6,AD78.5 u
M1 · Resolver on the estate backlogAD8,AD9,AD10,AD11,AD1215 u
M2 · External refereeAD13,AD14,AD1523 u
M3 · Bounded self-improvementAD16,AD17,AD18,AD19,AD2032.5 u
Layer 3 · Engineering
How this is scored. Every Nishi mark is measured: the generator reads the real organ source on disk and requires the implementing symbol to exist (no self-grading). A watching tag names the organ and symbol contracted to close a gap — the mark flips itself on the next compare beat when that workstream ships, and the comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).

Capability matrix — measured against source

leads / measured exceed present partial absent · click any capability for its evidence

CapabilityNishiClaude CodeCodex CLIOpenHandsLive-SWE-agent
INTAKE
the estate's own backlog as executable tasks -- the debt board (4,223 rows, 945 at sev 7 or above)Measured: db_open_count exists in runtime/nx_debt.nx, verified at emit. The rivals take tasks from GitHub issues and chat: Claude Code routes bug reports from Slack to pull requests and triages issues in CI [claudecode], Live-SWE-agent consumes SWE-bench instances [liveswe]; ours reads a plane whose every row already carries a severity and a scope, and the same board is where a resolved task closes Adoption: LIVE — fully adopted (top of its ladder).
foreign-bug instances admitted only if the bug REPRODUCES on a fresh build (equivalent-mutant discard)Measured exceed: has_failing in runtime/nx_autofix_intake_gate.nx, verified at emit. SWE-bench Verified instances were human-validated [swebench] and SWE-rebench V2 filters unsound instances with an ensemble of LLM judges [swerebenchv2]; no rival HARNESS validates the task it is handed. Ours refuses any instance with no failing per-function row, so a task that cannot fail never enters the resolve rate. PROMOTED and UNREGISTERED today Adoption: GATE:LIVE trial=RED — fully adopted (top of its ladder).
LOCALIZE
the failing function found from the grader's own per-function rows, never toldMeasured: fn_failing exists in runtime/nx_autofix_auto.nx, verified at emit. The agent-computer-interface lineage localises by navigating the repository and running tests [sweagent]; the ADI adds a function-level Frame Lifetime Trace and lifts existing agents 6.2 to 18.5 percent [adi]. Ours reads FNRES rows from a fresh grader run -- exact, cheap and single-function, which is the whole of its scope today Adoption: LIVE — fully adopted (top of its ladder).
PROPOSE
the fix-loop episode -- localise, generate, apply, fresh compile plus run as the sole judge, revert proven exactMeasured: episode exists in runtime/nx_autofix_auto.nx, verified at emit. Every rival column resolves repository-level issues [openhands] [liveswe]; ours is one function at a time with a 1.5B local maker and the binary is BUILT-UNPROMOTED. The mechanism is the same shape as SWE-agent's generate-run-revise loop [sweagent] with a deterministic judge in place of a test suite Adoption: LIVE — fully adopted (top of its ladder).
sampled retries after both greedy rounds miss (three samples at temperature 0.8, a deterministic seed ladder)Measured: lf_gen_t exists in runtime/nx_autofix_auto.nx, verified at emit. Test-time compute is where the cheap gains are: tournament voting and distill-refine lifted Claude-4.5-Opus from 70.9 to 77.6 on Verified [ttcagentic], and 32 rollouts with a generative verifier gained up to 7.9 points [swezerohero]. Ours samples exactly three, a literal in source that this board names as debt, and selects only by the judge Adoption: LIVE — fully adopted (top of its ladder).
VERIFY
the SWE-bench contract as the sole judge -- FAIL_TO_PASS flips AND PASS_TO_PASS holdsMeasured: swb_resolved exists in runtime/nx_swebench_local_gate.nx, verified at emit. The contract is the benchmark's own [swebench]; the gate proves the cheat-proof tooth that a regressive fix is NOT resolved. Rivals are graded by the same contract on the Python set; ours grades a sovereign NishiLang analog and says so on the page. PROMOTED and UNREGISTERED today Adoption: GATE:LIVE trial=SKIP — fully adopted (top of its ladder).
byte-exact deterministic build plus judge -- the verify signal is noise-freeMeasured exceed: fg_run in runtime/nx_forge.nx, verified at emit. Every rival judges by a test suite, which can flake and which a replay script can beat when verification is not privileged [statprecipice]; a GREEN here is a provable equality. This is the estate's largest structural advantage and every open rung below leans on it Adoption: LIB-WIRED importers=8 nonval=5 — fully adopted (top of its ladder).
non-vacuity by mutation -- can this gate ever fail? the subject is rebuilt and deployed per mutant and restored by hashMeasured exceed: gb_try in runtime/_hdl_build/nx_gate_bite.nx, verified at emit. The Darwin Godel Machine validates each self-modification empirically on a benchmark [dgm] but never asks whether the benchmark can fail; no rival harness mutates its own oracle. Ours does, and the failure-record ruler below found nine gates that have never once gone red Adoption: LIVE — fully adopted (top of its ladder).
the promote ruler -- a capability-loss diff naming every lost run before a candidate replaces the live binaryMeasured exceed: cd_show_lost in runtime/nx_contentdiff.nx, verified at emit. Regression prevention is a named open challenge of the harness field [codeharness]; ours refuses a promote that loses a run and prints the run. Its known blindness -- a code-only delta with no literal loss -- is why the verify chain below adds behaviour and size legs beside it Adoption: LIVE — fully adopted (top of its ladder).
SHIP
the ship loop as one organ -- build, prove by gate, contentdiff, behaveprobe, promote, journal with the shaMeasured: os_jrnl exists in runtime/nx_organ_ship.nx, verified at emit. Claude Code stages changes, writes commits and opens pull requests [claudecode]; Codex CLI runs locally on the developer's machine [codex]; OpenHands ships patches from a sandbox [openhands]; Live-SWE-agent emits a patch per instance [liveswe]. Ours ships to a serving root with a banked rollback and journals every stage; the rivals' git-native flow is the better product surface and the reason the cell reads Best for them Adoption: LIVE — fully adopted (top of its ladder).
GUARD
patch-cascade tooth -- a second failed fix refuses a third (rule 3 mechanised, the maximum a conf row)Measured exceed: PG_DEFMAX in runtime/_hdl_build/nx_patch_guard.nx, verified at emit. Agents average 8.8 test runs per task and range from 2 to 19 [runornot]; none of the four refuses a cascade by construction. Registered and dark today -- the tooth exists and the loop does not call it yet Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-DARK
the egress warden -- terms ledger plus classify-at-the-door before any byte reaches an external modelMeasured exceed: xl_egress_class in runtime/nx_extllm.nx, verified at emit. The rival products send the repository to the model by design [claudecode] [codex]. Ours refuses any prompt byte not on the operator's public-release manifest, because the ox-alpha page says prompts are retained and not used for training [oxalpha] while the contract that binds it says they are collected for training [stealtheula] -- the contract governs, and the ledger row carries its pin Adoption: LIVE — fully adopted (top of its ladder).
SAMPLER
an external frontier sampler behind the warden -- ox-alpha, 1,048,576-token context, freeMeasured: xl_call exists in runtime/nx_extllm_call.nx, verified at emit. The model page describes a reasoning model for coding and sustained agentic work [oxalpha]; the call organ ledgers provider, model, prompt hash, bytes, class, terms pin and verdict per call and attributes every 429 to its source. Capability facts only on this page, by the EULA [stealtheula] Adoption: LIVE — fully adopted (top of its ladder).
the local sovereign no-float Coder tier (0.5B and 1.5B) loaded bits-upOpen — no implementing organ is measured for this axis yet. No rival runs a model through its own integer stack; this is the sovereign bottom rung of the sampler ladder and the tier the July resolve rate was measured on. Scale is the capability lever the record already measured
CONTEXT
verified context pack emitted from sources, regenerableMeasured: fc_emit exists in runtime/nx_forge_ctx.nx, verified at emit. A harness is the code that decides what to store, retrieve and present to the model [metaharness] and its component list is now standard across the field [lilharness]; Claude Code reads CLAUDE.md and builds auto memory [claudecode]. Ours is vetted from sources with every rule witnessed and auto-retired when the compiler heals Adoption: LIB-WIRED importers=7 nonval=4 — fully adopted (top of its ladder).
sovereign semantic code retrieval (jina recipe, integer embedding path)Measured: ci_query_top1 exists in runtime/nx_code_index.nx, verified at emit. Base-weight quality, 3 of 6 top-1, a ratchet; the coding domain owns the contrastive tune that closes it. Here it is the retrieval leg of the localiser Adoption: LIB-WIRED importers=4 nonval=2 — fully adopted (top of its ladder).
FAILURE RECORD: 880 verdict streams classified against a published taxonomy; nine gates that have fired GREEN on every run of their history named as a bite worklistMeasured exceed: fc_alert_score in runtime/nx_failclass.nx, verified at emit. HAL inspected 2.5B tokens of agent logs with an LLM to find behaviours nobody reported, including agents searching for the benchmark instead of solving it [hal]; ours classifies the whole record deterministically and prints a partition that sums. A perfect history is the absence of evidence, and this ruler is how the loop chooses which oracles to bite first Adoption: LIVE — fully adopted (top of its ladder).
EVOLVE
the evolution harness with an external fitness (a gate permil) and a known-optimum controlMeasured: ev_fit_ext exists in runtime/_hdl_build/nx_evolve.nx, verified at emit. AlphaEvolve's loop is an evaluator feeding an evolutionary search [alphaevolve]; the Darwin Godel Machine keeps an archive of agents validated on benchmarks [dgm]; Live-SWE-agent rewrites its scaffold on failure [liveswe]. Ours validates the machinery on a fitness with a known optimum before trusting it on a gate, and it is PROMOTED and UNREGISTERED Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-DARK
GRADE
the SWE-bench Verified grader owning its denominator -- python is the oracle, never the scorerMeasured: sg_verdicts exists in runtime/_hdl_build/nx_swebv_grade.nx, verified at emit. The Verified leaderboard's numbers are computed by the benchmark harness and dominated by industry entries on proprietary models [swebenchcase]; ours computes the number itself over a results plane. Registered and unauthorised today Adoption: REGISTERED-DARK — PARTIAL: callable, authorised, no MCP invocation on record (a direct fork logs the runner, so this is not proof it never ran); no execution surface runs it either (clock, cron, daemon, roster, actlog and surfaced forks checked).
adoption REGISTERED-DARK
HEAL
the heal plane -- service status and the remediation log the loop closes againstMeasured: nh_metrics exists in runtime/nx_heal.nx, verified at emit. No rival has a notion of the estate it runs in; a resolved runtime defect here is one whose heal-plane row goes quiet. Registered and dark today Adoption: LIVE — fully adopted (top of its ladder).
RIGOR
every published number carries n, a Wilson interval with hierarchical bootstrap over the nested task structure, seed, harness hash and model pin -- and the emit REFUSES a bare percentageMeasured exceed: gv_envelope in runtime/nx_gate_verdict.nx, verified at emit. Posterior means with credible intervals and a declare-a-difference-only-on-non-overlap rule replace pass at k [dontpassk]; naive aggregation over correlated tasks under-covers badly and Wilson plus hierarchical bootstrap corrects it [statprecipice]. Today our autograde page prints 92 percent with no n. The envelope lives in the gate base class so that /compare/gen's referee run (G5) and this board publish under ONE ruler Adoption: LIB-WIRED importers=1506 nonval=108 — fully adopted (top of its ladder).
harness disclosure manifest -- tools, prompts, context builder, verifier, budgets and sampler tier, hashed and printed beside every resultMeasured exceed: os_harness_manifest in runtime/nx_organ_ship.nx, verified at emit. Harness configuration governs more variance than model choice and undisclosed harnesses make leaderboards misleading [harnessthesis]; OpenHands and Live-SWE-agent publish their code [openhands] [liveswe] but no per-result manifest, and the two products publish neither. Ours prints the hash the result was produced under, so two numbers are comparable only when their manifests match Adoption: LIVE — fully adopted (top of its ladder).
INTAKE
one task plane from every board -- debt, gate roster RED, failclass LATCHED, unwired, magic, artifactdrift BEHIND -- each task carrying its executable oracle and a RED-before receipt, partition summingMeasured exceed: af_intake_board in runtime/nx_autofix_auto.nx, verified at emit. The estate's boards hold thousands of tasks with machine oracles already attached, which no rival's intake can claim. ChainSWE shows that a stream of related defects in one codebase is the hard and realistic setting, with up to a 70 percent drop as chains lengthen [chainswe]; the plane keeps the chain order Adoption: LIVE — fully adopted (top of its ladder).
LOCALIZE
retrieval-backed localisation -- code index plus failclass plus gate stderr -- with hit at 1 measured on held-out tasksOpen — watching runtime/nx_autofix_auto.nx : af_localize_index, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Dynamic function-level evidence beside static retrieval is what the ADI measured as the cheapest lift [adi]; the contract publishes hit at 1 with its interval, never a bare rate
watching af_localize_index
SAMPLER
sampler tiers as DATA -- local 1.5B, ox-alpha through the warden, Claude for judged work -- selected per task class by a conf row, never by codeOpen — watching runtime/nx_autofix_auto.nx : af_sampler_tier, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OpenHands runs many model backends behind one agent [openhands]; ours adds the clearance: the free tier sees only U bytes [stealtheula], the local tier sees everything, and the Claude tier is the escalation whose output may never grow the sovereign corpus
watching af_sampler_tier
REPRO
a bug-reproduction tooth cogenerated with every fix -- a neg-control that fails before and passes after, and the planted mutant must dieOpen — watching runtime/nx_autofix_auto.nx : af_repro_tooth, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Google's study of 120 bugs found cogenerating the reproduction test in the same patch costs no fix rate and beats a separate test agent [brtcogen]. Ours makes the tooth a gate row, so the loop leaves prevention behind it every time it fixes something
watching af_repro_tooth
VERIFY
the verify chain as one pipeline -- build, gate, bite, contentdiff, behaveprobe, magic and unwired ratchets -- with an execution budget measured per task classOpen — watching runtime/nx_autofix_auto.nx : af_verify_chain, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prohibiting execution cost only 1.25 points of resolve rate across 3,000 attempts and the benefit concentrated on a subset of instances [runornot]; the budget is a conf row derived from the measured benefit per class, and the ratchets make a fix that adds a magic number or an unwired function REFUSE
watching af_verify_chain
PROPOSE
best-of-N against the free verifierOpen — watching runtime/nx_forge.nx : fg_best_of_n, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The same contract coding names as CD2 and extllm as XL5, so three boards flip on one organ; pass at N monotone in N and NONE reported when nothing is GREEN [ttcagentic] [swezerohero]
watching fg_best_of_n
multi-file coordinated patch set, the judge builds the whole closureOpen — watching runtime/nx_forge.nx : fg_multi_file, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. SWE-bench Pro tasks take a professional hours to days and patch multiple files [swebenchpro]; the same contract coding names as CD5 -- our loop is single-function until it lands
watching fg_multi_file
CONTROL
null resolvers -- empty patch, revert patch, verbatim echo, replay of a prior solution -- must score zero, printed beside every runMeasured exceed: af_null_control in runtime/nx_autofix_auto.nx, verified at emit. A replay script that never observes the screen beat frontier models on static benchmarks, with expected success exactly the source agent's pass at k [statprecipice]; the echo class already cost this loop three misses in July, so the null controls are the board's own history made mechanical Adoption: LIVE — fully adopted (top of its ladder).
GUARD
untrusted-input admission -- no external bug, patch or trajectory enters the tree or a training set unless it is DATA-ONLY, provenance-pinned and maintainer-merged; deny by defaultMeasured exceed: af_admit_untrusted in runtime/nx_autofix_auto.nx, verified at emit. Two sev-8 debt rows (1784661149, 1784661346) demand this before any ingestion or training. HAL found agents that searched HuggingFace for the benchmark rather than solving it [hal]; a poisoned bug-fix pair is the same class from the other side. Neg-control: a planted unmerged patch must be refused by name Adoption: LIVE — fully adopted (top of its ladder).
SANDBOX
every candidate builds and runs in an isolated scratch root with time and resource limits, never the serving rootMeasured: af_sandbox_root exists in runtime/nx_autofix_auto.nx, verified at emit. OpenHands executes in sandboxed environments by design [openhands]; the DGM ran every self-modification under sandboxing and human oversight [dgm]; the Codex README as mirrored says nothing about sandboxing so its cell is coded from silence [codex]. Ours today edits in place with a banked copy and a proven revert, which is a rollback and not an isolation Adoption: LIVE — fully adopted (top of its ladder).
LEDGER
one row per episode -- task, tier, samples, executions, verdict, tokens, wall, cost -- and the page reads the plane, never a hand numberMeasured: af_episode_ledger exists in runtime/nx_autofix_auto.nx, verified at emit. HAL released 2.5B tokens of logs so that behaviour could be audited [hal]; ours writes the row the rigor envelope aggregates and the cost the engine-shift gauge reads Adoption: LIVE — fully adopted (top of its ladder).
REFEREE
SWE-rebench V2 and SWE-bench Pro post-cutoff multilingual subset, run bench-only in an isolated sandbox host, graded by our graderOpen — watching runtime/_hdl_build/nx_swebv_grade.nx : sg_rebench_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. 32,079 executable tasks across 20 languages with pre-built images [swerebenchv2] and a held-out long-horizon set [swebenchpro] are the referees; Verified is saturated and industry-skewed [swebenchcase], so it stays a sanity check. The run reports n, the interval and the harness hash or it does not publish
watching sg_rebench_run
the Aider-Polyglot runnerOpen — watching runtime/nx_forge.nx : fg_polyglot_bench, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The same contract coding names as CD6; the first number is whatever it is
watching fg_polyglot_bench
EVOLVE
harness evolution as PROPOSALS -- harness parameters are conf rows, the loop proposes edits scored on a HELD-OUT split against a matched-budget test-time-scaling baseline, and nothing auto-admitsOpen — watching runtime/_hdl_build/nx_evolve.nx : ev_harness_search, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Meta-Harness and AHE automate the search [metaharness] [ahe] and the essay names it the near-term path to self-improvement [lilharness]; the July-2026 re-evaluation found evolved harnesses rarely beat simple test-time scaling at matched budget and generalise poorly [harnessevalrethink]. That finding is the accept rule: a proposal ships only if it beats the matched-budget baseline on held-out tasks with non-overlapping intervals
watching ev_harness_search
MEMORY
bounded failure memory -- a per-class failure library with eviction, feeding prevention rather than promptsOpen — watching runtime/nx_failclass.nx : fc_failmem, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Claude Code builds auto memory as it works [claudecode]; Live-SWE-agent keeps its original context to prevent drift while evolving [liveswe]. Ours bounds growth by construction because a store that outgrows its reader fails silently and tail-first
watching fc_failmem
CHAIN
sequential dependent tasks on one organ family without a reset, resolve rate published per chain lengthOpen — watching runtime/nx_autofix_auto.nx : af_chain_run, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. ChainSWE is the first benchmark of this setting and reports the drop with chain length [chainswe]; our debt board already holds chains (the same defect filed five times in three weeks), so the number is ours to publish first
watching af_chain_run
DISTILL
terms-gated harvest of verified episodes into the sovereign corpus -- never from ox-alpha or ClaudeOpen — watching runtime/nx_extllm.nx : xl_distill_corpus, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The same contract extllm names as XL9. SWE-HERO's recipe used 300,000 execution-free and 13,000 execution-based trajectories [swezerohero] and SWE-rebench V2 exists to feed RL [swerebenchv2]; our corpus grows only from episodes the byte-exact judge passed and only from providers whose terms permit it, and with every external sampler disabled it must still strictly grow
watching xl_distill_corpus
PREVENT
every resolved class emits a new gate or lint, measured as recurrence-rate reduction on the debt boardOpen — watching runtime/nx_failclass.nx : fc_prevent_emit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Prevention outranks self-healing outranks post-analysis, and post-analysis exists to build new prevention; no rival column emits an oracle after a fix. The measurement is the board's own recurrence count per defect class before and after
watching fc_prevent_emit

Risk register

Debt register

RiskLikelihood x impactMitigation
A harness that wins on the split it searched and loses on held-out tasks reads as progressprobable x highAD16 pre-declares the held-out matched-budget accept rule from the July-2026 re-evaluation [@harnessevalrethink] and journals every refusal
A stub carrying a contract symbol flips a watch cell green with no capability behind itpossible x highthe flip is the receipt, the rung gate is the proof; every AD rung names its neg-control
The resolver games its own oracle -- edits a gate to pass, or a fix that only satisfies the toothpossible x highAD8 requires the planted mutant to die under nx_gate_bite, AD9 refuses gate edits outside the fix's declared scope, AD4 keeps null controls on every batch [@statprecipice]
A free tier trains on what it is shown and a probe number leaks to a public pageprobable x highthe warden classifies every byte at the door and the publish guard keeps ox-alpha numbers at class C [@stealtheula]; this page carries capability facts only
A poisoned external bug-fix pair enters the corpus or the treepossible x highAD6 deny-by-default admission, data-only, provenance-pinned, maintainer-merged; neg-control planted and bite-proven before any ingestion
The number keeps changing and the reader learns to distrust the instrumentprobable x mediumAD1 publishes the method with every count and AD2 the manifest hash, so a moved number names what moved
An interrupted run yields nothingprobable x mediumAD5 appends each episode as decided; a run interrupted at 95 percent keeps 95 percent
IdSevWhat it isUnblock
autodev-loop-unpromoted7nx_autofix_auto is BUILT and UNPROMOTED (325,366 B staged) and nx_autofix_intake_gate, nx_swebench_local_gate and nx_evolve are PROMOTED and UNREGISTERED; the flagship loop of this board cannot be called from the MCP surface, which is the built-and-unwired class the operator named the deliverable.Promote through nx_organ_ship with the gate named in organ_gate.conf, register the three via /api/tools/register, mint a cap, invoke once, and re-read nx_catalog until all four read LIVE.
autograde-no-denominator6/compare/autograde/api.json publishes 92 percent resolved with no n, no interval and no run date beyond generated_unix; the July record shows n was 14, and a percentage without a denominator is the shape every rigor law on this estate forbids.AD1: the row is emitted from the episode ledger through gv_envelope or it reads REFUSED; until then the page says sovereign analog and nothing about confidence.
autofix-bon-literal4AF_BON_N, AF_BON_TEMP_PM, AF_BON_TOPP_PM and AF_BON_TOPK are literals in nx_autofix_auto.nx; the sampler shape is code, not data, so the best-of-N rung cannot be tuned without a rebuild.AD12 reads them from a conf row; the literals retire with a byte-identical-behaviour proof on the July manifest.
autodev-gates-unrunnable-from-root6MEASURED 2026-08-27 by nx_swcompare_evidence over autodev.gates: nx_swebench_local_gate exits 3 SKIP pass 0 of 0 from the serving root because the fix-loop ledger it reads lives under a home directory (nx_stage/autofix_ledger.log), and nx_autofix_intake_gate exits 1 RED pass 0 of 0 with no proposal rows to admit -- so the two promoted gates of the loop cannot serve as executed evidence for this page today and the evidence stamp reads RED, which this board publishes rather than hides.AD5 moves the ledger to a plane under knowledge/ that the gate reads by an estate-relative path, and AD3 gives the intake gate a standing proposal plane; both gates then run GREEN from the root and get bitten (C4 never-bitten today), and only then does the stamp turn.
autodev-dark-guards5nx_patch_guard, nx_gated_edit, nx_editjudge, nx_heal, nx_triage and nx_dr_run are registered and dark; the loop has guards it does not call.AD9 composes the patch guard and the gated editor into the verify chain; invocation is then measured by the actlog rather than asserted.
On these two registers. Rows are declared in the domain's plan file and carry the debt id, which is the join key back to the sovereign debt plane — that plane, not this page, is the authority on state. Reconciling them automatically (the regen reading the plane and refreshing these rows) is a named, owed rung; until it lands, treat an id here as a pointer to look up, not a status to trust.
Honest verdict. Honest position, measured 2026-08-27. The pieces of an autonomous resolver already exist here and most of them are dark. The fix loop (nx_autofix_auto: self-localising from the grader's own per-function rows, a local 1.5B maker, fresh compile plus run as the sole judge, revert proven exact, sampled retries after greedy misses) is BUILT and UNPROMOTED; its intake gate and its SWE-bench-contract grader are PROMOTED and UNREGISTERED; the evolution harness, the patch-cascade guard, the gated editor and the heal plane are registered and dark. The published scoreboard says 92 percent of our foreign-bug instances resolve under the SWE-bench contract [@swebench] and prints no denominator -- the first thing this board's rigor rung refuses is our own page. What the estate genuinely has that the field does not is the ORACLE: a deterministic byte-exact judge so the verify signal is noise-free, 2,317 gate sources that are executable specifications, a mutation harness that asks whether a gate can fail at all, and a promote ruler that names every lost run before a binary goes live -- regression prevention, which the harness survey lists as an open challenge [@codeharness]. What the field has that we do not is repository scale and a public number: the rival columns resolve real multi-file issues [@sweagent] [@openhands] [@liveswe] while ours is single-function scope with no external-leaderboard result. The August-2026 record sets the design: harness configuration explains more variance than model choice and must be disclosed [@harnessthesis]; test-time compute buys the largest cheap gains when a verifier exists [@ttcagentic] [@swezerohero]; execution should be budgeted, not defaulted [@runornot]; the reproduction test belongs in the patch [@brtcogen]; harness self-evolution must be scored on held-out tasks against a matched-budget baseline or it is overfitting wearing a win [@harnessevalrethink] [@metaharness] [@ahe]; Verified is saturated and sociologically skewed [@swebenchcase], so the referee is SWE-rebench V2's post-cutoff multilingual pool [@swerebenchv2] and SWE-bench Pro's held-out set [@swebenchpro], with chained maintenance as the honest next axis [@chainswe]; and every number carries an interval that respects the nested structure of agent benchmarks [@dontpassk] [@statprecipice] and ships beside a null control that a replay script cannot beat. The free frontier sampler is ox-alpha with a 1,048,576-token context [@oxalpha], reached only through the egress warden because the contract that binds it collects prompts for training [@stealtheula]. Nothing external is trusted: the debt board carries two sev-8 rows demanding deny-by-default admission of foreign bugs and fixes before any ingestion or training, and this board makes that a rung with a neg-control rather than a promise.

Person · product · place — not yet measured for this domain

Every compare carries this layer. Declare knowledge/compare/autodev.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain autodev, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).

References

Beyond a link list. Every reference below resolves twice — the publisher's copy and, where banked, the estate's own non-rottable library mirror with a content pin — and carries its evidence class plus the exact claim on this page it grounds. Keyed marks like [key] in the matrix notes jump here. A dash means honestly absent, never assumed.
  1. [harnessthesis] Zhang, Wang, Ge, Xu, Hamm and Reddy. Stop Comparing LLM Agents Without Disclosing the Harness (position paper, submitted 2026-05-07). The Binding Constraint Thesis: for long-horizon tasks across models of comparable frontier capability, harness configuration governs more performance variance than model choice, with measured cases of model ranking reversal; proposes a harness disclosure standard and a variance decomposition protocol. arXiv 2605.23950. publisher · read in our library knowledge/fetched/cmp_autodev_harnessthesis.html · pin he6f975369dc5da48550dbf83894d8baad17eedf57200a9ee5481094df989aad3 · accessed 2026-08-27 · published-paperGrounds: RIGOR: harness disclosure manifest hashed beside every result
  2. [metaharness] Lee, Nair, Zhang, Lee, Khattab and Finn. Meta-Harness: End-to-End Optimization of Model Harnesses (submitted 2026-03-30). The harness is the code that determines what information to store, retrieve and present to the model; an agentic proposer with filesystem access to prior candidates automates harness optimisation and surpasses hand-engineered baselines on TerminalBench-2. arXiv 2603.28052. publisher · read in our library knowledge/fetched/cmp_autodev_metaharness.html · pin he25f14e1e507216df47702e5410dc3d208b8e0d0d3dc4a875c2a2911f7d280c1 · accessed 2026-08-27 · published-paperGrounds: EVOLVE: harness evolution as proposals scored on a held-out split
  3. [ahe] Lin et al. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses (submitted 2026-04-28, revised 2026-05-18). Closed-loop harness evolution on three observability pillars; Terminal-Bench 2 pass at 1 lifted from 69.7 to 77.0 over ten iterations, and on SWE-bench Verified top performance with 12 percent fewer tokens. arXiv 2604.25850. publisher · read in our library knowledge/fetched/cmp_autodev_ahe.html · pin h62fcbc1aebd90defa93c346fa1e0d3b518e569c20845e3116a945cb800c53d3a · accessed 2026-08-27 · published-paperGrounds: EVOLVE: harness evolution as proposals scored on a held-out split
  4. [harnessevalrethink] Wang et al. Rethinking the Evaluation of Harness Evolution for Agents (submitted 2026-07-14). Two flaws named in the field: harness search is not compared against matched-budget simpler baselines, and it is evaluated on the same benchmark it searched on; measured on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, automatic harness evolution does not consistently outperform simple test-time scaling and generalises poorly to held-out tasks. arXiv 2607.12227. publisher · read in our library knowledge/fetched/cmp_autodev_harnessevalrethink.html · pin h85f0b5cf005a157cb4f7b866ff58d7a445fe22cc73dffb738fb6b5ddde3660ac · accessed 2026-08-27 · published-paperGrounds: EVOLVE: the matched-budget and held-out accept rule this board pre-declares
  5. [liveswe] Xia, Wang, Yang, Wei and Zhang. Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? (submitted 2025-11-17). An agent that starts with bash tools and modifies its own scaffold implementation while solving problems; 77.4 percent on SWE-bench Verified without test-time scaling and 45.8 percent on SWE-Bench Pro. arXiv 2511.13646. publisher · read in our library knowledge/fetched/cmp_autodev_liveswe.html · pin h478579f688f859d81f4022a55b971b5846aa6992631ca0055254b09647fc4af3 · accessed 2026-08-27 · published-paperGrounds: The Live-SWE-agent column and the self-evolving-scaffold rung
  6. [lilharness] Weng. Harness Engineering for Self-Improving Coding Agents (blog essay, 2026-07-04). A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans; near-term recursive self-improvement proceeds by evolving the harness layer through automated search rather than by rewriting weights; names Meta-Harness, Self-Harness, AHE, AlphaEvolve and ADAS as the systems doing it. publisher · read in our library knowledge/fetched/cmp_autodev_lilharness.html · pin h45bc3793c36e42f79749f84af510761798b85960ad18d3123d86f3a403e247ac · accessed 2026-08-27 · published-courseGrounds: CONTEXT: the harness component list this board measures itself against
  7. [ttcagentic] Kim et al. Scaling Test-Time Compute for Agentic Coding (submitted 2026-04-16). Rollouts converted to structured summaries so that Recursive Tournament Voting (parallel) and Parallel-Distill-Refine (sequential) apply to long-horizon agents; Claude-4.5-Opus from 70.9 to 77.6 percent on SWE-Bench Verified and Terminus 1 from 46.9 to 59.1 on Terminal-Bench v2.0. arXiv 2604.16529. publisher · read in our library knowledge/fetched/cmp_autodev_ttcagentic.html · pin h6ceef45a19fcda26c8bb90d55d365427ab65a991a1eb02d17f30466e5b221316 · accessed 2026-08-27 · published-paperGrounds: PROPOSE: sampling and selection against the verifier
  8. [runornot] Lin et al. To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair (submitted 2026-06-25). 7,745 SWE-bench leaderboard traces and 3,000 end-to-end attempts over 200 instances with Claude Code, Codex and OpenCode under four execution paradigms: agents average 8.8 test runs per task ranging 2 to 19, prohibiting execution narrows the resolve-rate gap by only 1.25 points while saving tokens and time, and the benefit concentrates on a subset of instances. arXiv 2606.26978. publisher · read in our library knowledge/fetched/cmp_autodev_runornot.html · pin hf05b212a46e782b0c126b8c9c40cafc28d95a124330956e48a6b2c51f7012ff4 · accessed 2026-08-27 · published-paperGrounds: VERIFY: the execution budget measured per task class
  9. [adi] Xiang, Xu, Chu, Tian and Zhang. Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis (submitted 2026-04-27, FSE 2026). The Agent-centric Debugging Interface with a function-level Frame Lifetime Trace resolves 63.8 percent of SWE-bench Verified at 1.28 dollars per task on Claude-Sonnet-3.7 and lifts existing agents by 6.2 to 18.5 percent. arXiv 2604.24212. publisher · read in our library knowledge/fetched/cmp_autodev_adi.html · pin h1d19de173a1db32df5a5a4be610701e886f991b1b70c34a7648d1695a98af07d · accessed 2026-08-27 · published-paperGrounds: LOCALIZE: function-level dynamic evidence beside static retrieval
  10. [brtcogen] Cheng et al. Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair (Google, submitted 2026-01-27, revised 2026-08-21). On 120 human-reported bugs, instructing the repair agent to emit the bug-reproduction test in the same patch as the fix yields reproduction tests for at least as many bugs as a dedicated test agent without lowering the plausible-fix rate, and patch selectors that account for test change select better patches. arXiv 2601.19066. publisher · read in our library knowledge/fetched/cmp_autodev_brtcogen.html · pin h635af7312b522354983d86304fa02c6ae30b53426753bb8bc460ce63ec5fb400 · accessed 2026-08-27 · published-paperGrounds: REPRO: the reproduction tooth cogenerated with every fix
  11. [chainswe] Jin et al. ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance (submitted 2026-07-01). 304 chronologically chained issues across 54 Python projects mined from six SWE-bench-family datasets; agents evaluated on sequential dependent fixes in one shared codebase drop by up to 70 percent as the chain lengthens. arXiv 2607.02606. publisher · read in our library knowledge/fetched/cmp_autodev_chainswe.html · pin hfcca15585b04cae26be323f5b4c90713094a4cc53f69cfd89edf74c67ded85d0 · accessed 2026-08-27 · published-paperGrounds: CHAIN: dependent tasks without a reset and the per-chain-length number
  12. [swerebenchv2] Badertdinov, Nekrashevich, Shevtsov and Golubev. SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale (submitted 2026-02-27, revised 2026-06-01). An automated language-agnostic pipeline yielding 32,079 executable tasks across 20 languages from 3,617 repositories with pre-built images, unsound instances filtered by an ensemble of LLM judges, plus 120,000 further tasks with metadata; built to relieve the task scarcity that constrains RL training of agents. arXiv 2602.23866. publisher · read in our library knowledge/fetched/cmp_autodev_swerebenchv2.html · pin h8bbd76e19cdd1626096a22d67ee80eb0572ab67c894d489cbd54e6c30cb39106 · accessed 2026-08-27 · published-paperGrounds: REFEREE: the post-cutoff multilingual task source
  13. [swebenchpro] Deng et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? (submitted 2025-09-21, revised 2025-11-14). 1,865 problems from 41 actively maintained repositories split into public, held-out and commercial sets; long-horizon tasks that take a professional hours to days and patch multiple files; human-verified and presented as a contamination-resistant testbed. arXiv 2509.16941. publisher · read in our library knowledge/fetched/cmp_autodev_swebenchpro.html · pin h21f69a25e615c4223357fe653663aeb94f962732cbbb1277e827bd3014f9df53 · accessed 2026-08-27 · published-paperGrounds: REFEREE: the held-out long-horizon bar and the multi-file rung
  14. [hal] Kapoor et al. Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation (submitted 2025-10-13). A standardised harness running 21,730 rollouts across 9 models and 9 benchmarks for about 40,000 dollars; higher reasoning effort reduced accuracy in the majority of runs; LLM-aided log inspection found agents searching for the benchmark on HuggingFace instead of solving the task; all 2.5B tokens of logs released. arXiv 2510.11977. publisher · read in our library knowledge/fetched/cmp_autodev_hal.html · pin hd1184a420b01d05e12edfc73b31048c0957348296eb5cf6c34b420267ababcfb · accessed 2026-08-27 · published-paperGrounds: LEDGER and GUARD: logs as the audit and benchmark-gaming as a named failure
  15. [dontpassk] Samandar et al. Don't Pass at k: A Bayesian Framework for Large Language Model Evaluation (arXiv 2510.04265, under review ICLR 2026). Replaces pass at k and avg at N with posterior success probabilities and credible intervals under a Dirichlet prior, closed-form for any weighted rubric; declares a difference only when intervals do not overlap and allocates further samples until intervals reach a pre-specified width; faster convergence and more stable rankings than pass at k on synthetic and math benchmarks. publisher · read in our library knowledge/fetched/cmp_autodev_dontpassk.html · pin h8cf03e8def93f54ba6c87e717816361daa40eb863a8379b2013688dd3a286ced · accessed 2026-08-27 · published-paperGrounds: RIGOR: the interval and the decision rule every published number carries
  16. [statprecipice] D'Oro et al. Computer Use at the Edge of the Statistical Precipice (submitted 2026-05-07). A 1 MB replay script that never observes the screen outperforms frontier models on static benchmarks, with expected success exactly the source agent's pass at k in deterministic environments; the PRISM principles (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, multifactorial variability) and Wilson score intervals with hierarchical bootstrap for the nested structure of agent benchmarks. arXiv 2605.08261. publisher · read in our library knowledge/fetched/cmp_autodev_statprecipice.html · pin hf1c4c6ed1d17b82a952ebcd3af187473c0b7afc840e861159499f1145e242e19 · accessed 2026-08-27 · published-paperGrounds: RIGOR and CONTROL: the replay null control and the hierarchical interval
  17. [swebenchcase] Martinez and Franch. What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair (submitted 2026-02-04). First study of the Lite (79 entries) and Verified (133 entries) leaderboards: most entries come from industry, proprietary LLMs dominate with the Claude family at the top, and academic open-source entries remain competitive. arXiv 2602.04449. publisher · read in our library knowledge/fetched/cmp_autodev_swebenchcase.html · pin h45932060b651ab7ae1db2288e4da062979b7be244394bbe92b525bc963958218 · accessed 2026-08-27 · published-paperGrounds: REFEREE: why Verified is a sanity check and not the frontier for this board
  18. [swezerohero] Ludwig, Ahmad, Majumdar and Ginsburg. From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents (submitted 2026-04-02, revised 2026-05-06). Two-stage SFT over 300,000 execution-free and 13,000 execution-based trajectories distilled from open-weight frontier models; SWE-HERO-32B resolves 62.2 percent of SWE-bench Verified and 44.1 percent of SWE-bench Multilingual zero-shot from Python-only training. arXiv 2604.01496. publisher · read in our library knowledge/fetched/cmp_autodev_swezerohero.html · pin h2fe810a07d9da7fbc919617f7dfeb5979b17f5a1d140baa999b561712c1ba926 · accessed 2026-08-27 · published-paperGrounds: DISTILL: what a verified-trajectory corpus is worth and where the terms let us build one
  19. [sweagent] Yang, Jimenez, Wettig, Lieret, Yao, Narasimhan and Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (submitted 2024-05-06, final 2024-11-11). A custom agent-computer interface for creating and editing files, navigating repositories and executing tests; pass at 1 of 12.5 percent on SWE-bench and 87.7 percent on HumanEvalFix at the time. arXiv 2405.15793. publisher · read in our library knowledge/fetched/cmp_autodev_sweagent.html · pin h1589f706b0380d637196a65970b23cb89037465fe6d883af1007e34a41d8bdbf · accessed 2026-08-27 · published-paperGrounds: LOCALIZE and VERIFY: the agent-computer-interface lineage every rival column shares
  20. [openhands] Wang et al. OpenHands: An Open Platform for AI Software Developers as Generalist Agents (submitted 2024-07-23, revised 2025-04-18, ICLR 2025, MIT licence). Agents that write code, use a command line and browse; sandboxed environments for code execution; multi-agent coordination; benchmarks incorporated in the platform and evaluated over 15 tasks including SWE-bench and WebArena. arXiv 2407.16741. publisher · read in our library knowledge/fetched/cmp_autodev_openhands.html · pin h2b9f10c47d1964713d513d8652d75cc988bbd3d7415f89abd8446bfab6e4e3c5 · accessed 2026-08-27 · published-paperGrounds: The OpenHands column: sandbox, tools and evaluation harness
  21. [claudecode] Anthropic. Claude Code documentation, Overview (fetched 2026-08-27). An agentic coding tool that reads the codebase, edits files, runs commands and integrates with development tools across terminal, IDE, desktop and web; git staging, commits and pull requests; MCP connections; CLAUDE.md instructions, auto memory, skills and hooks; parallel subagents and an Agent SDK; headless piping in CI; scheduled routines. publisher · read in our library knowledge/fetched/cmp_autodev_claudecode.html · pin h48897aa2d9b166167b54b333c2c440a6e6c58daf526f9e83363cbef66c6a18c6 · accessed 2026-08-27 · vendor-docGrounds: The Claude Code column: the product surface as the vendor describes it
  22. [codex] OpenAI. Codex repository README (openai/codex, main, fetched 2026-08-27). Codex CLI is a coding agent from OpenAI that runs locally on your computer; installation for Mac, Linux and Windows, editor integrations, sign-in with ChatGPT plans or an API key. The README as mirrored does not describe sandboxing, approval modes, cloud tasks or MCP, so those cells are coded from silence and say so. publisher · read in our library knowledge/fetched/cmp_autodev_codex.md · pin hba4e1f69ff48386e72a9c5e1edaf76aad64a475c2d51af79ccba6d1128261ba7 · accessed 2026-08-27 · vendor-docGrounds: The Codex column: a thin mirror, coded honestly as such
  23. [dgm] Zhang, Hu, Lu, Lange and Clune. Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents (submitted 2025-05-29, revised 2026-03-12). A self-improving system that modifies its own code and validates each change empirically on coding benchmarks from an archive of agents; SWE-bench from 20.0 to 50.0 percent and Polyglot from 14.2 to 30.7; all experiments under sandboxing and human oversight. arXiv 2505.22954. publisher · read in our library knowledge/fetched/cmp_autodev_dgm.html · pin h30ff41b60abf5df7cdd7af79d2b22eb92a2777afedb85dd297874327961f4649 · accessed 2026-08-27 · published-paperGrounds: EVOLVE and SANDBOX: empirical validation of self-modification under isolation
  24. [alphaevolve] Novikov et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery (DeepMind white paper, submitted 2025-06-16). An evolutionary coding agent that continuously receives feedback from one or more evaluators and iteratively improves an algorithm; a 4x4 complex matrix multiplication in 48 scalar multiplications, the first improvement over Strassen in 56 years, and deployed data-centre scheduling and accelerator-design results. arXiv 2506.13131. publisher · read in our library knowledge/fetched/cmp_autodev_alphaevolve.html · pin h7a0d9acc7002b1960fea119a059a925a6c20d27edd8233777aa0179cb5e8f451 · accessed 2026-08-27 · published-paperGrounds: EVOLVE: the evaluator-driven loop our gate-permil fitness already mirrors
  25. [codeharness] Ning et al. Code as Agent Harness (survey, 42 authors, submitted 2026-05-18). Code as the operational substrate for agent reasoning, acting, environment modelling and execution-based verification; open challenges named include evaluation methodology, verification with incomplete feedback, regression prevention, shared state consistency and human oversight. arXiv 2605.18747. publisher · read in our library knowledge/fetched/cmp_autodev_codeharness.html · pin h21feb4689d1535ec99d95edf34e5b363fb68165b7572f9bb45761f89687ad66b · accessed 2026-08-27 · published-paperGrounds: VERIFY: regression prevention as a named open challenge the promote ruler answers
  26. [swebench] Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv 2310.06770). 2,294 real GitHub-issue problems across Python repositories; a patch is RESOLVED only when the FAIL_TO_PASS tests pass and the PASS_TO_PASS tests still pass, with the test suite as the sole judge. publisher · read in our library knowledge/fetched/cmp_autodev_swebench.html · pin hcd170cab3a56f891ee94070d5748d923685f2a84a91201d6336a221436fd0f9f · accessed 2026-08-27 · published-paperGrounds: VERIFY: the contract our sovereign analog adopts verbatim
  27. [oxalpha] OpenRouter. Ox Alpha model page (stealth/ox-alpha, released 2026-08-20; mirror pinned by the extllm domain 2026-08-24). A reasoning model designed for coding, sustained agentic work and production workloads; 1,048,576-token context; the page banner states prompts and completions are retained by the provider and not used for training and defers all other use to the Stealth Model Terms. publisher · read in our library knowledge/fetched/cmp_extllm_oxalpha.html · pin h108b8c85dfed1a9b5726af14f6ca955c1a7f5d56c3d6161f70009d88a42586e2 · accessed 2026-08-24 · vendor-docGrounds: SAMPLER: the free frontier sampler behind the warden, capability facts only
  28. [stealtheula] OpenRouter. Stealth Program End User License Agreement (updated 2026-07-06; mirror pinned by the extllm domain 2026-08-24). Stealth models are offered free of charge for purposes of collecting User Content for use in Stealth Model training and improvement, with an irrevocable perpetual licence to the provider, and Exhibit A forbids publicly disseminating confidential technical information regarding the performance of the models. publisher · read in our library knowledge/fetched/cmp_extllm_stealtheula.html · pin hfed941390d8fbcb7b06d92337a4b367997e6de39452edec152dfee899f2f6a31 · accessed 2026-08-24 · vendor-docGrounds: GUARD: why no ox-alpha performance number appears on this page and why the warden classifies every prompt byte

generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/autodev.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers