Nishi Family › Compare › Autonomous Team and Software Agents
Nishi Compare · measured, not asserted
Autonomous Team and Software Agents
Nishi vs the field — every Nishi cell is measured against real organ source at emit time; each gap names the watch contract that will close it.
Nishi vs SWE-agent and OpenHands and Aider and AutoGen
Where we are. Measured 2026-08-19. The team is a real governance-and-verification shell -- zero-token execute-and-verify loop, anti-drift coach, heartbeat boards, governed pulses, lineage, honest autonomy metering, refusal over fabrication -- and the 2026-07-05 diagnosis still stands: it is not a builder. The forge lane (local Coder seat with a byte-exact judge, measured 0 of 3) is the incumbent for authoring and is now named on the headline row; SWE-bench Verified ingest and grading exist for the issue-to-patch rung; the container host exists for the sandbox.
Where we need to go. Give real workstreams machine-executable targets, move authoring off the zero floor under the coach and gates the shell already provides, then repo loops, governance UX and agents -- the governance the field lacks wrapped around the authorship we lack.
8 of 20 capabilities measured|1 of them measured exceeds|12 open|coverage 400/1000|adoption 3 full / 5 partial
Do this next — computed by the ranker, never chosen by a seat
Order from nx_compare_rank (nx_dr_ocm: (deficit + cost-of-delay + option + enables) x sponsor x self-sufficiency x momentum / cost). FINISH rows are rungs whose symbol is present but whose organ is short of full adoption: the cheapest closures on this board, listed before any new work. Stamp: # asof=1787883522 domain=team target_version=0.1 rungs=10 done=0 open=10 finish=0 ranker=nx_dr_ocm
| # | Stage | Rung | Priority | Derivation |
|---|---|---|---|---|
| #1 | 0.1 | Machine-executable targets on real workstreams (R0) ee_target_declare | 900 | v=9 m=1 c=10 |
| #2 | 0.1 | Forge authoring past the zero floor (R1) fe_author_organ | 400 | v=8 m=1 c=20 |
| #3 | later | Budget meter (R9) ab_meter | 4800 | v=24 m=1 c=5 |
| #4 | later | Agent memory (R6) am_recall | 2500 | v=25 m=1 c=10 |
| #5 | later | Human-in-the-loop approvals (R5) th_approve | 1200 | v=12 m=1 c=10 |
| #6 | later | Sandbox per task (R4) as_spawn | 933 | v=14 m=1 c=15 |
| #7 | later | Agent session UI (R8) au_page | 466 | v=7 m=1 c=15 |
| #8 | later | Multi-agent conversation (R7) ac_turn | 350 | v=7 m=1 c=20 |
| #9 | later | Git-native pair loop (R3) gp_edit | 333 | v=5 m=1 c=15 |
| #10 | later | Issue-to-patch over SWE-bench Verified (R2) ip_resolve | 250 | v=5 m=1 c=20 |
Critical path — contract, done-rule, executor, cost
| Rung | Closes with | Definition of done (pre-declared) | Executor | Est. |
|---|---|---|---|---|
| Machine-executable targets on real workstreams (R0) | ee_target_declare | Stalled workstreams declare a target row (gate or bench plus accept rule) the exec engine can run; the engine's own coverage line moves from 0 of 104 and the number is the proof | Organ | 1 u |
| Forge authoring past the zero floor (R1) after R0 | fe_author_organ | The local forge cascade authors a novel organ that compiles, runs and passes its declared gate; nx_forge_local_bench moves from 0 of 3 by a PRE-DECLARED floor and the floor ratchets | Local model | 2 u |
| Issue-to-patch over SWE-bench Verified (R2) after R1 | ip_resolve | The forge loop resolves banked SWE-bench Verified instances (the swebv ingest and grader already exist) end to end; pass rate published beside the field's, never claimed above it | Organ | 2 u |
| Git-native pair loop (R3) after R1 | gp_edit | Repo map, diff-style edits and auto-commit over the sovereign git; a planted failing test is fixed by the loop and committed with the gate's evidence | Organ | 1.5 u |
| Sandbox per task (R4) | as_spawn | Each agent task runs in the container host the containers domain already ships, with the ocap grant of the swarm; escape attempts in the fixture are refused | Organ | 1.5 u |
| Human-in-the-loop approvals (R5) | th_approve | DENY-WINS approvals from the workflow domain gate every external action of an agent; the approval is a recorded row | Organ | 1 u |
| Agent memory (R6) | am_recall | Per-agent memory stream with recall over it (nx_charmem and em_recall are the substrate); a planted false memory is refutable by the reflection primitive | Organ | 1 u |
| Multi-agent conversation (R7) after R6 | ac_turn | Two or more sovereign-seat agents converse through the conductor with turn budgets; the transcript is additive and replayable; outcome graded by the referee | Local model | 2 u |
| Agent session UI (R8) | au_page | Live agent sessions on the portal with the board's liveness semantics; a stalled session reads STALLED, never ACTIVE | Organ | 1.5 u |
| Budget meter (R9) | ab_meter | Tokens, wall time and any external spend metered per session as rows; the zero-token loop reads zero, which is the point | Organ | 0.5 u |
Milestones
| Milestone | Rungs | Cumulative |
|---|---|---|
| M0 · Targets and authoring | R0,R1 | 3 u |
| M1 · Repo loops | R2,R3,R4 | 8 u |
| M2 · Governance UX | R5,R8,R9 | 11 u |
| M3 · Agents | R6,R7 | 14 u |
comparewatch- plane row flips with it. The flip is necessary, not sufficient: it proves the symbol exists, never that the capability is good. The bar is the rung's pre-declared done-rule, proven by its gate — a symbol shipped without the behaviour behind it is a defect, and the flip is exactly what makes that defect visible instead of quiet. Competitor marks record documented capability presence — presence, not depth or scale. Adoption is measured too: every measured row carries where its organ stands on the estate's ladder (source → built → promoted → registered → invoked; libraries by importer reach minus validation importers; gates by the execution surfaces that run them). A row is fully adopted only at the top of its ladder; anything short is tagged partial with the exact remedy, so a build nobody promoted can no longer read as shipped. Census stamps: importers asof 1787849099, gate census asof 1787855507 (unix seconds; -1 = census absent).Capability matrix — measured against source
◉ leads / measured exceed● present◐ partial○ absent · click any capability for its evidence
| Capability | Nishi | SWE-agent | OpenHands | Aider | AutoGen |
|---|---|---|---|---|---|
Autonomous execute-and-verify loop on machine targetsMeasured:run_tool exists in runtime/nx_exec_engine.nx, verified at emit. LIVE-gated 2026-07-06: 3/3 targets advanced (2 evo-synth held-out + compiler-health witness); its own output prints the honest coverage: 0 of 104 stalled workstreams declare a machine-executable target yet Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ● | ◉ | ● | ● |
Anti-drift coach gating rung claims (liar-killer)Measured:wc_count exists in runtime/_hdl_build/nx_ws_coach.nx, verified at emit. Gate GREEN then census then coach; only COACH-OK claims a rung -- checks U-axes, dated bars, claim-vs-proof ratio, freshness; no field agent gates its own success claims against measured censuses Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ○ | ○ | ○ | ○ |
Work board with heartbeat livenessMeasured:wb_count_ids exists in runtime/_hdl_build/nx_ws_board.nx, verified at emit. The board refuses to trust a stale ACTIVE label (imports nx_heartbeat_str, wellformed-checked); OpenHands and AutoGen track task state without staleness refusal [openhands-repo] Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ● | ○ | ◐ | ○ | ◐ |
Governed scheduler (warden-owned conductor pulses)Measured:cr_seed exists in runtime/_hdl_build/nx_conductor_registry.nx, verified at emit. Ratchet and training pulses run ONLY governed, never ad-hoc; AutoGen is the orchestration-framework bar [autogen2023] Adoption: LIB-WIRED importers=5 nonval=4 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ◉ |
Organ lineage genealogistMeasured:gen_register exists in runtime/_hdl_build/nx_genealogist.nx, verified at emit. Ancestry and relatedness of organs tracked as data; the field tracks conversation history, not artifact lineage Adoption: LIB-WIRED importers=4 nonval=1 — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Honest autonomy meteringMeasured:am_read exists in runtime/_hdl_build/nx_autonomy_meter.nx, verified at emit. Measures the Claude-touch bar toward 0 -- the team scores its OWN dependence honestly; no field agent measures its own autonomy. 07-15 LIVING: nx_zero_claude_ledger persists the local-vs-Claude lift trend (0.5B 38%->1.5B 53% zero-token, sound because the LCF-kernel verifier only certifies correct answers) + the hybrid cascade routes local-first, escalates only the residual to the lightest Claude Adoption: LIVE — fully adopted (top of its ladder). | ● | ○ | ○ | ○ | ○ |
Teacher and practice loopOpen — no implementing organ is measured for this axis yet. Toy-practice honest (the 07-05 diagnosis); OpenHands and AutoGen have tutorial-ish agent training flows; NOTE this organ is one of the 19 live shadow-name collisions (runtime + _hdl_build copies differ) | ○ | ○ | ◐ | ○ | ◐ |
LLM-driven code authoring (novel organ building)Open — watchingruntime/nx_forge_engine.nx : fe_author_organ, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. THE headline absent: every novel organ falls to Claude (diagnosis 2026-07-05, unchanged today); LLM code-authorship is the field's entire product | ○ | ● | ◉ | ● | ● |
Issue-to-patch repository pipelineOpen — watchingruntime/nx_issue_patch.nx : ip_resolve, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. SWE-agent was born from SWE-bench issue resolution [sweagent2024] | ○ | ◉ | ● | ● | ○ |
Git-native pair-programming edit loopOpen — watchingruntime/nx_git_pair.nx : gp_edit, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Aider is the bar: repo-map + diff edits + auto-commit [aider] | ○ | ● | ● | ◉ | ○ |
Sandboxed per-task agent runtimeOpen — watchingruntime/nx_agent_sandbox.nx : as_spawn, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OpenHands docker runtime is the bar [openhands2024]; our container host exists as a SEPARATE domain (950/1000) but is not wired under the team [sweagent-repo] | ○ | ◐ | ◉ | ○ | ◐ |
Multi-agent conversation frameworkOpen — watchingruntime/nx_agent_chat.nx : ac_turn, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. AutoGen is the bar; our conductor schedules pulses, it does not orchestrate agent dialogue | ○ | ○ | ◐ | ○ | ◉ |
Long-horizon planning and memory for agentsOpen — watchingruntime/nx_agent_memory.nx : am_recall, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Hot frontier axis; our WMS registry and memory files serve Claude, not the team organs | ○ | ◐ | ◐ | ◐ | ◐ |
Benchmark-scored coding evaluationMeasured:lb_one exists in runtime/nx_forge_local_bench.nx, verified at emit. 07-15 NOW PRESENT: nx_forge_local_bench scores the local Coder-0.5B via nx_forge_judge BYTE-EXACT compile+run (provable GREEN, exceeds flaky-test verifiers) -- MEASURED 0/3 honest floor (candidates compile+run, bytes differ). We HAVE the objective grader + a real number; SWE-bench pass-rate is the field bar [swebench2023] and we are far behind (present, not best) Adoption: SOURCE-ONLY — PARTIAL: source exists, never compiled: /api/build it. | ● | ◉ | ◐ | ◐ | ◐ |
Human-in-the-loop approval flowsOpen — watchingruntime/nx_team_hitl.nx : th_approve, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Operator gating exists by DOCTRINE and the workflow domain ships DENY-WINS approvals, but no team organ carries approval flows [autogen-repo] | ○ | ◐ | ● | ● | ◐ |
Web UI for agent sessionsOpen — watchingruntime/nx_agent_ui.nx : au_page, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. OpenHands app is the bar; our /project portal shows projects, not live agent sessions | ○ | ○ | ◉ | ○ | ◐ |
Cost and token budget managementOpen — watchingruntime/nx_agent_budget.nx : ab_meter, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. Every field agent meters provider spend; ours has nothing to meter (see the zero-token exceed) -- absent as a feature, moot by construction | ○ | ◐ | ◐ | ◐ | ◐ |
Zero-LLM-token sovereign autonomous loopMeasured exceed:run_tool in runtime/nx_exec_engine.nx, verified at emit. The whole execute-verify-advance loop is compiled sovereign organs: 0 tokens, 0 API spend, deterministic, offline; every field agent burns provider tokens per action Adoption: BUILT-UNPROMOTED — PARTIAL: compiled, never promoted to the serving root: /api/promote it. | ◉ | ○ | ○ | ○ | ○ |
Refusal-over-fabrication by constructionOpen — no implementing organ is measured for this axis yet. Gate-proven neg-control: a non-zero exit NEVER advances -- novel work is refused (NEEDS_TUTOR), not faked; the field's agents are notorious for claiming unverified success | ○ | ○ | ○ | ○ | ○ |
Machine-executable targets on stalled workstreamsOpen — watchingruntime/nx_exec_engine.nx : ee_target_declare, re-measured on every compare beat. Ship that symbol and this mark flips itself; the comparewatch- plane row flips with it. The headline coverage number the engine prints itself: 0 of 104 stalled workstreams declare a target; a declared target is a row the engine can run | ○ | ● | ◉ | ● | ● |
Person · product · place — not yet measured for this domain
knowledge/compare/team.ppp (rows surface|nishi or c1..c4|label|url|connect naming OUR live surface and each rival's front door), run nx_ppp_probe domain team, and this section fills itself on the next beat: the same ruler on both sides — privacy and CX (third-party hosts, tracker classes, cookies, security headers), design and longevity (design hygiene, computed WCAG contrast, render-blocking resources, unsized media, script weight, theme and motion queries), findability (landmarks, skip link, on-site search, breadcrumb, headings, internal links).References
- [sweagent2024] Yang, Jimenez, Wettig, Lieret, Yao, Narasimhan, Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. arXiv:2405.15793. publisher · read in our library
knowledge/fetched/cmp_team_sweagent2024.html· pinh1589f706b0380d637196a65970b23cb89037465fe6d883af1007e34a41d8bdbf· accessed 2026-08-18 · published-paperGrounds: The SWE-agent column: Best on 'Issue-to-patch repository pipeline' (born from SWE-bench issue resolution, per the note) and Yes on 'LLM-driven code authoring (novel organ building)' -- the agent-computer-interface design that makes LLM code authorship the product. - [sweagent-repo] SWE-agent/SWE-agent (GitHub repository; the v1.1.0 bar of 2025-05-22 was banked from its releases feed). Accessed 2026-08-18. publisher · read in our library
knowledge/fetched/cmp_team_sweagent-repo.html· pinh9d88d10d6325dbd4e1be5b89224922f08f8d867490d0d8bdbc02f7a39968522f· accessed 2026-08-18 · vendor-docGrounds: The SWE-agent version bar in the matrix header (v1.1.0, 2025-05-22) and its Part code on 'Sandboxed per-task agent runtime'. - [openhands2024] Wang, Li, Jiang, Xie, Chen et al. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv:2407.16741, 2024 (ICLR 2025). publisher · read in our library
knowledge/fetched/cmp_team_openhands2024.html· pinh2b9f10c47d1964713d513d8652d75cc988bbd3d7415f89abd8446bfab6e4e3c5· accessed 2026-08-18 · published-paperGrounds: The OpenHands column: Best on 'Autonomous execute-and-verify loop on machine targets', 'LLM-driven code authoring', 'Sandboxed per-task agent runtime' (the docker runtime the note calls the bar) and 'Web UI for agent sessions'. - [openhands-repo] All-Hands-AI/OpenHands (GitHub repository; the cloud-1.45.1 bar released 2026-07-09 was banked from its releases feed). Accessed 2026-08-18. publisher · read in our library
knowledge/fetched/cmp_team_openhands-repo.html· pinhfb28c25cdbdd3cef48f4c4c0562cbd2098e206160fc9ef066e5cf3b6e8e5e2ba· accessed 2026-08-18 · vendor-docGrounds: The OpenHands version bar in the matrix header (cloud-1.45.1, 2026-07-09) and its Part codes on 'Work board with heartbeat liveness' and 'Teacher and practice loop'. - [aider] Aider -- AI pair programming in your terminal (aider.chat docs; the v0.86.0 bar of 2025-08-09 was banked from its releases feed): repo map, diff edits, automatic git commits. Accessed 2026-08-18. publisher · read in our library
knowledge/fetched/cmp_team_aider.html· pinh5c4361d45bfdf3a54ec8f3f17c8c8bf7da1aa975e34db2e36d02b6f5ab977770· accessed 2026-08-18 · vendor-docGrounds: The Aider column: Best on 'Git-native pair-programming edit loop' -- the note names repo-map plus diff edits plus auto-commit as the bar, and this is the page that documents them. - [autogen2023] Wu, Bansal, Zhang, Wu, Li et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:2308.08155, 2023. publisher · read in our library
knowledge/fetched/cmp_team_autogen2023.html· pinh67472921c0c51d0c9e39f3c51a32ab6cc617d181bfe478dccccd9b02ea5775db· accessed 2026-08-18 · published-paperGrounds: The AutoGen column: Best on 'Multi-agent conversation framework' and 'Governed scheduler (warden-owned conductor pulses)' as the orchestration-framework bar the notes name; our conductor schedules pulses and does not orchestrate agent dialogue. - [autogen-repo] microsoft/autogen (GitHub repository; the python-v0.7.5 bar of 2025-09-30 was banked from its releases feed). Accessed 2026-08-18. publisher · read in our library
knowledge/fetched/cmp_team_autogen-repo.html· pinh4607e57b4c2b7afba6f96e6b3dadd087f34191a0da106b875422c2df5e78d2ea· accessed 2026-08-18 · vendor-docGrounds: The AutoGen version bar in the matrix header (python-v0.7.5, 2025-09-30) and its Part codes on 'Sandboxed per-task agent runtime' and 'Human-in-the-loop approval flows'. - [swebench2023] Jimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. arXiv:2310.06770. publisher · read in our library
knowledge/fetched/cmp_team_swebench2023.html· pinhcd170cab3a56f891ee94070d5748d923685f2a84a91201d6336a221436fd0f9f· accessed 2026-08-18 · published-paperGrounds: The 'Benchmark-scored coding evaluation' row: SWE-bench pass rate is the field bar the note names, against which nx_forge_local_bench's byte-exact compile-and-run grader is present-not-best (measured 0/3 honest floor).
generated by nx_swcompare_matrix (sovereign NishiLang organ) from knowledge/compare/team.matrix · every Nishi cell verified against organ source at emit time · watch cells re-measured on every compare beat · zero JS, zero trackers