nishi code wiki / research / npc minds
NPC intelligence, behaviour and social simulation
SOTA census · compiled 2026-08-01 · 7 axes · ~45 sourced claims · 8 declared gaps
lineage: forked from research_rtrender parent domain: embodiment target: first-person farming-life village sim on the sovereign voxel engine
Claims are labelled SHIPPING / RESEARCH / ANNOUNCED / ABANDONED. Gaps are declared UNVERIFIED rather than guessed.
1. Classical shipping SOTA — the 25-year canon
| System | Status | Mechanism | Source (date) |
|---|---|---|---|
| The Sims — smart-object advertisements | SHIPPING | Objects broadcast motive-satisfaction adverts; the sim scores adverts against its current needs and picks by weighted utility. Behaviour lives in the WORLD, not the agent. | Forbus & Wright, “Under the Hood of The Sims” — users.cs.northwestern.edu/~forbus/c95-gd/lectures/The_Sims_Under_the_Hood_files/v3_document.htm (2001); qrg.northwestern.edu/papers/Files/Programming_Objects_in_The_Sims.pdf (2001) |
| RimWorld — ThinkTree + thought ledger | SHIPPING | Hierarchical priority tree of ThinkNodes selects the current job; mood = base + a ledger of signed, timestamped thoughts that decay. Legible: every mood has an itemised receipt. | rimworldwiki.com/wiki/Thoughts; github.com/roxxploxx/RimWorldModGuide/wiki/SHORTTUTORIAL:-How-Pawns-Think (accessed 2026-08-01; game 1.0 2018-10-17) |
| Dwarf Fortress — facet model | SHIPPING | Personality = beliefs (values) + goals + ~50 facets each scored 0–100; facets gate which thoughts fire and which needs exist; memories shift facets over time; breakdowns are personality-dependent. | dwarffortresswiki.org/index.php/DF2014:Personality_facet (accessed 2026-08-01; emotion rewrite v0.42 2015-12-01) |
| Crusader Kings 3 — traits + stress | SHIPPING | Trait set fixed at adulthood; acting against traits generates stress → mental breaks. Documented weakness: traits never change after ~16 and their effects reduce to stat/opinion modifiers — mechanically transparent, shallow as simulation. | forum.paradoxplaza.com/forum/threads/personality-traits-never-change.1602268/ (accessed 2026-08-01; game 2020-09-01) |
| Shadows of Doubt — per-NPC 24 h schedules + AI-LOD | SHIPPING | Every citizen has a home, a job, and an individualised 24 h schedule (typically 4–10 journeys/day). AI update rate is high only near the player: “the performance cost of 95% of the citizens at any one time [is] relatively insignificant”. | colepowered.com/shadows-of-doubt-devblog-8-simulating-a-city/ and devblog-15-moving-in-the-citizens (accessed 2026-08-01); en.wikipedia.org/wiki/Shadows_of_Doubt (1.0 2024-09-26) |
| Nemesis system | SHIPPING, patented | Enemies remember encounters and rise through a procedural hierarchy on vendettas. Patent US10926179B2 (granted 2021-02-23) runs to 2036-08-11; used in exactly two games; Monolith shut down 2025-02. | patents.google.com/patent/US10926179B2 (2021-02-23); engadget.com “locked behind a patent until 2036” (2025-02) |
The pattern across all five healthy systems: integer/enum state, utility scoring, and a legible ledger — no neural nets anywhere in the shipped canon. The Sims' inversion (objects advertise, agents score) is the single highest-leverage trick: content scales by ADDING OBJECTS, not by editing agent code. Shadows of Doubt is the existence proof that hundreds of fully-scheduled citizens fit in a real-time budget if simulation fidelity follows player proximity.
2. Social simulation — the gossip lineage nobody shipped past
The academic lineage: Comme il Faut (McCoy, Treanor, Mateas, Wardrip-Fruin — ojs.aaai.org/index.php/AIIDE/article/view/12454, AIIDE 2011), a rule-based playable social model, shipped as Prom Week (dl.acm.org/doi/10.1145/2159365.2159425, FDG 2011; game released 2012). Then Talk of the Town (Ryan, Summerville, Mateas, Wardrip-Fruin, “Toward Characters Who Observe, Tell, Misremember, and Lie” — EXAG @ AIIDE 2015; jamesryan.world/talktown): characters form knowledge from observation, propagate it through discrete conversations, misremember, forget, and lie — information flow is agent-driven, not abstractly diffused. Open-source successor: Neighborly (Johnson-Bey, Nelson, Mateas — IEEE CoG 2022; github.com/ShiJbey/neighborly), a ToT-style emergent-narrative sandbox. All four: RESEARCH.
3. LLM NPCs — the honest ledger, 2023–2026
| Case | Status | What actually happened | Source (date) |
|---|---|---|---|
| inZOI “Smart Zoi” | SHIPPING | The first life sim shipping an on-device LLM: a 0.5 B-parameter Mistral NeMo Minitron SLM via NVIDIA ACE, running 100% locally, no server dependency; feeds ~400 mental traits/needs/karma values into autonomous decisions and rumour propagation. | playinzoi.com/en/news/8419; nvidia.com/en-us/geforce/news/nvidia-ace-naraka-bladepoint-inzoi-launch-this-month/ (2025-03; early access 2025-03-28) |
| Fortnite Darth Vader | SHIPPING | Cloud stack (Gemini 2.0 Flash + ElevenLabs Flash v2.5, licensed James Earl Jones voice). Jailbroken into slurs and profanity within hours of launch; Epic hotfixed in ~30 min and added a parental control for AI conversations. | pcgamer.com; kotaku.com/fortnite-star-wars-ai-darth-vader-james-earl-jones-1851781018 (2025-05-16) |
| Suck Up! | SHIPPING | Cloud inference with metered economics: players get 10,000 AI tokens ≈ 40–50 h of gameplay, because “personal computers today are simply not powerful enough” (Proxima). Marginal inference cost passed to the player as a consumable. | themagicrain.com/2024/04/suck-up-is-a-vampire-game-that-uses-a-i-to-interact-with-its-players/ (2024-04; game 2023-11) |
| Inworld AI | ANNOUNCED pivot | The flagship cloud character-engine vendor pivoted to distilling models for local execution with hybrid cloud fallback (GDC 2025), then to a general runtime. The cloud-character-as-a-service business did not hold. | inworld.ai/blog/gdc-2025 (2025-03); wccftech.com Inworld GDC 2025 Q&A (2025-03) |
| Replica Studios | ABANDONED | AI voice-for-games vendor, first SAG-AFTRA AI voice deal (2024-01-09) — service ended 2025-06-30 after seven years, citing funding and competition. | replicastudios.com (farewell page); nbcnews.com/tech/video-games/sag-aftra-replica-studios-voice-actors-video-games-rcna133162 (2024-01-09); multilingual.com (2025) |
| Ubisoft NEO NPCs | ANNOUNCED, never shipped | GDC 2024 prototype (press release 2024-03-19) built on Inworld; nothing reached a shipped game. Evolved into another demo, “Teammates” (2025-11) — still a prototype two years on. | staticctf.ubisoft.com PRESS_RELEASE_GDC (2024-03-19); aiandgames.com/p/ubisofts-teammates-demo-and-their (2025-11-26) |
Read the ledger honestly: the survivors are small and on-device (inZOI's 0.5 B), the cloud voice/character vendors died or pivoted to local, the highest-profile cloud NPC was jailbroken in hours, and the studio with the biggest budget shipped nothing in two years of demos. Cloud LLM NPCs failed on cost, on safety, and on latency — simultaneously.
4. Sizing the local model — and the hybrid canon
The shipped precedent is exactly one data point: 0.5 B parameters on-device (inZOI, 2025-03). The emerging consensus architecture — visible in inZOI, in Inworld's pivot, and in the research line of §5 — is the hybrid canon: classical utility AI at frame rate, the LLM consulted at low frequency (dialogue moments, day-boundary planning, reflection), never in the per-tick loop.
Against our substrate: the sovereign integer LLM path runs a 0.5 B model at ~1.06 tok/s on CPU today (in-house, 2026-08-01) — a 20-token villager line takes ~19 s, which is ambient/asynchronous territory only (generate during sleep ticks, deliver when done). The measured GPU path (362 GB/s DRAM-resident, Q4_K kernel bit-exact vs CPU reference) projects ~42.1 tok/s on a 14 B Q4_K — sub-second short lines and day-tick planning. Both in-house measurements, 2026-08-01.
5. Memory and planning — the cost wall and the 10–100x door
Generative Agents (Park et al., arxiv.org/abs/2304.03442, 2023-04-07) is the reference architecture: a memory stream, retrieval scored by recency × importance × relevance, periodic reflection into higher-level beliefs, and recursive day planning. It produced believable emergent behaviour (party invitations propagated, agents showed up) — and it cost thousands of US dollars in tokens for 25 agents over two simulated days (paper's own limitations section; commonly cited as ~$2,000 — see UNVERIFIED #3; forbes.com Joon Sung Park interview, 2024-02-20). RESEARCH.
The reductions came fast: Lyfe Agents (arxiv.org/abs/2310.02172, 2023-10-03) — option-action decision hierarchy, asynchronous self-monitoring, Summarize-and-Forget memory — at 10–100x lower cost than prior generative agents in real-time 3D social scenarios. Affordable Generative Agents (arxiv.org/abs/2402.02053, 2024-02-03) — lifestyle-policy reuse and social-relationship compression, same believability benchmarks at a fraction of the calls. Both RESEARCH.
The transferable lesson: the believability came from the memory architecture, not the model — and cost collapses when the LLM is consulted only at decision boundaries while cached policies replay routine days. For a farming sim, most villager days ARE routine: that is the AGA case in its purest form. The classical ledger of §1 (thoughts, facets, schedules) doubles as the memory stream the retrieval layer reads — build it once, serve both masters.
6. Evaluation — there is no yardstick
No industry-standard NPC believability benchmark exists (as of 2026-08-01; declared, not proven — see UNVERIFIED #5). The closest things: academic believability rubrics (“Towards an Understanding of Character Believability”, dl.acm.org/doi/10.1145/3582437.3582466, FDG 2023-04) and Generative Agents' human-judged TrueSkill comparisons (2023-04). Nobody ships against a number. Consequence for us: define our own gates — a replay-determinism gate (same seed + save ⇒ identical village, LLM lines included), a schedule-coherence gate (every villager reachable at their scheduled place), and a gossip-provenance gate (every belief traceable to an observation or a retelling chain). All three are mechanical, bit-exact checkable, and match the house gate discipline — a believability bar that CAN fail.
7. What this means for the Nishi stack — the ranked ladder
Constraint filter: proprietary language → WASM; integer/fixed-point deterministic core; no third-party libraries; substrate already has 12 NPCs, a 20-bit appearance genome, inherited bold/shy personality bits, WANDER/SOCIALIZE/REST/FORAGE/FLEE states, an entity store, save/load, and bit-exact sim. Rungs 1–5 are classical and cheap; rungs 6–7 are LLM and expensive. Climb in order — each rung's data structures feed the next.
| # | Rung | Payoff / cost | Exit criterion |
|---|---|---|---|
| 1 | Smart-object advertisements (The Sims, 2000) | Objects advertise (need, magnitude); NPC picks integer argmax weighted by needs. Content scales by adding objects. Days of work, pure integer. | An NPC chooses FORAGE vs REST from advertised utilities; decision trace logged; replay bit-exact. |
| 2 | Thought ledger (RimWorld) | Ring buffer of signed, timestamped thoughts → mood integer feeding state weights. Legibility for free: every mood has a receipt. Days. | A social snub measurably changes an NPC's day; ledger survives save/load. |
| 3 | 24 h schedules + AI-LOD (Shadows of Doubt) | Per-NPC daily plan over the village; full tick near player, coarse tick far. The proven scaling trick (“95% of citizens… insignificant”). ~1–2 weeks. | 100+ NPCs at the current 12-NPC tick budget; determinism preserved across LOD transitions (player position is replay input). |
| 4 | Facet expansion (Dwarf Fortress) | Widen bold/shy bits to ~8 integer facets (0–255) in the genome, feeding utility weights, thought susceptibility, gossip propensity. Inheritance machinery already exists. Days. | Two villagers with different facets visibly diverge on the same day's stimuli; facets inherited and mutated at birth. |
| 5 | Symbolic gossip (Talk of the Town, 2015) — THE DIFFERENTIATOR | Knowledge records: (subject, claim, source, tick, confidence); mutate on retell; decay; lies. No shipped game has surpassed the 2015 prototype — this rung alone is a marketable feature. ~2–4 weeks. | Player traces a false rumour to its origin through the provenance chain; chain survives save/load; gossip-provenance gate GREEN. |
| 6 | LLM ambient dialogue (hybrid canon; inZOI precedent) | 0.5 B local, greedy integer decode, generated asynchronously at today's ~1.06 tok/s; state from rungs 1–5 is the prompt. First deterministic replay-safe LLM NPC in any engine. ~2–4 weeks after rung 5. | A villager line regenerates bit-exactly from (seed, save); replay-determinism gate GREEN with LLM enabled. |
| 7 | LLM day-planner + reflection (Generative Agents → AGA) | Day-boundary planning and reflection on the 14 B Q4_K GPU path (~42 tok/s projected); AGA-style plan caching keeps cost per NPC-day bounded. Blocked on the GPU daemon shipping. ~1–2 months. | Plan-cache hit rate >80% on routine days; tokens per NPC-day bounded and logged; behaviour still passes the schedule-coherence gate. |
Patent note: the Nemesis-style procedurally-promoted vendetta hierarchy is claimed by US10926179B2 until 2036-08-11. The gossip/provenance design above is mechanically distinct (information propagation, not rank promotion on player-death events), but any future faction-promotion feature should be checked against the claims first.
Declared UNVERIFIED — do not treat as measured
- The commissioned primary source was lost. The prior deep-research agent's transcript (tasks/a2c4372cc6834815c.output) was 0 bytes — killed by the same shared 200-call search budget documented in research_rtrender's method note. Every claim here was re-verified from scratch on 2026-08-01; the original report may have contained sourced claims this brief lacks.
- Publication dates for wiki/blog sources (RimWorld wiki, Dwarf Fortress wiki, Paradox forum thread, ColePowered devblogs 8/15) — pages are undated or rolling; cited as accessed 2026-08-01.
- The precise $2,000 figure for Generative Agents — the paper states “thousands of US dollars” for 25 agents/2 days; the exact number circulates in secondary coverage and was not located in the paper itself.
- Prom Week's often-cited ~5,000 social rules — not re-verified; omitted from the body.
- “No shipped game has surpassed Talk-of-the-Town-depth gossip” and “no industry believability benchmark exists” — absence claims; search sweeps found no counterexample, but a negative cannot be proven by search.
- Aggregate cloud-NPC cost projections ($0.5–2 M/yr at 100 k DAU) — single secondary source (Medium, 2026), uncorroborated; excluded from the body, Suck Up!'s first-party numbers used instead.
- inZOI's Smart Zoi consultation cadence (how often the 0.5 B model is actually invoked per Zoi) — not publicly documented; the “hybrid canon” framing rests on the architecture pattern, not a measured inZOI rate.
- github.com/ShiJbey/neighborly cited from prior knowledge; the CoG 2022 paper was verified, the repo URL was not re-fetched this session.
Method. 13 targeted web searches on 2026-08-01, one per claim cluster (patents, inZOI, Fortnite/Vader, Replica/Inworld, Shadows of Doubt, Generative Agents cost, Lyfe/AGA, Suck Up!, Ubisoft NEO, ToT/Neighborly, CK3, Sims, RimWorld/DF); each load-bearing number traced to a first-party or primary source where one exists. In-house numbers (1.06 tok/s, 362 GB/s, 42.1 tok/s projected, 2.985e-08) are session-measured on the sovereign stack, 2026-08-01. What would change the conclusions: a shipped game demonstrating ToT-depth gossip, or evidence of a deterministic-inference engine shipping elsewhere — either would demote the two moat claims.