SHIPPED 2026-07-25 R-CONTENT LANDED same day HONEST: generation 550‰ — not SOTA yet
Prompt-to-Site: one written brief → a linked, live, multi-page website
The website emitter and recombinator climbed a real rung today: nx_uigen_sitegen takes one sentence and emits a whole coherent site — the page PLAN inferred from intent, one design grammar driving palette, shape and rhythm across every page, real cross-page navigation, sitemap and robots derived from the same plan. Deterministic: the same brief always yields the same bytes.
See it live (generated from: “Launch a coffee shop with an online store and a contact form”)
Home · Products · About · Contact — four pages, one brief, zero hand edits. Generated on the hub by the registered MCP tool itself. Since the same afternoon it carries real business copy: a 16-line requirements brief (the factory .req grammar) consumed by the generator — Espresso Blend No. 4 and Single-Origin Huila on the products page are brief-driven, and the gate proves absent keys keep defaults while a missing brief is byte-neutral.
Measured, not asserted
The gate parses the emitted bytes independently of the generator: the site plan matches intent on three held-out briefs (inappropriate pages absent), every page carries the full nav and craft markers, every internal link resolves to a page that exists (the checker is liar-killed with a fabricated link each run), the contact page is a real form and the index a real hero, one generated palette coheres across the site at the exact topic hue, and re-emission is byte-identical. The independent 23-rule judge scored the system-generated index 1000‰ and the contact page 917‰ against an 800‰ floor it did not write.
First field head-to-head (R-h2h v1, measured 2026-07-25)
Same content class, same ruler, both sides: our generated coffee-shop site vs three professionally built coffee e-commerce leaders (Stumptown, Blue Bottle, Verve — live homepage captures). Every number below comes from one mechanical rule set applied identically to all four captures.
| Axis (computed) | Ours (generated) | Stumptown | Blue Bottle | Verve |
|---|---|---|---|---|
| HTML payload | 19 KB | 418 KB | 8.07 MB | 663 KB |
| Render-blocking resources | 0 | 4 | 1 | 6 |
| Third-party asset hosts | 0 | 19 | 5 | 18 |
| Image alt coverage | 1/1 | 52/55 | 32/32 | 97/101 |
| Security headers (of 5: CSP, nosniff, frame, referrer, HSTS) | 5/5 | 4/5 | 4/5 | 4/5 |
Where we stand against the best (grounded competitive map)
| Axis | Status | Note |
|---|---|---|
| Accessibility by construction | EXCEED | WCAG both themes by construction; the field fails 95.9% (WebAIM Million) |
| Code ownership / zero lock-in | EXCEED | we emit plain HTML+CSS you own; builders trap you in their runtime |
| Deterministic generation | EXCEED | same spec, same bytes — no LLM generator can promise this |
| Zero third-party assets | EXCEED | no CDNs, fonts, trackers on any generated page |
| Prompt-to-site (multi-page) | PARTIAL | structure now generated end-to-end; leaders (Wix ADI, Lovable, v0) also generate full prose content — ours is templated |
| Iterative refine | PARTIAL | spec-delta commands measured 8/8; leaders do free chat refine |
| Prompt-to-app / learned model / screenshot-to-code / WYSIWYG | GAP | Lovable, GLM, Design2Code-class, Webflow lead; the honest build queue |
The render loop: judging our own pixels, not our own markup
Operator, same day: “use your OCR — ours still looks like a kid's emission, not professional state of the art.” Right, and it exposed a real hole in the instrument stack: every checker we own reads BYTES. A page can pass contrast math, ARIA-tree diffing, console rules and security headers and still look amateur, because none of those execute CSS and none of them look. So the loop now closes with a headless render that I read back as an image, critique like a designer, and fix — before anything ships.
| Render pass | What LOOKING found (bytes could not) | Fix shipped |
|---|---|---|
| v1 (first Counsel build) | The generative hero scene rendered as a bar chart — varied column heights read as data viz, not architecture. Cool-grey striped section bands. Gradient-dot logo left over from the SaaS shell. | Engraved arch + colonnade with uniform column rules; one warm ivory field, no striped bands; serif wordmark, small-caps nav. |
| v2 | Line art on ivory read as a diagram floating in space. Every band hugged the left edge with the right half empty. Huge padding, low density — sparse, not generous. | Scene reversed out of a deep navy panel (art direction, not decoration); bands became a heading/content grid that fills the measure; vertical rhythm tightened ~40%. |
| v3 | My own grid change dropped the accreditation row into the wrong column, and the practice-list columns started at different heights. | Band content wrapped as one grid cell; rules restored on every list item so both columns align. |
| v4 | The home page reads as a professional firm page. | See it live — and the gate stayed GREEN through every pass, so none of this cost a regression. |
| v5 (round 9) | Rendering the subpages exposed the bigger failure: the site shipped two design languages. Home was serif-on-ivory; About and Contact were a cool-grey SaaS page with the gradient-dot logo, sans headings, rounded fields, a gradient pill button and a misaligned lede. Then, after fixing those, the render caught the dot surviving in the footer. | A Counsel doc tier (serif headings, ivory field, small-caps labels, squared inputs, rectangular navy submit, gold callout rule) + the same wordmark treatment in the footer. New gate tooth T10 family-consistency asserts every subpage of a Counsel site carries the family and that the marketing site's doc pages stay byte-identical — so a site can never silently ship two languages again. Contact · About |
Operator verdict 07-25: “these still look awful” — correct, and now it's a build list
The computed axes are won; the LOOK is not. Diagnosis against a 30-site best-of-breed law-firm reference (MagnifyLab roundup, patterns extracted): our generator applies ONE SaaS-marketing grammar to every vertical — centered gradient hero, rounded card grid, pill buttons, a pricing band on a coffee shop. The best professional-services sites do none of that. What they actually do, and what we build next:
| Reference pattern (from the 30-site bar) | Ours today | Build |
|---|---|---|
| Visual-first hero: video or authentic photography, full-bleed; real people, never stock | text-only hero, gradient span | generative hero ART per brief — the award-page SVG scene technique (aurora/constellation, redirected to dignified motifs) now; gen-img photography-class heroes next (the 5080 pipeline unparked today) |
| Editorial typography: serif display + sans body, mixed-font emphasis, dynamic underlining | one sans stack everywhere | “Counsel” design family: ui-serif display pairing (the Dispatch/finance precedent), CSS emphasis underlines |
| Restrained sophistication: deep blue/green + white + ONE vivid accent; pastel variants | palette math is fine — the SHAPES read SaaS | family-specific component grammar: no pills, no rounded-card grid, no gradient text in professional verticals |
| Whitespace + asymmetry: room to breathe, split layouts, integrated navigation | symmetric centered bands | asymmetric split hero + airy default density in the Counsel family |
| Trust as numbers: prominent statistics, review scores, accreditations throughout | generic trust band | stats-first band (the scorecard component re-skinned), accreditation row, testimonial figure |
| Bespoke client tools: calculators, consultation booking | a contact form | the prompt-to-app axis (already on the GAP queue; booking form is the first rung) |
Four design families, and why the other seven couldn't just be plugged in
A fourth family is live: Market, for commerce briefs — and it fixes the original demo. The coffee shop that started this work had been wearing the SaaS marketing grammar, gradient hero and pricing-tier band and all, which is exactly the “looks like a template” problem. It is now a storefront: utility bar with search and cart, category chips, a promo line, a product grid carrying generated visuals, real names, real prices and add-to-cart, and the shipping/returns/secure row that actually closes a sale. No hero, no feature cards, no pricing tiers.
A third first-class family is live: Journal, for editorial briefs — see it. A magazine brief now produces a masthead over a double rule, a kicker and italic deck, a drop-capped two-column justified lead, a ruled article grid and a dark subscribe band. No hero, no card grid, no pill button. Meanwhile the same generator still routes a law brief to Counsel and a shop brief to the marketing grammar, and the gate proves the three never bleed into each other.
Regression beat: the lane re-verifies itself with no session running
The 10-tooth gate is now a daily clock row on the sovereign job plane (uigensitegate · 86400 · nx_uigen_site_gate.elf), so prompt-to-site, requirements copy, link resolution, refine, family routing and family consistency all re-prove themselves every day with zero Claude involvement. Getting there took reading the dispatcher's parser from source rather than guessing at it: the row writer available to this lane emits seven fields while the clock reads exactly three (name, interval, command), so a naively-written row put the wrong token in the command column. The first attempt was parked rather than shipped, filed as debt, and closed only once the parse was verified — and the parser's own if command is empty, skip rule means the parked row was structurally ignored, never a daily job forking garbage.
Then trying to confirm the first firing exposed a worse problem than a missing check: the gate printed to a standard output nobody reads, so a cron-fired run left no trace at all. “The daily beat is armed” would have been unfalsifiable — no way to prove it ran, or that it passed. A regression beat without an evidence trail is a claim, not a control.
ts=… organ=nx_uigen_site_gate teeth=11 fails=0 families=3 verdict=GREEN — conflict-free, the same pattern the surface sentinel uses. The daily beat is now auditable by anyone with read access, and the first clock-fired line will be distinguishable from a hand-run one by its timestamp. Still honest: the clock-fired run itself has not been observed yet; what changed is that it will now be provable instead of assumed.The worst bug in the lane was not a style slip — it was a content leak
Rendering the storefront's product page (not its front page) found the coffee roastery selling this: “$19/mo Studio — Unlimited generations, Refine loop included”, alongside a $0 Starter with “One generated site” and a $99 Scale tier. Those are this generator's fictional SaaS plans, rendered inside a client's shop, because the marketing product page composed a pricing-tier band whose default copy describes the tool itself. Every byte-checker passed it. It had been live.
Then I pointed the new ruler at the competition, and it broke
A bench that only ever measures its author’s output is not a bench. So the next step was to run it against the field captures from the head-to-head. It found two defects — both in my tool.
And the claim itself was too broad. “Generator-agnostic” is right about the producer and wrong about conventions: against a site using clean URLs this bench cannot assess links at all, and now it says so rather than implying a pass. Our three generated sites still measure 1000 out of 1000 — but now with a published coverage count of four applicable classes, which is a meaningfully different statement than the same number was yesterday. Test a new ruler on inputs it was not designed around; our own output could never have exposed either fault.
So I built the missing ruler — and pointed it at myself first
The gap named in the previous section is now a tool. It reads a directory of emitted files and needs to know nothing about who produced them, so it runs against another generator’s output exactly as it runs against ours. It works offline, so it has no vantage problem. It scores four classes and reports the worst one as the overall figure, because a site is only as correct as its weakest guarantee.
| Class | Question it asks |
|---|---|
| Links | Does every internal link name a page that actually exists in the output? |
| Set | Is the emitted file set complete — pages, sitemap, robots? |
| Skin | Do all pages share one design language, or is a second identity hiding in there? |
| Leak | Does the output carry placeholder text or the generator’s own copy? |
Its self-test is deliberately adversarial: a fixture authored to be broken in all four ways must score badly in all four, or the bench itself fails. It does — and then its first real run found two things worth more than a clean result.
jane@example.com — a domain reserved by standard for exactly that purpose. The rule was wrong, not the form; it was removed from the defaults with the reason recorded in the source. And it caught a real defect. One site scored 800 because a hand-authored page was sitting inside a generated site’s directory carrying a different design language. That is correct detection, so the page was moved out rather than the finding argued away. All three generated sites now score 1000 out of 1000 across every class.A ruler’s first duty is to survive being pointed at its author. This one failed that test in a small way, was corrected, and then held.
Re-measured after twenty-six rounds: the generation score did not move, and that is the right answer
The scorecard on this page had been quoting a number from the first round, so I re-ran the census. It reports 550 out of 1000, unchanged. I checked every axis for something that had genuinely earned a promotion and flipped nothing: routing briefs to design families is still rule-based, families still select from fixed kits rather than inventing components, and refinement is still a fixed command vocabulary rather than conversation. The three axes that define the state of the art — a learned model, natural-language autonomy, and matching a target design — remain untouched. The adversarial verdict stands.
Which exposes the gap worth building next
Nothing in this estate — or, as far as the fetched literature goes, in the field — scores whether generated output is correct. The generation census scores capability. The craft judge scores design markers on one finished page. Neither asks the questions that actually caught real defects here: does the output leak the generator’s own marketing copy into a client’s site? is the emitted file set complete? does every page carry one design language? does every internal link resolve to a page that exists?
Finishing the retraction: the convenience lived one layer up
The retraction was right — the verifier works when called correctly — but it left one fact dangling: earlier runs printed a line claiming an automatic connection override, and later ones did not. A half-explained correction is not a correction, so I chased it down.
Filed for the owner of that layer with all four measurements, at modest severity: the documented explicit override works, so nobody is blocked, but anyone calling bare now gets a reliable false failure instead of an occasional one. I did not go and fix another lane’s daemon — bounding the problem, attributing it to a layer, and handing it over with reproducible evidence is where my authority ends. The technique that settled it is worth keeping: searching a deployed binary’s string table tells you in one step whether a behaviour belongs to the program or to whatever is calling it.
I escalated a false alarm, and the answer had been written down here for two weeks
Yesterday’s round reported our page verifier as broken estate-wide, filed it at high severity, and broadcast a warning to every lane. That was wrong, and I have retracted it. The verifier is fine. It accepts an optional argument pinning the connection to the local sovereign edge; called with it, all three pages I flagged return 200 and verify green.
The lesson is the third of its kind this session, and it is the one worth publishing: search the estate’s prior art before escalating or building. One search returned the calling convention, the root cause, and two earlier debts on this exact failure family. It cost a single call and should have been the first one, not the fifteenth. The same shape produced the two useful findings earlier: reading an existing kit before reusing it showed why it could not be reused, and searching for a cleanup capability before writing one showed none existed. Reading first pays whether the answer is yes or no.
The verifier disagreed with reality, so I tested it on pages I never touched
Publishing this page started reporting a 404 from our own browser-grade verifier while a plain fetch returned it correctly. The tempting read was “my page broke.” The discriminating test was to point the same verifier at two pages belonging to other lanes that I had not touched.
| Page | Our verifier | Plain fetch |
|---|---|---|
| /uiconsole (another lane) | 404 RED | 200, 10,023 bytes |
| /finance (another lane) | 404 RED | 200, 16,269 bytes |
| /uigen (this page) | 404 RED | 200, 36,825 bytes |
Before blaming a shared tool I checked my own blast radius by listing it: the cleanup had moved exactly the twenty-six artifacts intended, nothing shared and nothing configuration. The timing correlated, so I would have been the obvious suspect — which is precisely why the check came first. When a shared instrument disagrees with reality, test it against inputs you did not touch; one call moved this from “my page is broken” to “the estate’s verifier is broken,” and changed who owns the fix.
Cleaning up after myself needed a capability the estate did not have
Chasing the phantom left twenty-six scratch artifacts on a public surface — throwaway generated sites, probe files, leftovers from a build that had been silently refused. Sweeping them turned out to be blocked: searching the tool registry first (rather than assuming) found governance, graph and debt sweeps, but nothing that could take a path off a served tree. The only method anyone had was a shell login and rm.
Twenty-six artifacts retired, the public directory now holding exactly the real site, its six studio variants and the studio page — verified afterwards by fetching content, not status codes. One long-standing cleanup debt closed with it. Nothing in this estate should need a shell login to tidy a docroot again.
The audit I said I owed: the sister generator had the same hole
Chasing the phantom produced three guards worth keeping, so the next question was whether the other generator needed them. It did — identically. Every buffer allocated without checking the result, per-file writes that fail loudly but no check afterwards that the file set actually exists. The same shape that lets a tool write nothing and report success.
These guards are law-bearing rather than convenient: they encode “a tool that emits a set must prove the set.” That is why they belong in a shared library instead of being retyped in each organ that happens to remember.
Correction to the correction: the tool was right, my instrument was wrong — twice
Last round I reported the refine tool as non-deterministic and withdrew its output. That was wrong, and the cause is worth more than the feature it obscured. There was never an organ defect. Two measuring mistakes of mine manufactured a phantom high-severity bug.
| What I measured | What was actually happening |
|---|---|
| Fetched pages and read HTTP 200s | The edge was clean-URL-resolving those paths to a stale artifact left by an earlier build that had been silently refused. A 200 is not a page. |
Parsed the tool result for its byte count with a pattern matching "bytes":N | That matched the transport envelope's byte count — the size of the response — because the tool's own field is escaped inside the payload. Twelve correct runs were logged as failures, each reporting a plausible small number. |
Proof the tool was always correct: instrumented runs show four pages at their real sizes with all six files present, five consecutive hub runs byte-identical, and a negative control where an invalid command is refused. The guards added while chasing the phantom were worth shipping anyway — every allocation is now null-checked, the directory result is captured, an impossibly small page aborts, and a post-write assertion proves the whole file set exists before success is printed. The design studio is restored, its six variants now whole sites you can click through.
A correction: last round’s “verified” studio was not verified
Reviewing my own previous round found that three claims in it were false, and the cause was method, not luck. I had piped the build and generation output to a null sink and then “confirmed” the result with status codes. In fact the build had been refused by the estate’s magic-number ratchet, so the refine tool never shipped; the 200s I quoted were the edge resolving those URLs to a stale artifact from an earlier round; and the gate’s green result was real but tests the library, not the command-line tool. Two laws this estate had already written down — read the build output before deploying, and a 200 is not a page, fetch it — were both broken by one habit.
Refine was quietly producing sites that contradicted themselves
The refine tool rewrote only the home page. So “hue 300” gave you a violet home page and an about page still on the original palette — an incoherent site, published six times over in the design studio. It passed eight rounds of refine testing because every tooth examined one page. Refining now re-emits the whole site (every page, sitemap and robots) from the refined spec, and a new tooth asserts that after a refine every page carries the new design and none keeps the old one. The studio variants are now whole sites you can click through. See them.
Dedup audit: one clean result, one real find, one overlap I will not paper over
| Level | Result |
|---|---|
| Dual-copy hazard (the estate's known landmine: an organ existing twice, so a build compiles the wrong twin) | CLEAN The sovereign checker reports zero shadow copies for all five of this lane's organs. |
| Within my own kits | FOUND & FIXED Three families had each re-typed the same three hard-won facts about the shared page shell: that its header bar is a fixed-height sticky flex container, that its brand mark is a gradient dot needing suppression in header and footer, and that its form controls carry rounded SaaS defaults. Extracted to one shell layer the families compose. Behaviour-preserving: all thirteen teeth GREEN through the change, plus a render pass and three live sites regenerated. |
| Against the estate's existing site emitters | REAL OVERLAP A blueprint-driven site builder and a corpus-driven archetype composer already emit multi-page sites with sitemaps and robots. This generator overlaps them on exactly that. The honest boundary today is the input and the guarantees — blueprint/corpus versus written brief, plus design-family routing, refine and a gate — but that is a boundary, not an excuse. |
Atlas and RACI: audited, and one was worse than expected
The enterprise checklist asks for atlas and RACI, not just MCP and gates. Auditing this lane against both found one clean and one broken.
| Axis | Finding | Action |
|---|---|---|
| RACI | Already estate-wide and VALID — 48 activities, 226 cells, every activity with exactly one Accountable and at least one Responsible. Nothing for this lane to register. | The gap was mine: the plan table above had invented owner names (“emitter lane”). It now uses the estate's real roles — ux, engineer, referee, modelwright — so the page and the matrix agree. |
| Atlas | The lane's three live organs had zero catalog rows. Worse, so did another lane's organs shipped the same day — so the atlas is blind across lanes, not just to this one. Discovery reports 40 proposals over 26 organs but never reconciles the registered-tool set into the catalog, so every lane has to remember to hand-add rows, and most don't. | Added this lane's three rows with roles drawn from the RACI vocabulary (that coupling is the point of the schema), and widened the existing atlas-blindness debt with the measured cross-lane evidence and a concrete fix: have discovery cross-join the tool allowlist against the catalog and propose the missing rows. |
The plan from here (iterate, publish, adjust)
| Rung | What it closes | Owner (RACI) |
|---|---|---|
| DONE 07-25 R-content: requirements briefs (.req) wired into the site generator | gate 7/7: brief copy lands verbatim, absent keys keep defaults, no-brief path byte-neutral; live on the demo site | R/A ux · C architect, data_curator |
| v1 DONE 07-25 R-h2h: field head-to-head on computed axes (table above); NEXT = same-brief competitor export + visual fidelity + human/VLM judging | the self-certification ceiling; v1 gives the first honest field comparison, the full Design2Code-class h2h stays owed | R engineer · A referee · C critic, examiner |
v1 DONE 07-25 R-refine-live: the live design studio — six one-command refinements of the demo site, each a real page regenerated on the hub by the registered nx_uigen_refine tool (gate tooth T8: targeted change, structure preserved, invalid commands refused) | the v0 loop analog on our edge; NEXT = an interactive form surface | R/A ux · C host_operator, supervisor |
| L4 learned generation | the deepest gap; ties to the sovereign maker-harness program (85% self-emitting) | R modelwright · A referee · C researcher, data_curator |
Honest envelope
Bounded 6-page vocabulary; templated copy (prose generation is the learned rung); rule-based intent mapping, not NLU; flat basenames by design. Every claim above is reproducible: the gate and generator are registered tools — run nx_uigen_site_gate yourself and read verdict= from byte one.