What happened
After ten weeks of building, the question was asked plainly: does the system actually do what its documents say it does? An independent auditing pass went through the source, the runtime, the host machine, the game bench, and the browser controller, read-only, and compared every declared behavior against the live wiring.
The test shape was the same every time. Something is declared or configured. Does anything read it? A border is declared universal. Where does it actually attach? An acceptance API exists. Who calls it? A regression test exists. What does it compare against?
The answer came back on 2026-09-02 with twenty-six numbered findings, each with a severity and a concrete piece of evidence: a commit, a timestamp, a count of code sites, a receipt’s contents, a pair of config files that contradict each other.
The verdict
Control maturity was scored on a five-level ladder. Level 1, invocation, was solid. Levels 2 through 4 were partial. Level 5, an operator that can be trusted to run the house, was not reached.
Verdict: FAIL. She could not take control of her own house yet.
The audit also said what not to do: no rebuild. The kernel engine, the resolver, the lounge, the UI security, custody, the memory ledger, promise enforcement, and the read-only envelope for worker models were explicitly worth keeping. The repair was ten clusters of targeted work, in a stated order, one repair per commit with a negative-control test each, and acceptance only when the owner drove it through his own interface.
The findings
Severity runs P0 (worst) to P3. Status is as of 2026-09-14. “Fixed” is claimed only where a later document says so.
| ID | Finding | Sev | Status |
|---|---|---|---|
| F01 | The production checkout had been silently moved to a months-old branch | P0 | Fixed 09-03 |
| F02 | Owner authority was asserted by comparing a name string at about thirty code sites | P0 | Fixed 09-05, signed owner token |
| F03 | The containment border attached at three entry points; roughly forty other routes reached execution around it | P0 | Open, next cluster |
| F04 | Eight completion and acceptance APIs existed with zero live callers; prose closed work items | P1 | Open |
| F05 | Lifecycle state projections nested recursively, producing multi-hundred-megabyte files | P1 | Fixed 09-03 |
| F06 | Identity stores disagreed about who did the work; job receipts carried no commander or mission | P1 | Fixed 09-04, forward-only |
| F07 | Tests were not fenced from the live runtime; dozens touched production state | P1 | Fixed 09-03, backlog remains |
| F08 | Three independent risk classifiers disagreed; the regression test compared a function to its own output | P2 | Open |
| F09 | Three timestamp conventions across the control plane | P2 | Fixed 09-04 |
| F10 | Not preserved in the surviving condensed record | Unknown | |
| F11 | Operator-brain configuration contradicted itself across three files | P2 | Open |
| F12 | The chat prompt received flat dumps with no room, mission, or lifecycle context | P2 | Open |
| F13 | Not preserved in the surviving condensed record | Unknown | |
| F14 | “Completed with insufficient evidence” was a terminal state; about twenty tasks sat there | P2 | Open |
| F15 | No brain-tier escalation governor; one model per worker | P2 | Partially fixed 09-04, governor still owed |
| F16 | Dependency lists hand-maintained; interpreter identity not part of the currency check | P2 | Not re-examined |
| F17 | Acceptance staleness tracked in prose; a gate document stamped before the commit it cited | P2 | Not re-examined |
| F18 | About 140,000 receipt files and 19 GB, with multi-gigabyte video duplicated into proof folders | P2 | Fixed 09-03, 51.9 GB reclaimed |
| F19 | Incidents not retrieved by relevance | P3 | Not re-examined |
| F20 | Owner corrections stored as free text only; the constraint compiler unwired | P2 | Not re-examined |
| F21 | Both backup mirrors lagged production by about thirty-nine commits | P1 | Fixed 09-02 |
| F22 | A third-party browser extension acted as an ungoverned browser controller | P3 | Not re-examined |
| F23 | Not preserved in the surviving condensed record | Unknown | |
| F24 | A reconciliation fix was a special case rather than a general rule | P3 | Not re-examined |
| F25 | The general room is backbone-only by design | P3 | Open, owner decision |
| F26 | One service shutdown was gated only by a free-text reason field | P3 | Not re-examined |
Three findings are marked unknown because the full report was delivered in conversation and only a condensed record survived. That is itself a finding about the process, and it is why every audit since has been filed to the tree.
The clusters
| Cluster | Theme | Status |
|---|---|---|
| S-A | One production source pointer | Green |
| S-B | Signed owner authorization | Accepted |
| S-C | One execution border | Open, partial wiring on the language lanes |
| S-D | Claim and acceptance authority | Open |
| S-E | Identity and provenance spine, one time convention | Accepted |
| S-F | Environment identity, fail closed | Green |
| S-G | Projection hygiene and storage governance | Green |
| S-H | Operator context packet | Open |
| S-I | Brain governor and residency | Partial |
| S-J | One risk source | Open |
What was proven
That the method works on a system its owner is attached to. Every fixed item above was fixed by a single commit with a test that fails when the fix is removed, and accepted only when the owner ran it from his own interface, three times.
Next target
S-C, the single execution border. It is the largest open item and the one that decides whether the verdict flips. Everything else in the open column is a bounded fix behind it.
This is the method. It is the same one offered to anyone running an agent they cannot yet prove. The contact button opens with this report as the subject.