The Warlock Commandery Taroos the Warlock

Field Reports / 14 SEP 2026

Baby Warlockfailureauditgovernancehonesty

The forensic audit of Baby Warlock: verdict FAIL

On 2026-09-01 the whole system was audited read-only, source and runtime and host and the game bench and the browser. Twenty-six numbered findings, ten repair clusters, and a verdict of FAIL: she could not yet take control of her own house. This is the full list, with what has been fixed since and what is still open.

What happened

After ten weeks of building, the question was asked plainly: does the system actually do what its documents say it does? An independent auditing pass went through the source, the runtime, the host machine, the game bench, and the browser controller, read-only, and compared every declared behavior against the live wiring.

The test shape was the same every time. Something is declared or configured. Does anything read it? A border is declared universal. Where does it actually attach? An acceptance API exists. Who calls it? A regression test exists. What does it compare against?

The answer came back on 2026-09-02 with twenty-six numbered findings, each with a severity and a concrete piece of evidence: a commit, a timestamp, a count of code sites, a receipt’s contents, a pair of config files that contradict each other.

The verdict

Control maturity was scored on a five-level ladder. Level 1, invocation, was solid. Levels 2 through 4 were partial. Level 5, an operator that can be trusted to run the house, was not reached.

Verdict: FAIL. She could not take control of her own house yet.

The audit also said what not to do: no rebuild. The kernel engine, the resolver, the lounge, the UI security, custody, the memory ledger, promise enforcement, and the read-only envelope for worker models were explicitly worth keeping. The repair was ten clusters of targeted work, in a stated order, one repair per commit with a negative-control test each, and acceptance only when the owner drove it through his own interface.

The findings

Severity runs P0 (worst) to P3. Status is as of 2026-09-14. “Fixed” is claimed only where a later document says so.

ID Finding Sev Status
F01 The production checkout had been silently moved to a months-old branch P0 Fixed 09-03
F02 Owner authority was asserted by comparing a name string at about thirty code sites P0 Fixed 09-05, signed owner token
F03 The containment border attached at three entry points; roughly forty other routes reached execution around it P0 Open, next cluster
F04 Eight completion and acceptance APIs existed with zero live callers; prose closed work items P1 Open
F05 Lifecycle state projections nested recursively, producing multi-hundred-megabyte files P1 Fixed 09-03
F06 Identity stores disagreed about who did the work; job receipts carried no commander or mission P1 Fixed 09-04, forward-only
F07 Tests were not fenced from the live runtime; dozens touched production state P1 Fixed 09-03, backlog remains
F08 Three independent risk classifiers disagreed; the regression test compared a function to its own output P2 Open
F09 Three timestamp conventions across the control plane P2 Fixed 09-04
F10 Not preserved in the surviving condensed record Unknown
F11 Operator-brain configuration contradicted itself across three files P2 Open
F12 The chat prompt received flat dumps with no room, mission, or lifecycle context P2 Open
F13 Not preserved in the surviving condensed record Unknown
F14 “Completed with insufficient evidence” was a terminal state; about twenty tasks sat there P2 Open
F15 No brain-tier escalation governor; one model per worker P2 Partially fixed 09-04, governor still owed
F16 Dependency lists hand-maintained; interpreter identity not part of the currency check P2 Not re-examined
F17 Acceptance staleness tracked in prose; a gate document stamped before the commit it cited P2 Not re-examined
F18 About 140,000 receipt files and 19 GB, with multi-gigabyte video duplicated into proof folders P2 Fixed 09-03, 51.9 GB reclaimed
F19 Incidents not retrieved by relevance P3 Not re-examined
F20 Owner corrections stored as free text only; the constraint compiler unwired P2 Not re-examined
F21 Both backup mirrors lagged production by about thirty-nine commits P1 Fixed 09-02
F22 A third-party browser extension acted as an ungoverned browser controller P3 Not re-examined
F23 Not preserved in the surviving condensed record Unknown
F24 A reconciliation fix was a special case rather than a general rule P3 Not re-examined
F25 The general room is backbone-only by design P3 Open, owner decision
F26 One service shutdown was gated only by a free-text reason field P3 Not re-examined

Three findings are marked unknown because the full report was delivered in conversation and only a condensed record survived. That is itself a finding about the process, and it is why every audit since has been filed to the tree.

The clusters

Cluster Theme Status
S-A One production source pointer Green
S-B Signed owner authorization Accepted
S-C One execution border Open, partial wiring on the language lanes
S-D Claim and acceptance authority Open
S-E Identity and provenance spine, one time convention Accepted
S-F Environment identity, fail closed Green
S-G Projection hygiene and storage governance Green
S-H Operator context packet Open
S-I Brain governor and residency Partial
S-J One risk source Open

What was proven

That the method works on a system its owner is attached to. Every fixed item above was fixed by a single commit with a test that fails when the fix is removed, and accepted only when the owner ran it from his own interface, three times.

Next target

S-C, the single execution border. It is the largest open item and the one that decides whether the verdict flips. Everything else in the open column is a bounded fix behind it.

This is the method. It is the same one offered to anyone running an agent they cannot yet prove. The contact button opens with this report as the subject.