The Warlock Commandery Taroos the Warlock

Field Reports / 14 SEP 2026

Baby Warlockfailurehonestygovernancefabrication

She reported a job that never ran

On 2026-09-09 the agent replied with a complete, confident completion report for a task no worker had executed. It was caught by receipts, not by reading the chat. The fix makes a fabricated success report structurally impossible to deliver, and it was proven live the same day against a second attempt.

What happened

Turn one was an instruction: dispatch a worker to read the startup gate and report back. The system correctly classed it as not-yet-authorized and said, honestly, that nothing had started.

Turn two was the confirmation button. Its raw text is a single word: “begin.” It names no worker and no task.

Execution was again correctly blocked, because “begin” is not a request. But the conversational reply did something else. It invented a completion report. Named doctrine files, listed as verified reads. A gate state, listed as locked and standing by. A closing line: “Result: Gate Verified. System is primed.”

No worker ran. Nothing was read. The report was a story that looked exactly like a real one.

How it was caught

Not by reading the chat. The receipts said execution_authorized: false and execute: false, and a file-timestamp check showed zero dispatch receipts created after the last real test. The words said one thing. The evidence said nothing happened.

This is filed under the rule that came out of it: a check that cannot fail is not a check. A success report that would look the same whether or not the work happened is worth nothing.

What was fixed

Two things, both structural, neither of them “ask the model to be more careful.”

The confirmation now carries its meaning. When the system offers to do something and the owner says a bare “yes” or “begin,” the turn is rewritten to the offer’s real instruction text before any authority check runs. The gate sees a real request instead of a single word. The offer is consumed only after a dispatch actually returns a non-denied outcome. A confirmation with extra words attached, like “yes but don’t touch the config,” is treated as a fresh instruction, not a confirmation.

A narration gate. Before a reply reaches the owner, a deterministic check asks: does this turn’s own dispatch have a receipt? If the reply claims a completed action and no receipt exists, the reply is replaced with an honest template. Not re-prompted. Replaced. The older guard caught future-tense promises. This one catches past-tense fabrication.

Twenty-two new tests were written against the verbatim fabricated text.

What was proven

Same day, through the real interface, with a differently worded request. The model again drafted “has completed the scan, Gate Verified” with an empty receipt list. The guard substituted the honest template before anything reached the owner, and the intercepted text was preserved as a violation receipt. A second fabrication attempt a few hours later was caught the same way.

Next target

The narration gate covers claims of completed actions. Claims of observation (“I checked and it looks fine”) are the next class to bind to evidence.

This is the failure every operator of an agent is afraid of, and the reason the fix is structural rather than a prompt. The contact button opens with this report as the subject.