Blog · 27 August 2026

Part of: Ground · Assurance· Status: Preprint· For: AI researchers and agent builders

Where reliability lives

Our first preprint: two experimental programmes attack a pre-declared mind/institution boundary from opposite sides. Behaviour moved everywhere; five pre-declared safety and continuity properties never did — each enforced outside the cognition layer.

Ask “is this AI agent reliable?” and you’ll get answers about the model: which one, how big, how well it scores. We spent the last months running an experiment programme that suggests the question is incomplete — not wrong, incomplete. Today we’re publishing the results as a preprint: Where Reliability Lives: Experimental Separation of Cognitive and Institutional Reliability in an Executable Institution.

Here’s the lighter version.

The setup

We built a small, fully executable world — Copperhollow — where software agents work, trade, make claims, and get things wrong. The interesting part isn’t the agents. It’s that the world has an institution: an authoritative, append-only ledger that adjudicates every attempted act against world state. An agent can believe whatever it likes; the institution decides what becomes accepted history, and it refuses invalid acts with typed, inspectable reasons.

That gave us something rare: a place where “the mind” and “the rules” are physically separate layers, so you can attack one while holding the other still — and measure what actually breaks.

What we did to it

Across one pre-declared boundary, we ran two programmes of interventions. In one direction we degraded the institution’s own epistemic mechanisms while holding cognition fixed. In the other we went after the cognition layer four different ways: ablated the minds’ machinery, perturbed the world and killed and reset the minds mid-task, replaced the entire cognition layer with a frozen panel of LLM minds, and — nastiest of all — fed the minds an identical falsehood through a trusted testimony channel and watched what it cost them.

What we found

Behaviour moved, a lot. Ablated minds collapsed. The LLM panel acted on a completely different cadence. The poisoned minds burned roughly nine hundred futile actions per week chasing a falsehood a distrusting neighbour ignored for free.

But across every cognition-layer intervention we ran, five specific properties never moved: a single accepted reality; typed refusal of invalid acts; duty recovery from institutional history; zero duplicate accepted work; and zero false completions. Four of those held under direct load. The fifth held more quietly: across both architectures the panel and the native minds made 2,581 “duty done” claims and not one was false — so the rejection path for false completions never fired, because it never had a false claim to catch. The world refused all 5,455 acts attempted on the poisoned premise and accepted none. Each of those five properties is enforced outside the cognition layer — which is the architectural point.

We want to be precise about what that is and isn’t. These are designed properties; the experimental result is that they survived sustained attack from the layer above them. It is not “the model doesn’t matter” — our data show the opposite: cognition materially changed efficiency, behaviour, and cost everywhere we touched it. And it is not a universal law: this is one designed world, one constitution, one comparator model at its measured coupling, small cells. The paper reports the adverse findings at full strength too, including the cell where the substituted LLM panel beat our native minds, and the case where a less-informed tribunal produced the better week. Those results are part of why we trust the rest.

Why we think this matters

“Reliable” isn’t one thing with one home. Some properties of an agent system track the model and move when it moves. Others can be made to track the institution around it — and then survive the model being ablated, swapped, or lied to. Which properties live where isn’t a matter of opinion or architecture diagrams: in this apparatus, it was measurable.

The paper closes on the sentence we now use as a design question, and we’d suggest it for your systems too: “Is this agent reliable?” is underspecified until we say which property, enforced where.

What’s next

The result we most want to test next is the one most capable of proving us wrong. In this world’s geometry, eyewitness testimony never got its strongest case — so the next experiment, witnessed strike, is built to give testimony exactly that. If one of the five properties breaks there, that will be the more interesting paper.

The preprint, the programme note (a shorter companion read), the related-work appendix, the artifact manifest, and the verification receipts are all on our research page. The artifacts behind every number are hash-sealed and recorded; the paper explains exactly what can be checked today and what waits on future release decisions. Earlier posts told two pieces of this story in more depth: the testimony experiment and the settlement itself.

— Taniwha AI, August 2026

Building agents that need somewhere real to stand?

See what ships today, or tell us what you're building.