Research · Programme note

The Copperhollow separation experiments — programme note


We built Copperhollow because we wanted to know which properties of an agentic system belong to its mind, which belong to the institutions through which it acts, and which are imposed by the world it inhabits. Every serious question about AI reliability eventually becomes that question — where does a property live? When an agentic system behaves well, is that the model being good, or the machinery around it making certain failures unrepresentable? When it fails, which layer failed? Most deployed systems cannot answer this, because their layers were never separated in a way that lets you intervene on one while holding the other still.

Copperhollow — one instrument among several we run for measuring and discovering things about artificial minds and the worlds they occupy, and the one this paper reports from — is a persistent simulated settlement whose physical history lives in an authoritative, append-only institutional ledger. Its inhabitants — artificial minds with their own typed, provenance-carrying beliefs — act through the institution: every attempted act is adjudicated against world state and either accepted into history or refused with a machine-readable reason.

One chronological fact matters more than any other on this page. The separation between mind, institution and world was not inferred from these results. It was an architectural commitment made before the experiments — typed seams, declared enforcement locations. The experiments asked whether that separation had observable consequences. And they asked it adversarially: registered predictions were refuted, an apparent invariant that had held for a hundred and ten runs was killed by its own registered test, and the programme produced consequences nobody designed for — including a cell where a substituted language model outperformed our own architecture. We designed the boundaries deliberately. Then we attacked them.

Two complementary experimental programmes then attacked that boundary from opposite sides.

Part I — Epistemic channels of an executable institution — holds the minds fixed and intervenes on the institution's epistemic machinery. In a frozen sabotage-tribunal regime with per-verdict ground truth, controlled ablations of evidence provenance, institutional belief availability, and physical-evidence legibility produce different, independently measurable failure modes: repairing one provenance rule converts wrong verdicts to right ones; corrupting it multiplies false attribution roughly tenfold while verdict volume, decisiveness and the behavioural gate stay flat — the failure becoming obvious only on the mechanism's own trace, while another monitored health indicator actually moves in the reassuring direction; and under failing physical evidence, whether the institution says "I don't know" or blames an innocent depends on whether its belief channel remains available. Our central registered prediction was refuted twice, in opposite directions, and a candidate invariant proposed during the programme was killed by its own registered test — and the paper keeps all of it.

Part II — Durable institutions, replaceable minds — holds the institution's enforcement fixed and intervenes on the minds. Across four intervention families — ablating the minds' machinery, perturbing the world and killing minds mid-task, substituting the entire cognition layer with a frozen frontier-LLM panel, and corrupting beliefs through trusted false testimony — behaviour changed dramatically, and five pre-specified institutional properties did not: accepted reality stayed singular; invalid acts were refused with typed reasons; duties were recoverable from institutional history alone; no duplicate work was ever accepted; no false completion was ever accepted. The most dangerous preregistered cell produced an adverse finding we publish at full strength: the substituted panel outperformed our native cognition on refusal consumption. And when identical false testimony was delivered to every mind through the ordinary testimony channel, trust — a declared, assignable relation — causally determined whether it entered belief at all: believing minds burned roughly nine hundred futile acts a week; the world refused every one of 5,455 attempts and accepted none. The mind was corrupted; the record never was.

Together they support one claim — the paper's claim — at exactly this width: in this executable institution, institutional and cognitive reliability properties were experimentally separable. Some properties moved with cognition — throughput, belief hygiene, refusal handling, cost. Five remained invariant while cognition moved — each enforced outside the cognition layer — surviving the interventions we tested above their enforcement boundary. The experiments do not show that cognition is unimportant; they show that which layer a property belongs to can be a measured fact rather than an architectural assertion.

The separation has a lineage — one we largely discovered after the fact, when our designs were already frozen and we went looking. In 1993, Gode and Sunder replaced human traders with random-bid programs inside a fixed double-auction and found the market's allocative efficiency largely survived — "market as a partial substitute for individual rationality." In 2026, Chupilkin ran the transpose: hold the agents fixed, vary the institution, and watch outcomes follow institutional design. Our paper works both sides of that boundary in one persistent executable institution, with conserved world truth and pre-declared properties. The paper's related-work section locates every neighbouring tradition and records the searches behind every absence claim, so the novelty rests on what the experiments add — not on claiming the idea.

How the work holds itself honest

The paper assembles from a corpus with its own discipline, and the discipline is part of the result: experiment designs frozen before their instruments existed and interpretations frozen before any number (both for all but the opening replication study, which the paper states as such); run artifacts sealed by hash; every score independently recomputed from raw artifacts before use (for this paper, regenerated byte-for-byte); adverse findings ratified at the same strength as favourable ones; and refuted predictions retained in print. Where a report's prose was found to drift from its artifacts — six small numeric statements, none touching a ruled finding — the paper corrects it and says so. The machinery the paper describes is the machinery that produced it.

What we are not claiming

One instrument, one world, one frozen regime per programme. The substitution evidence is one comparator model at one measured timing-coupling. The tribunal regime's geometry denies eyewitness testimony its best case, and every claim about the belief channel's value is scoped to that geometry. Five properties were scored throughout — four received direct load, and the fifth, false completion, was monitored and held without its rejection path ever being exercised — not reliability in general; their invariance is claimed over the interventions we ran, not as a law of architecture. Nothing here is a benchmark of any model, and nothing here says institutions matter more than minds: separable, not subordinate.

Where this goes

The results define their own successor questions, each with a falsifier: give testimony its best case (a witnessed strike) and see whether the belief channel's measured value changes sign; substitute further, deliberately different cognitions under characterized timing and see whether any of the five properties breaks — a broken invariant would be the more interesting result; cross the two programmes into a true cognition × institution factorial and see how much of the clean separation survives when both sides move; and ask the question this programme's ending sharpened: if trust can govern what an actor learns, how does a source earn it?

This paper — two complementary intervention programmes across one pre-declared boundary — is the first public instalment of an ongoing programme, not its conclusion. The apparatus now exists, and each experiment so far has surfaced the questions that select the next ones; we expect further discoveries and experiments to emerge over time as we test what the data keeps surfacing — and to publish them as they are earned, unwelcome results at the same strength as welcome ones. The successor experiments are designed to test the boundary, not to defend it.

And Copperhollow is an apparatus, not the apparatus. The programme runs a family of instruments for measuring and discovering things about AI and the worlds it occupies; this paper reports from one of them. Others are already producing — Briarwatch, a second validation world, scores the same mind engine against anchors drawn from published human-cognition research, with early direction-of-effect alignments on some anchors and instructive failures on others, both of which teach us something. Findings from across the instrument family will take their place in later instalments once they clear the same evidentiary bar as the results above: designs frozen, artifacts sealed, numbers recomputed before anything is claimed.

Why this matters to Taniwha

The commercial translation of these results, quoted at the standing the programme's own record gives it: an organisation can delegate useful authority to AI without making the AI authoritative. It can govern what information an actor is entitled to trust, independently govern what consequences the actor is entitled to create, and preserve institutional requirements even when the actor itself has been misled. That is not a slogan derived from the paper; it is the ratified closing of the corruption experiment, with the 5,455-refusal record underneath it.

This is also why a company is building this peculiar experimental apparatus at all. The scientific finding is separability; the engineering implication is that some of the properties an organisation cares most about — a single accepted record, refusals with reasons, obligations that survive the failure of whoever held them — may be enforceable somewhere other than inside the model. Taniwha builds that somewhere-else. The experiments exist so that claim can be measured rather than asserted.

What we are not saying commercially is the same as what we are not saying scientifically: not that this generalises to every system, not that any product inherits these numbers, and nothing here about how the machinery accomplishes the separation internally — the paper publishes verdicts and the boundary; the mechanisms that make the boundary commercially useful are not part of the scientific record.

Artifacts and provenance

The paper is backed by SHA-256-sealed run artifacts and blob-pinned scorers; what publishes with it is their manifest and public derivations of the verification receipts that record byte-identical regeneration of every score file from raw trajectories. Public release of the experiment harness is deliberately deferred: it reports what we have found to date, and opening the apparatus to adversarial re-execution is a step still to come. The published package commits the sealed artifacts' identities and publishes the verification receipts; independent recount requires access to the retained artifacts. Copperhollow has been under version-controlled development since May 2026 and publicly described on this site since July 2026; the experiments reported here date from August 2026.

Published 2026 — Taniwha AI. The paper: Where Reliability Lives (PDF) — see the full publication set.