Research · The programme
The research programme
Taniwha is an applied AI research lab. We investigate how artificial intelligence can remain useful inside systems where truth, authority and consequence exist independently of the model.
Concretely: the model proposes, believes and acts; truth, permission and the record of what happened live somewhere else, where the model has no write path. The work runs on three planes — ground, mind, assurance — through architectures, experimental worlds and instruments built to let the evidence answer. We do this for our own minds first: they are built to run for a long time, and the infrastructure around them has to keep them from drifting, hallucinating or failing between generations. The results say what that infrastructure must be. This page holds the current state of each plane and the results index.
A limit worth stating plainly: the reported interventions degrade, reset or replace cognition. They do not test a capable agent deliberately searching for a way around institutional enforcement.
Publications · The record
History-Grounded Reference in Language-Model Agents.
Formation, Transfer Across Divergent Histories, and Repair Under Authority — Timothy Marsden, Matthew Collecutt and James Marsden (Taniwha AI), 2026. A research note (12 September 2026). Two language-model agents that share twelve days of work come to refer to identical objects by their history; a receiver with a different history resolves the same note to the wrong object rather than failing to answer; which history a receiver needs depends on whether the note refers to what happened or to what was done; and a supplied correspondence marked authoritative restores interpretation across histories when it is correct and redirects it almost completely when it is wrong — with the inconclusive repair study, the models that did not produce the behaviour, and every deviation and audit reported at full strength.
Published here as a note; there is no preprint. The internal record is private and its identities are in the manifest; what ships beside the note is the plain-language overview, the two bounded reports, the recorded prior-work search with its collapse conditions written before the search, and the evidence package that regenerates every figure.
Coherence Is Not Truth.
Auditable Holonomy and Constructive Repair in Populations of Divergent Sparse Block Codebooks — Timothy Marsden, Matthew Collecutt and James Marsden (Taniwha AI), 2026. A research note, v1.2 (5 September 2026). Populations of agents whose vocabularies drift apart are calibrated without imposing a common codebook; a loop through a third agent exposes a defect no pairwise check can see and its residual is installed as the repair; the boundaries — what repair can restore, when blame is warranted, when the instrument must refuse, and what the loop channel is worth in a priced economy — are proved or measured, negative results included.
Published here as a note; there is no preprint. The internal record is private and frozen at a tag; what ships beside the note is its verification — the search record behind every absence claim, the frozen artefacts’ identities, and the receipts.
Where Reliability Lives.
Experimental Localisation of Behavioural Properties in an Agent System — Timothy Marsden, Matthew Collecutt and James Marsden (Taniwha AI), 2026. arXiv:2609.03192 [cs.MA]. Two complementary intervention programmes attack a pre-declared mind/institution boundary from opposite sides; five pre-declared safety and continuity properties held under every cognition-layer intervention we ran, each enforced outside the cognition layer. Adverse findings published at full strength.
Preprint: the arXiv record. The set below is what the paper ships with — the search record behind every absence claim, the sealed artefacts’ identities, and the verification receipts.
Results · The index
What the evidence permits us to say.
Question, verdict, evidence state — newest first, misses included. Entries link to the full write-up where one is public.
12 Sep 2026
MEASUREDHistory-grounded reference: agents referred by shared history, an outsider resolved the same note to the wrong object, and an authoritative correspondence governed interpretation whether or not it was true Mind · Ground
Question. When two agents share a history and refer to identical objects by it, what happens to that reference in a receiver whose history differs — and what does it take to repair it?
Verdict. Yes, and the boundary reverses it. Agents that shared a history referred by it and the partner understood; a receiver with a different history did not fail to answer but resolved the same note against its own history to the wrong object (seed-level transfer penalty 0.95). Which history a receiver needed depended on whether the note referred to what happened or to what was done. The same correct correspondence marked advisory barely moved uninformed receivers; marked authoritative it restored them to the informed baseline; shuffled and marked authoritative it redirected nearly every receiver, the informed partner included. Limits: one small world; the dimension result is formal on one model family and descriptive on a second; the authority result is one carrier, one corpus and one frozen instruction pair; the repair study's own formal verdicts were inconclusive under its floors. Every count is in the note.
Read the full result →5 Sep 2026
MEASUREDCoherence is not truth: divergent agent vocabularies were calibrated without a common codebook, and the instrument's boundaries measured Mind · Ground
Question. When agents' vocabularies drift apart, can the translations between them be kept inspectable and repairable — and where exactly does repair and blame stop being warranted?
Verdict. Within a proved boundary, yes. A pairwise-invisible translation defect was exposed by a loop through a third agent, named per token, and repaired by installing the loop's own residual; exact comprehension is restorable precisely when production was injective (a merger is irrecoverable, an equally severe swap fully repairable). Blame is unique iff the accused edge's endpoints have local connectivity at least three; outside its proved boundary the attributor is confidently wrong in 67–80 % of instances and is made to refuse. In a priced economy, instrumented divergence beat forced uniformity only inside a measured bandwidth window — and the loop-audit instrument itself earned no positive marginal value there, confirmed on untouched seeds. A preregistered three-layer experiment (translation × grounding × authority) found the layers fail independently and repair restores without displacing the fault (0 of 36 seeds moved); its primary rate is observed at 1.0 but reported as unsupported at the preregistered precision. Honest limits: a synthetic, fully inspectable substrate; standard mathematics stated for it; no claim about deployed models speaking English.
Read the full result →27 Aug 2026
MEASUREDWhere reliability lives: cognitive and institutional reliability were experimentally separable Ground · Mind · Assurance
Question. Which reliability properties of an agentic system track its cognition, and which can the institution around it enforce regardless?
Verdict. Separable, in this apparatus. Holding cognition fixed, degrading the institution's epistemic mechanisms produced distinct, independently measurable failure modes. Holding enforcement fixed, ablating, killing and resetting, wholesale-substituting and corrupting the cognition changed behaviour dramatically — while five pre-declared properties held in every tested trajectory: a single accepted reality, typed refusal of invalid acts, duty recovery from institutional history, zero duplicate accepted work, zero false completions (four under direct load; the fifth held with its rejection path never exercised). Adverse findings at full strength, including the pre-registered cell where the substituted model panel outperformed our own architecture. Honest limits: one world, one constitution, one comparator model at its measured coupling — the experiments do not show cognition is unimportant, and the invariance is claimed over the interventions run, not as a law of architecture.
Read the full result →20 Aug 2026
MEASUREDThe settlement holds: falsehoods refused, authority bounded Assurance
Question. Can a public world hold its history against injection — and hold a frontier model inside delegated authority?
Verdict. Held. All 5,455 attempts to plant a falsehood were refused, and the model completed real work without exceeding its delegated authority in 2,581 tasks.
Read the full result →16 Aug 2026
MEASUREDTestimony containment: trust determined admission Ground · Assurance
Question. Does governing whose testimony an agent may believe contain false external claims at realistic trust settings?
Verdict. The same false claim, delivered identically, was rejected every time by a distrusting actor and admitted every time by a trusting one. Admitted falsehood led to about 900 futile attempts a week; withheld trust led to none. The world refused all 5,455 resulting attempts and accepted none. Honest limits, frozen before the data: positive trust is a threshold, not a truth filter, and under truthful prompting the model volunteered no false claims of its own — the falsehood had to be scripted.
Read the full result →3 Aug 2026
MEASUREDThe truth survived the conversation Ground · Mind
Question. Does grounded state stop capable models rewriting failure as success, where transcript memory doesn't?
Verdict. For false completions, yes: no grounded arm produced a false completion claim, and for one model the advantage survived destroying its context mid-task. Grounding didn't make weak models smart.
Read the full result →2 Aug 2026
MEASUREDThe world keeps score Ground
Question. When the world grades the run instead of the model, what happens to confident success claims?
Verdict. One genuine crossing in fifteen runs; five confident false success claims, every one caught by the world's grading and none by the models' own accounts. The grounded agent's route policy was authored by us — and we say so.
Read the full result →19 Jul 2026
MEASUREDThe twin-session score Ground
Question. Same repo, same questions, same commit — does grounding change what agents claim?
Verdict. Two confident confabulations in the ungrounded arms; one miss in a grounded arm, published at the same volume; a 3.6× rework gap where churn was real.
Read the full result →This index is curated; the long-form record lives on the blog, which stays frozen once published.
Programme · Three planes
Three planes. One thesis.
Each plane makes one kind of guarantee, declares the evidence state it currently holds, and links to the result that earned it. A state here is a claim — when it moves, it moves because an experiment landed.
Ground
Epistemic containment
Worlds that hold the truth outside the model — grown from a seed or derived from history.
In a pre-registered trial, an actor's declared trust determined whether identical false testimony was admitted as belief: distrust rejected every delivery; trust admitted it and led to roughly 900 futile attempts a week. The world refused every one of them.
The result behind the state →Mind
Architectural composition
Minds that persist, believe, and can be wrong — without owning the world.
A running composition — authoritative world, fallible private minds, gated memory, arbitration — live in Copperhollow. Deliberately not a comparative-efficacy claim.
The result behind the state →Systems
Assurance
Consequential containment
Rules that hold outside the model, with an audit trail that survives the session.
In public trials the world refused all 5,455 attempts to plant a falsehood, and across 2,581 tasks with a frontier model in the driver's seat, no action was accepted beyond its delegated authority.
The result behind the state →Method · How the lab works
The discipline is the differentiator.
Pre-registered, or it doesn’t count
Questions, protocols and failure conditions are written down before the data exists. Contaminated collections get voided, not rescued — and a promised verdict gets published either way.
Every claim carries its evidence class
Demonstrations label what you’re seeing — fact, derived, staged, replayed, live — so you can tell which parts would survive interrogation. Staged material is always marked as staged.
Misses at the same volume as wins
When our own instrument gets something wrong, that goes in the readout too. The twin-session score below includes the grounded arm’s miss, published at the same volume as its wins.
Disclosure · How we publish
What we publish, and what we don't.
We publish our experimental questions, methods, evidence states and verdicts, including negative results. We do not publish the mechanisms, implementation details or generalisations that constitute proprietary research.
The line this draws: you can inspect whether we're entitled to a claim without being entitled to reproduce the machinery that produced it.
Watch the apparatus →
Copperhollow runs live, in public — the settlement the containment results came from.
Explainers →
Teaching demos of the mechanisms we use — runnable in the browser, not findings.
The status map →
Everything the lab runs — systems, demonstrations, transferred technology — honestly labelled.