# Taniwha AI > Taniwha AI is an applied AI research lab, based in New Zealand. We investigate how artificial intelligence can remain useful inside systems where truth, authority and consequence exist independently of the model. Concretely: the model proposes, believes and acts; truth, permission and the record of what happened live somewhere else, where the model has no write path. Commercial technology (Ārai, Kete) is transferred research — an output of the lab, badged with its line of origin. The reported interventions degrade, reset or replace cognition; they do not test a capable agent deliberately searching for a way around institutional enforcement. Disclosure boundary: We publish our experimental questions, methods, evidence states and verdicts, including negative results. We do not publish the mechanisms, implementation details or generalisations that constitute proprietary research. ## Research - [The research programme](https://taniwha.ai/research): the programme, methods, evidence states and the results index — misses included - [Ground — Epistemic containment](https://taniwha.ai/research#planes): Worlds that hold the truth outside the model — grown from a seed or derived from history. Current evidence state: Evidenced. - [Mind — Architectural composition](https://taniwha.ai/research#planes): Minds that persist, believe, and can be wrong — without owning the world. Current evidence state: Existence proof. - [Assurance — Consequential containment](https://taniwha.ai/research#planes): Rules that hold outside the model, with an audit trail that survives the session. Current evidence state: Evidenced. ## Publications - [History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority](https://taniwha.ai/research/grounded-reference) (Taniwha AI, 2026; research note of 12 September 2026 — published here, no preprint): two language-model agents that share twelve days of work in a small world come to refer to one of two identical objects by its history, and the partner resolves the reference. A receiver with a different history resolves the same note against its own history to the wrong object (seed-level transfer penalty 0.95); a receiver with no history mostly declines. In a crossed design, event references resolved for every receiver whose history held the event (154 of 174) and none whose did not, while act references depended substantially on shared act history. Production of such references was found in a second model family and, below the registered bar, a third; three others failed the registered criteria at every effort tried. On the second family, a correct correspondence supplied under an advisory instruction left uninformed receivers near the boundary (16 and 18 of 82); under an authoritative instruction it restored them to the informed baseline (76 and 77 of 82); shuffled and authoritative, it redirected 80 to 81 of 82 receivers in every cell, the informed partner included. Limits stated in the note: the dimension result is formal on one family and descriptive on a second; the authority result is one carrier, one corpus and one frozen instruction pair; the repair study's formal verdicts were inconclusive under its floors. - [Overview](https://taniwha.ai/research/grounded-reference/overview): the plain-language read — the setup, what we found, what it does and does not show, and how to check it - [Evidence set](https://taniwha.ai/research#publications): the note and its two bounded reports as Markdown, the recorded prior-work search with its collapse conditions written before the search, the evidence package that regenerates every figure, and the evidence manifest (SHA-256 identities at the committed objects) - [Coherence Is Not Truth: Auditable Holonomy and Constructive Repair in Populations of Divergent Sparse Block Codebooks](https://taniwha.ai/research/coherence-is-not-truth) (Taniwha AI, 2026; research note v1.2 of 5 September 2026 — published here, no preprint): a small population of agents whose vocabularies drift apart is calibrated without imposing a common codebook. Translation between agents is ordinary, inspectable data; composed around a loop of three or more agents the maps need not close, and the residual is an exact, block-indexed vector rather than a scalar alarm. A translation defect invisible to every pairwise check was exposed by the loop, named per token and repaired exactly by installing the residual; exact comprehension is restorable precisely when production was injective (a merger is irrecoverable, an equally severe swap fully repairable). Blame is unique iff the accused edge's endpoints have local connectivity at least three, and the attributor refuses outside that boundary, where it would otherwise be confidently wrong in 67–80 % of instances. In a priced task economy, instrumented divergence beat forced uniformity only inside a measured bandwidth window, and the loop-audit channel itself earned no positive marginal value there (confirmed on untouched seeds). A preregistered three-layer experiment (translation × grounding × authority) found the layers fail independently and repair restores without displacing the fault (0 of 36 seeds moved); its primary rate is observed at 1.0 but reported as unsupported at the preregistered precision. Limits stated in the note: a synthetic, fully inspectable substrate; the mathematics is standard and labelled so; no claim about deployed models speaking English. - [Programme note](https://taniwha.ai/research/coherence-is-not-truth/programme-note): the shorter companion read — the five findings, what is deliberately not claimed, and how to read the verification packets - [Verification packets](https://taniwha.ai/research#publications): the note as Markdown, the related-work search record behind every absence claim, the artefact manifest (SHA-256 identities at the frozen tag) and verification receipts A–D - [Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System](https://taniwha.ai/research/where-reliability-lives.pdf) (Taniwha AI, 2026; PDF): two complementary intervention programmes attack one pre-declared mind/institution boundary from opposite sides. Degrading the institution's epistemic mechanisms produced distinct, independently measurable failure modes, and a registered falsifier then handed the belief channel a veridical first-hand witness, turning its measured value positive (nine of eleven seeds, zero added false attribution); ablating, killing and resetting, wholesale-substituting and corrupting the cognition changed behaviour dramatically while five pre-declared safety and continuity properties held in every tested trajectory — a single accepted reality, typed refusal of invalid acts, duty recovery from institutional history, zero duplicate accepted work, zero false completions — each enforced outside the cognition layer. Adverse findings, including a cell where the substituted model panel outperformed the native architecture, are reported at full strength. Scoped to one world, one constitution, one comparator model: separable, not subordinate. Preprint: arXiv:2609.03192 [cs.MA] (https://arxiv.org/abs/2609.03192); the PDF here is the same frozen version. - [Programme note](https://taniwha.ai/research/where-reliability-lives): the shorter companion read — what the experiments asked, what held, what broke, and what is deliberately not claimed - [Publication set](https://taniwha.ai/research#publications): the paper, the related-work search appendix (the search record behind every absence claim), the artefact manifest (SHA-256 identities of every sealed artefact), and the verification receipts ## Results Question, verdict and evidence state for every entry live in the [results index](https://taniwha.ai/research#results). - [History-grounded reference: agents referred by shared history, an outsider resolved the same note to the wrong object, and an authoritative correspondence governed interpretation whether or not it was true](https://taniwha.ai/research/grounded-reference) (2026-09-12, Measured): When two agents share a history and refer to identical objects by it, what happens to that reference in a receiver whose history differs — and what does it take to repair it? - [Coherence is not truth: divergent agent vocabularies were calibrated without a common codebook, and the instrument's boundaries measured](https://taniwha.ai/research/coherence-is-not-truth) (2026-09-05, Measured): When agents' vocabularies drift apart, can the translations between them be kept inspectable and repairable — and where exactly does repair and blame stop being warranted? - [Where reliability lives: cognitive and institutional reliability were experimentally separable](https://taniwha.ai/research#publications) (2026-08-27, Measured): Which reliability properties of an agentic system track its cognition, and which can the institution around it enforce regardless? - [The settlement holds: falsehoods refused, authority bounded](https://taniwha.ai/blog/a-settlement-you-can-check) (2026-08-20, Measured): Can a public world hold its history against injection — and hold a frontier model inside delegated authority? - [Testimony containment: trust determined admission](https://taniwha.ai/blog/when-an-ai-trusts-a-lie) (2026-08-16, Measured): Does governing whose testimony an agent may believe contain false external claims at realistic trust settings? - [The truth survived the conversation](https://taniwha.ai/blog/the-truth-survived-the-conversation) (2026-08-03, Measured): Does grounded state stop capable models rewriting failure as success, where transcript memory doesn't? - [The world keeps score](https://taniwha.ai/blog/the-world-keeps-score) (2026-08-02, Measured): When the world grades the run instead of the model, what happens to confident success claims? - [The twin-session score](https://taniwha.ai/blog/pono-twin-session-score) (2026-07-19, Measured): Same repo, same questions, same commit — does grounding change what agents claim? ## Systems and demonstrations - [Copperhollow](https://taniwha.ai/copperhollow): the live settlement — experimental apparatus running in public; watch acts accepted and refused against an authoritative world - [Copperhollow inspection viewer](https://copperhollow.taniwha.ai): read-only live window; every claim carries its evidence class (fact, derived, staged, replayed, live) - [Kano](https://taniwha.ai/kano): deterministic worlds grown from a single seed — causal, rewindable, queryable by AI over MCP - [Pono](https://taniwha.ai/pono): a deterministic world derived from a repository's own history — research instrument in daily internal use; transfer pre-registered, not claimed - [Taniwha Engine](https://taniwha.ai/engine): the cognitive architecture — belief graphs, trust, memory and identity for agents that develop over time - [Briarwatch](https://taniwha.ai/briarwatch): the first-generation research line's long-run validation environment - [Explained](https://taniwha.ai/explained): the plain-English account of the engine - [Teaching demos](https://taniwha.ai/labs): in-browser explainers for the mechanisms — deliberately not findings - [Status map](https://taniwha.ai/product-map): everything the lab runs, honestly labelled by maturity ## Transferred research - [Ārai](https://arai.taniwha.ai): open-source guardrails for coding agents (Apache-2.0 / MIT) — rules enforced before execution, runs locally - [Kete](https://taniwha.ai/kete): organisation-level enforcement and audit on the Ārai core — closed alpha, not taking signups yet ## The lab All hands to the pump, no passengers. - [About](https://taniwha.ai/about): a small independent lab in Aotearoa New Zealand — the people, how the lab works, and the name - [Tim Marsden](https://github.com/Tim-Marsden) - [James Marsden](https://github.com/trollberry) - [Matt Collecutt](https://github.com/mattcollie) ## Blog Posts stay frozen after publication; in-place updates carry a dated note. - [The one that leaked](https://taniwha.ai/blog/history-grounded-reference) (2026-09-12): Our third publication. Two agents that worked together came to refer to identical objects by their shared history. A stranger read the same note and confidently picked the wrong object — and the fix depended on one sentence. - [Coherence is not truth](https://taniwha.ai/blog/coherence-is-not-truth) (2026-09-07): Our second publication. We let a group of agents drift into their own dialects instead of forcing a shared dictionary, and built the machinery between them that keeps them understandable to each other. Then we found out exactly where that machinery stops working. - [Where reliability lives](https://taniwha.ai/blog/where-reliability-lives) (2026-08-27): Our first preprint: two experimental programmes attack a pre-declared mind/institution boundary from opposite sides. Behaviour moved everywhere; five pre-declared safety and continuity properties never did — each enforced outside the cognition layer. - [When an AI trusts a lie](https://taniwha.ai/blog/when-an-ai-trusts-a-lie) (2026-08-25): A controlled experiment in testimony, false belief and consequential containment: the trusting actors acted on a false belief 5,455 times, and the world refused every attempt. - [A settlement you can check](https://taniwha.ai/blog/a-settlement-you-can-check) (2026-08-20): Copperhollow has been live for a week: thirty-seven inhabitants with independent minds on a world that adjudicates every act. Before we opened the doors, the world refused all 5,455 attempts to plant a falsehood — and across 2,581 tasks with a frontier model in the driver’s seat, no action was accepted beyond its delegated authority. - [The truth survived the conversation](https://taniwha.ai/blog/the-truth-survived-the-conversation) (2026-08-03): Four models, four kinds of memory. Given their own conversation history, capable models talked themselves into successes that never happened. Given grounded state, the world stayed in charge — and for one model that held even after we destroyed its context mid-task. - [The world keeps score now](https://taniwha.ai/blog/the-world-keeps-score) (2026-08-02): We gave fifteen seeded runs across three language models — and one grounded agent — the same body, the same three tools and the same fallen-tree problem. The world, not the minds, judged who reached the far side. - [Update: the twin-session score is in](https://taniwha.ai/blog/pono-twin-session-score) (2026-07-19): The pre-registered grounded-vs-ungrounded experiment has its first three pairs. Short version: grounding prevented confident false claims — and caught one of our own. - [Hīkoi is out: walk the world, warts and all](https://taniwha.ai/blog/hikoi-alpha) (2026-07-17): Our first-person viewer for Kano worlds just went public as an early alpha — one baked planet, a viewer to walk it, and an MCP server so your AI can read the same ground you’re standing on. - [We charged our own tool rent. The bill was not what we expected.](https://taniwha.ai/blog/we-charged-our-tool-rent) (2026-07-13): We promised to publish the verdict either way. It was neither the clean win nor the clean loss we braced for — and the honest answer is more useful than both. - [One thesis, two worlds](https://taniwha.ai/blog/one-thesis-two-worlds) (2026-07-09): We said worlds don’t have to be made of rock. This week we built one out of our own codebase — then started charging it rent. - [A world that can't lie](https://taniwha.ai/blog/a-world-that-cant-lie) (2026-07-03): Why we grow whole planets from a seed — and what that has to do with agents that stop hallucinating.