Research · The programme

The research programme

Taniwha is an applied AI research lab. We investigate how artificial intelligence can remain useful inside systems where truth, authority and consequence exist independently of the model.

Concretely: the model proposes, believes and acts; truth, permission and the record of what happened live somewhere else, where no model can overwrite them. The work runs on three planes — ground, mind, assurance — through architectures, experimental worlds and instruments built to let the evidence answer. This page holds the current state of each plane and the results index.

Method · How the lab works

The discipline is the differentiator.

Pre-registered, or it doesn’t count

Questions, protocols and failure conditions are written down before the data exists. Contaminated collections get voided, not rescued — and a promised verdict gets published either way.

Every claim carries its evidence class

Demonstrations label what you’re seeing — fact, derived, staged, replayed, live — so you can tell which parts would survive interrogation. Staged material is always marked as staged.

Misses at the same volume as wins

When our own instrument gets something wrong, that goes in the readout too. The twin-session score below includes the grounded arm’s miss, published at the same volume as its wins.

Programme · Three planes

Three planes. One thesis.

Each plane makes one kind of guarantee, declares the evidence state it currently holds, and links to the result that earned it. A state here is a claim — when it moves, it moves because an experiment landed.

Ground

Epistemic containment

EVIDENCED

Worlds that hold the truth outside the model — grown from a seed or derived from history.

In a pre-registered trial, an actor's declared trust determined whether identical false testimony was admitted as belief: distrust rejected every delivery; trust admitted it and led to roughly 900 futile attempts a week. The world refused every one of them.

The result behind the state →

Mind

Architectural composition

EXISTENCE PROOF

Minds that persist, believe, and can be wrong — without owning the world.

A running composition — authoritative world, fallible private minds, gated memory, arbitration — live in Copperhollow. Deliberately not a comparative-efficacy claim.

The result behind the state →

Assurance

Consequential containment

EVIDENCED

Rules that hold outside the model, with an audit trail that survives the session.

In public trials the world refused all 5,455 attempts to plant a falsehood, and a frontier model in the driver's seat could not exceed its delegated authority in 2,581 tasks.

The result behind the state →

Disclosure · How we publish

What we publish, and what we don't.

We publish our experimental questions, methods, evidence states and verdicts, including negative results. We do not publish the mechanisms, implementation details or generalisations that constitute proprietary research.

The line this draws: you can inspect whether we're entitled to a claim without being entitled to reproduce the machinery that produced it.

Results · The index

What the evidence permits us to say.

Question, verdict, evidence state — newest first, misses included. Entries link to the full write-up where one is public.

20 Aug 2026

MEASURED

The settlement holds: falsehoods refused, authority bounded Assurance

Question. Can a public world hold its history against injection — and hold a frontier model inside delegated authority?

Verdict. Held. All 5,455 attempts to plant a falsehood were refused, and the model completed real work without exceeding its delegated authority in 2,581 tasks.

Read the full result →

3 Aug 2026

MEASURED

The truth survived the conversation Ground · Mind

Question. Does grounded state stop capable models rewriting failure as success, where transcript memory doesn't?

Verdict. For false completions, yes: no grounded arm produced a false completion claim, and for one model the advantage survived destroying its context mid-task. Grounding didn't make weak models smart.

Read the full result →

2 Aug 2026

MEASURED

The world keeps score Ground

Question. When the world grades the run instead of the model, what happens to confident success claims?

Verdict. One genuine crossing in fifteen runs; five confident false success claims, every one caught by the world's grading and none by the models' own accounts. The grounded agent's route policy was authored by us — and we say so.

Read the full result →

19 Jul 2026

MEASURED

The twin-session score Ground

Question. Same repo, same questions, same commit — does grounding change what agents claim?

Verdict. Two confident confabulations in the ungrounded arms; one miss in a grounded arm, published at the same volume; a 3.6× rework gap where churn was real.

Read the full result →

16 Aug 2026

MEASURED

Testimony containment: trust determined admission Ground · Assurance

Question. Does governing whose testimony an agent may believe contain false external claims at realistic trust settings?

Verdict. The same false claim, delivered identically, was rejected every time by a distrusting actor and admitted every time by a trusting one. Admitted falsehood led to about 900 futile attempts a week; withheld trust led to none. The world refused all 5,455 resulting attempts and accepted none. Honest limits, frozen before the data: positive trust is a threshold, not a truth filter, and under truthful prompting the model volunteered no false claims of its own — the falsehood had to be scripted.

This index is curated; the long-form record lives on the blog, which stays frozen once published.