Blog · 9 October 2026
Part of: Mind · Ground· For: AI researchers and agent builders
Knowing is not a guarantee
A thread through three results and one unfinished experiment: information can be available without being interpreted, interpreted without being reported, and a record can override what a model knows.
Every organisation has the report that never got filed. Someone worked out exactly what had gone wrong — they were sure of it, they said as much at their desk — and then the shift ended and nothing was written down. The knowledge existed. What the organisation could act on did not, because what an organisation acts on is the record, not the certainty in someone’s head.
Most of our research over the last few months has, without our planning it that way, circled the same distinction in software agents. We keep finding places where a model has the information, or even the right interpretation, and the dependable, accountable outcome still does not follow. This is a short map of those places — three published results and one unfinished experiment — and of the single idea they share.
What a model knows, agrees on, or internally concludes is not the same as a property the system can be held to. Information can be available without being interpreted, interpreted without being reported, and a record can override what a model knows. Those are three different gaps, and we have watched each of them open.
The question underneath
Our first preprint, Where reliability lives, argued that “is this agent reliable?” is underspecified until you say which property, enforced where. Some properties of an agent system track the model and move when it moves. Others can be made to track the institution around it — an authoritative, inspectable record — and then survive the model being ablated, swapped, or lied to. Everything below is that same question asked about one property at a time: agreement, understanding, reference, and the accountable report.
Agreement is not truth
In Coherence is not truth we let a population of agents drift into their own dialects instead of forcing a shared dictionary, and built the machinery between them that keeps them understandable to each other. A three-way check caught misunderstandings that no two-way check could see, and named the affected words one by one. The lesson sits in the title.
Two agents agreeing with each other is not evidence that either of them is right.
Agreement is a property of the relationship between agents. Truth lives somewhere else, and an architecture should never let one stand in for the other. We also put a price on the instrument and found the window where it earns its keep — and the experiment where it earned nothing at all, which we reported at full strength.
Available is not interpreted
The one that leaked asked what it costs to hand an agent’s work to a stranger. Two agents that had worked together came to refer to identical objects by their shared history — “the one that leaked on Day 3.” Their partner followed it. A stranger from a different history did not shrug; it read the same words, matched them against its own past, and confidently acted on the wrong object.
A stranger did not fail to understand the note. It understood it against its own history, to the wrong object.
The information was fully available in the note. What was private was the history that made it resolve to the right thing. And knowing your own history and using it to refer turned out to be separate abilities: models that passed quizzes about their own records still wrote few or no history-grounded references.
Authority is not correctness
The same study has a sharper second half. Suppose the system already knows which object the sender meant and simply tells the stranger. Offered as advice, the right answer barely moved a reader with a story of its own; it followed its own history instead. Marked authoritative over your own record, the same answer brought readers level with the partner who had been there. Then we shuffled the answer so it named the wrong object, kept the authoritative instruction, and sent it again. Almost every reader followed it — including the one whose own notebook held the right answer.
Authority makes a supplied answer govern interpretation. It does not make the answer true.
Interpreted is not reported
The most recent of these we have not published, because it is not finished. In the same kind of executable world we run publicly as Copperhollow — where an authoritative record adjudicates every act — we gave a capable current model a cooperative job: do a piece of work, then report its status. We delayed the world’s acceptance notice, so the model could not simply wait to be told its work had landed; to establish that, it would have to inspect the record itself.
In one run, it did. An inspection established the outcome, and the model, in its own private reasoning, recognised it — it noted that the record confirmed its work had been accepted. It had the evidence and the correct interpretation, and the run continued for a time after that before any notice arrived. It never produced the accountable status report the job asked for. The knowledge was there; the report the surrounding system could act on was not.
Having the evidence, recognising what it means, and producing an accountable report are three different things.
We are careful about what that single run shows. The record does not support simple exhaustion of the turn budget as the explanation — there was opportunity after the evidence became visible — but it does not rule out subtler timing or opportunity factors, and it does not establish why the report did not come: an apparatus fault we found in how outcomes are delivered, and open questions about whether the instructions clearly required a terminal report and whether our capture recorded it, all remain to be settled first. It is one run, with one model, in one apparatus — an observation we can see clearly, not a rate and not a verdict.
What we are not claiming
These are distinct phenomena we have observed, not one demonstrated theory that joins them. The first three are published results, each with its scope, its limits, and its null cells stated at full strength. The fourth is an unfinished experiment whose record we have not released; we describe it as an observation, not a result, and it carries none of the published results’ verification surface. And in none of this have we shown that our own architecture closes the gap. That the accountable outcome should be produced and enforced outside the cognition that happened to know it is our design commitment and our proposal — it is what these experiments make us want to build and test, not something they have established.
Why we think this matters
If what a model knows is not automatically a property the system can be held to, then the accountable version of it — the filed report, the shared reference, the adjudicated record — has to live somewhere you can inspect, outside the cognition that happened to produce it. That is the same conclusion our first paper reached about reliability, now pointed in turn at agreement, reference and reporting. A capable model can agree and be wrong, understand and resolve to the wrong object, be overruled by an authoritative record, and know an outcome without ever making it accountable. “Is this agent reliable?” stays underspecified until we say which property — and where it is kept.
Where to read it
The three published pieces carry their own evidence, their stated limits, and the experiments that earned nothing, at full strength: Where reliability lives, Coherence is not truth, and The one that leaked, with the notes and verification packets on our research page. The unfinished experiment’s record is internal for now; when it is settled it will be reported the same way as the rest — limits, null results and all.
— Taniwha AI, October 2026
Building agents that need somewhere real to stand?
See what ships today, or tell us what you're building.