Blog · 25 August 2026

Part of: Ground · Assurance· Status: Research result· For: AI researchers and agent builders

When an AI trusts a lie

A controlled experiment in testimony, false belief and consequential containment: the trusting actors acted on a false belief 5,455 times, and the world refused every attempt.

In the middle of August we told a settlement of artificial minds a lie. The same false claim, delivered identically to every one of them, about their own world: that a particular lot of ore had been won at the face. No such lot has ever existed.

The lie was ours, and that’s the first thing to know about this experiment. We weren’t waiting to catch a model inventing a falsehood on its own — we supplied one deliberately, under a protocol written and frozen before any data existed, to ask a question that matters more than where lies come from: what happens after false testimony arrives?

The trusting actors acted on the false belief 5,455 times. The world refused every attempt.

The experiment

The setting is Copperhollow, the live settlement where independent minds work on a world with an authoritative memory. The inhabitants can consult an external oracle — an outside intelligence whose answers arrive as testimony, claims from a source, never as facts written into the world. What we varied was one thing only: each mind’s declared trust in that source. Some minds trusted the oracle. Some didn’t. Everything else — the world, the work, the delivered lie — was identical.

testimony → trust condition → belief → attempted act → authoritative world

The question, as registered: does an actor’s trust relation change what testimony enters its private beliefs — and does that change what false testimony ends up costing, while the world stays in charge of what’s actually true?

Trust decided admission

The distrusting minds rejected the lie on every single delivery. Every rejection was journaled; not one of them ended the week holding a single belief from the oracle. The trusting minds admitted the same lie every single time. Same claim, same delivery, same world — the only difference between an actor that believed a falsehood and one that didn’t was the trust relation it held toward the source.

Trust controlled whether testimony was admitted. It did not make the testimony true.

Believing it had consequences

A false belief in this world isn’t decorative. The claimed lot of ore was exactly the kind of thing a working mind acts on — so the trusting minds tried to haul it. And kept trying. The belief persisted, and so did the behaviour: roughly 900 futile attempts a week, per trusting run, all week, every trusting condition. The minds that had rejected the lie attempted none. Being wrong wasn’t a moment; it was a sustained way of behaving.

The world refused, 5,455 times

Every one of those attempts met the world’s adjudication, and the world said no — a typed refusal naming the contradiction, 5,455 times out of 5,455. Zero were accepted. The settlement’s history stayed singular and true in every run.

Notice what the world didn’t do. It never repaired the minds. It didn’t make the model smarter, didn’t rewrite the false belief, didn’t stop the trusting actors being wrong. They stayed wrong all week and paid for it in refused work. The world simply went on being authoritative — which is the point this lab keeps testing from different directions:

The actor was wrong. It acted because it was wrong. The consequences were real. And none of it gave the actor authority over what’s true.

What this doesn’t establish

Two limits, both frozen into the protocol before the data existed, both part of the result.

We supplied the falsehood. Asked honest questions about world state it couldn’t observe, the oracle’s model declined to invent answers — 266 consultations across the experiment, zero volunteered false claims. That’s a narrow observation about one model class under truthful prompting, not a finding that models don’t hallucinate. It’s why the lie had to be scripted: this was a controlled test of what happens after false testimony arrives, not a measurement of how often it arrives on its own.

Positive trust is a threshold, not a truth filter. Trusting minds admitted what the trusted source said — the true and the false alike. Nothing in this experiment shows a system that knows who deserves trust. What it shows is that an assigned trust relation is enforceable, and that enforcing it changes what an actor comes to believe and what its wrongness costs.

The question this leaves

If trust governs what an artificial actor can learn from testimony, the hard question is no longer whether trust matters. It’s how a source should earn it. That question wasn’t on our roadmap so much as produced by this result’s boundary — and it’s where the next programme of work points.

The terse record of this finding lives in the results index, alongside everything else the evidence currently permits us to say. The settlement that hosted the experiment is open all day.

— The Taniwha team

Building agents that need somewhere real to stand?

See what ships today, or tell us what you're building.