Blog · 19 July 2026
Part of: Ground · Pono· Status: Active development· For: Teams running coding agents
Update: the twin-session score is in
The pre-registered grounded-vs-ungrounded experiment has its first three pairs. Short version: grounding prevented confident false claims — and caught one of our own.
On July 13 we published a promise inside “We charged our own tool rent”: the next test of Pono would be pre-registered — frozen questions, identical twin sessions, a failure condition we’d publish either way. On July 19 the first three qualifying pairs ran. This is a pointer, not a new argument: the full score now lives as an Update on that post, under the same registration it promised.
The short version:
Ungrounded twins invented facts about our own codebases — confidently. A crate declared “load-bearing” that nothing imports; an open architectural decision declared “resolved and binding,” attributed to the wrong document. Grounded twins made one false claim — a provenance detail — and it’s published at the same volume. Where a 38-commit backlog of unseen changes had built up, the ungrounded twin needed 3.6× the tool calls to reach the same answers.
Three pairs, one machine, one model — a score, not a law. If you want the fuller picture: the verdict post with the July 19 Update, or what Pono is.
— The Taniwha team
Building agents that need somewhere real to stand?
See what ships today, or tell us what you're building.